Audio Fingerprint Indexing via Time-Variant Spectrogram Transforms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio signal identification techniques fail to accurately identify noisy or distorted audio signals due to noise masking the underlying signal, leading to false negatives and inefficient processing requirements.

Innovation Solution

A system generates noise- and distortion-insensitive audio fingerprints by applying a time-variant transformation to the frequency spectrums of overlapping audio frames, selecting low-frequency components for indexing, and using phase differences to enhance matching accuracy and reduce computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional audio fingerprinting techniques are used to identify audio signals, then identification speed is improved through index-based searching, but noise and distortion in the signal cause false negatives and reduce identification accuracy

Engineering Contradiction:
Improveidentification speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transforms the audio signal from time domain to frequency domain using Fourier transform, then applies Mel-scale filtering to convert linear frequency to perceptual frequency. This parameter transformation allows the system to capture noise-insensitive features in the frequency domain that remain stable even when noise is present, thereby maintaining both fast index-based searching and high identification accuracy

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediate representation called audio fingerprint that extracts essential spectral features from the original audio signal. This fingerprint serves as a mediator between the raw audio signal and the identification process, capturing only the most discriminative features while filtering out noise. The fingerprint can be efficiently indexed and searched, enabling both fast retrieval and accurate matching even in noisy conditions

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the signal to noise ratio is very low (less than −6 dB), then noise completely masks the signal making identification difficult, but conventional techniques still attempt to match noisy portions resulting in false negatives

Engineering Contradiction:
Improvesignal detection accuracyVSAvoidnoise masking effect
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent converts the harmful effect of noise by transforming to the frequency domain where noise and signal occupy different spectral regions. The Mel-scale filtering and spectral feature extraction process selectively captures signal-dominated frequency regions while ignoring noise-dominated regions. This transforms the noise from a masking factor into a filterable component, allowing accurate signal identification even when noise completely masks the time-domain signal

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent extracts only the most discriminative spectral features from the audio signal to create the fingerprint, rather than using the complete signal. By selecting only the most stable and informative frequency components (through Mel-scale filtering and spectral peak detection), the system achieves accurate identification with minimal signal data, effectively ignoring the noisy portions that would otherwise prevent identification

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If tempo shifting occurs when audio signal is played faster or slower than original speed, then spectral content shifts along time axis causing noise to increasingly mask the original signal

Engineering Contradiction:
Improvetempo variation handlingVSAvoidnoise masking with tempo shift
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent uses dynamic time warping and elastic matching algorithms that can adapt to tempo variations. Instead of requiring exact temporal alignment, the system dynamically adjusts the matching process to accommodate speed changes. The fingerprint matching algorithm is designed to be elastic, allowing temporal stretching or compressing of the query fingerprint to match reference fingerprints at different playback speeds, thereby maintaining identification accuracy across tempo variations

Inventive Principle:
Principle #15Dynamics

4Speed

If index-based selection of reference fingerprints is used for matching, then searching speed is improved, but noise and distortion in the index cause failures to match against database indexes

Engineering Contradiction:
Improvesearching speedVSAvoidindex matching accuracy
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent pre-processes the audio signal to create a compact fingerprint representation that captures essential spectral features before the matching process. This preliminary action creates a noise-resistant summary of the signal that can be efficiently indexed. The fingerprint is designed to be stable under various transformations, ensuring that index-based searching remains both fast and accurate even when the original signal contains noise or distortion

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10418051B2Indexing based on time-variant transforms of an audio signal's spectrogram
Publication Date: 2019.09.17 META PLATFORMS INC
  • US10418051B2 patent drawing
  • US10418051B2 patent drawing
  • US10418051B2 patent drawing

AI summary

An audio identification system generates audio fingerprints and indexes associated with the audio fingerprints based on discrete and overlapping frames within a sample of an audio signal. The system applies a time-to-frequency domain transform to a time-sequence of frames, which may be filtered. The audio identification system then applies a time-variant transformation (e.g., a Discrete Cosine Transform) to the transformed frames and generates an audio fingerprint and index by selecting sets of coefficients of the time-variant transformation. The system selects coefficients that are less sensitive to possible noise and/or distortions in the underlying signal, such as low-frequency coefficients. The time-variant transformation provides sufficient sampling among the indexes by incorporating the phase information of the frames into the indexes. The system stores the audio fingerprint and other identifying information by index for efficient retrieval and matching of the retrieved fingerprints.