Audio Signal Fingerprinting With Exponential Mean Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Historical audio fingerprinting technologies often generate fingerprints based on background noise and not the audio of interest, particularly in the bass spectrum, leading to reduced effectiveness and increased memory and bandwidth requirements.
Innovation Solution
Generate fingerprints using exponential mean normalization, normalizing time-frequency bins based on adjacent audio regions' average characteristics, and comparing query fingerprints to reference fingerprints with adjusted sample rates to improve matching accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio fingerprinting methods are used to capture bass spectrum information, then fingerprinting coverage is improved, but memory and bandwidth requirements increase
Solution Approach 1:
The patent extracts only the essential spectral features needed for accurate fingerprinting while discarding redundant information. By using spectral flux to identify changing spectral components and focusing only on those regions, the system achieves accurate fingerprinting without processing or storing unnecessary data, thereby reducing memory and bandwidth requirements.
Solution Approach 2:
The patent applies different processing strategies to different frequency regions based on their characteristics. Instead of uniformly processing the entire spectrum, the system identifies and focuses on specific spectral regions that contain relevant audio information, applying localized analysis only where needed. This selective approach maintains fingerprinting accuracy while reducing overall computational and storage resources required.
2Ease of manufacture
If traditional audio fingerprinting methods are used, then fingerprint generation is simplified, but fingerprints are generated from background noise rather than audio of interest
Solution Approach 1:
The patent performs preliminary analysis of the audio signal to identify spectral regions containing actual audio content before generating fingerprints. By using spectral flux to detect changing spectral components and pre-identifying relevant frequency regions, the system ensures that fingerprints are generated from meaningful audio information rather than background noise, while maintaining an efficient streamlined process.
Solution Approach 2:
The patent introduces spectral flux as an intermediary measure to bridge the gap between simple fingerprint generation and accurate audio content identification. Spectral flux acts as a mediator that highlights changing spectral regions, allowing the system to automatically distinguish between relevant audio content and background noise without complex processing, thus maintaining simplicity while improving accuracy.
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
Methods, apparatus, systems, and articles of manufacture are disclosed to fingerprint an audio signal via exponential normalization. An example apparatus includes an audio segmenter to divide an audio signal into a plurality of audio segments including a first audio segment and a second audio segment, the first audio segment including a first time-frequency bin, the second audio segment including a second time-frequency bin, a mean calculator to determine a first exponential mean value associated with the first time frequency bin based on a first magnitude of the audio signal associated with the first time frequency bin and a second exponential mean value associated with the second time frequency bin based on a second magnitude of the audio signal associated with the second time frequency bin and the first exponential mean value. The example apparatus further includes a bin normalizer to normalize the first time-frequency bin based on the second exponential mean value and a fingerprint generator to generate a fingerprint of the audio signal based on the normalized first time-frequency bins.