Audio Fingerprinting With Local Normalization for Spectral Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio fingerprinting technologies often rely on the loudest parts of an audio signal, which can be associated with noise, leading to incomplete fingerprints that do not include samples from all parts of the audio spectrum, particularly higher frequency ranges.
Innovation Solution
The method involves normalizing time-frequency bins of an audio signal by an audio characteristic of the surrounding region, such as mean energy, and generating a fingerprint by selecting points from normalized audio signal frequency components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio fingerprinting relies on the loudest parts of the audio signal, then the fingerprint generation is simple and fast, but the fingerprint completeness and spectral coverage deteriorate
Solution Approach 1:
The patent changes the selection criterion from absolute energy level (loudest parts) to normalized energy distribution (parts above local threshold). This parameter transformation allows selection of spectrally diverse components while maintaining processing efficiency, resolving the contradiction between speed and completeness.
Solution Approach 2:
The patent applies local normalization by comparing each time-frequency bin's energy to the mean energy of its surrounding region. This local quality assessment ensures that fingerprints capture spectral variations across different regions rather than globally selecting only the loudest parts, improving spectral coverage while maintaining computational efficiency.
2Device complexity
If audio fingerprinting selects only the loudest parts of the audio signal, then the processing complexity is reduced, but the spectral coverage and identification accuracy deteriorate
Solution Approach 1:
The patent transforms the selection parameter from absolute energy magnitude to relative energy prominence (exceeding local mean). This change maintains computational simplicity while improving identification accuracy by ensuring selection of spectrally representative components across the entire frequency spectrum.
Solution Approach 2:
The patent adds a spatial dimension to the selection process by considering the distribution of energy across the time-frequency plane and selecting components that are prominent in their local regions. This dimensional approach ensures comprehensive spectral coverage without significantly increasing processing complexity.
3Loss of information
If audio fingerprinting uses normalization by surrounding region mean energy, then the spectral distribution evenness improves, but the computational complexity increases
Solution Approach 1:
The patent applies local normalization using the mean energy of surrounding time-frequency bins as a reference. This local quality metric efficiently captures spectral distribution characteristics without requiring global processing, achieving even spectral coverage with moderate computational overhead.
Solution Approach 2:
The patent computes normalization based on a limited surrounding region (local window) rather than the entire audio spectrum. This partial action approach provides sufficient spectral coverage information while significantly reducing computational complexity compared to global normalization methods.
Data Source
AI summary
Methods, apparatus, systems, and articles of manufacture are disclosed to fingerprint audio via mean normalization. An example apparatus for audio fingerprinting includes a frequency range separator to transform an audio signal into a frequency domain, the transformed audio signal including a plurality of time-frequency bins including a first time-frequency bin, an audio characteristic determiner to determine a first characteristic of a first group of time-frequency bins of the plurality of time-frequency bins, the first group of time-frequency bins surrounding the first time-frequency bin and a signal normalizer to normalize the audio signal to thereby generate normalized energy values, the normalizing of the audio signal including normalizing the first time-frequency bin by the first characteristic. The example apparatus further includes a point selector to select one of the normalized energy values and a fingerprint generator to generate a fingerprint of the audio signal using the selected one of the normalized energy values.


