Audio Fingerprinting Via Mean Normalization Across The Spectrum
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio fingerprinting technologies often fail to accurately identify audio signals due to reliance on loudest parts, which can be noise, and neglect higher frequency ranges, leading to incomplete fingerprints.
Innovation Solution
Generate audio fingerprints by normalizing time-frequency bins and audio signal frequency components using mean energy or other characteristics, and select points for fingerprint creation, ensuring inclusion of all audio spectrum parts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If audio fingerprinting relies on the loudest parts of the signal, then processing is simplified, but identification accuracy deteriorates because loud parts may be noise rather than meaningful audio content
Solution Approach 1:
The patent changes the selection criterion from amplitude-based (loudest parts) to spectral normalization-based (parts deviating from mean spectral energy). This parameter change allows the system to identify meaningful audio content regardless of overall volume levels, resolving the contradiction between processing simplicity and identification accuracy.
Solution Approach 2:
Instead of selecting audio parts based on maximum energy (conventional approach), the patent inverts the logic by selecting parts that deviate most from the mean spectral energy after normalization. This inversion enables the system to capture meaningful content that may not be the loudest but is spectrally distinctive.
2Device complexity
If audio fingerprinting focuses on lower frequency ranges, then processing is easier, but higher frequency information is lost resulting in incomplete fingerprints
Solution Approach 1:
The patent applies spectral normalization across the entire frequency spectrum, transforming the audio signal so that each frequency component is normalized by the mean energy of its surrounding frequencies. This enables uniform processing of all frequency ranges while preserving high-frequency information that would otherwise be lost.
Solution Approach 2:
The patent transitions from time-domain or simple frequency-domain analysis to a normalized spectral domain representation. By introducing spectral normalization as an additional processing dimension, the system can efficiently process the entire frequency spectrum without increasing overall complexity, as the normalization provides a unified framework for handling all frequencies.
3Productivity
If audio signals are not normalized, then processing is faster, but fingerprint accuracy deteriorates due to inconsistent energy distribution across frequency components
Solution Approach 1:
The patent performs spectral normalization as a preliminary step before fingerprint extraction. By pre-normalizing the audio signal's spectral components, the system establishes consistent energy distribution across all frequency ranges, ensuring that subsequent fingerprint processing operates on standardized data and achieves higher accuracy without significant speed penalty.
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
Methods, apparatus, systems, and articles of manufacture are disclosed to fingerprint audio via mean normalization. An example apparatus for audio fingerprinting includes a frequency range separator to transform an audio signal into a frequency domain, the transformed audio signal including a plurality of time-frequency bins including a first time-frequency bin, an audio characteristic determiner to determine a first characteristic of a first group of time-frequency bins of the plurality of time-frequency bins, the first group of time-frequency bins surrounding the first time-frequency bin and a signal normalizer to normalize the audio signal to thereby generate normalized energy values, the normalizing of the audio signal including normalizing the first time-frequency bin by the first characteristic. The example apparatus further includes a point selector to select one of the normalized energy values and a fingerprint generator to generate a fingerprint of the audio signal using the selected one of the normalized energy values.