Audio Fingerprint Normalization for Time-Scale Shift Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio identification processes fail due to skewed speed and frequency-domain representations of audio signals, often caused by intentional alterations, processing errors, or Doppler shifts, preventing accurate matching with reference frequency-domain representations.
Innovation Solution
A method to normalize the query frequency-domain representation by determining the peak frequency within a predefined frequency bin and transforming it to a predefined position, aligning it with normalized reference frequency-domain representations, thereby facilitating accurate audio identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio signals are processed through content-presentation devices, then content can be transmitted and presented to users, but the audio signals may undergo time-scale modifications that skew frequency-domain representations and prevent accurate identification
Solution Approach 1:
The patent applies preliminary action by performing normalization of the query frequency-domain representation before matching. The system pre-processes the audio signal to correct time-scale modifications, establishing a normalized representation that aligns with reference signals. This includes determining frequency peaks, calculating normalization factors, and adjusting the query representation in advance of the matching operation, thereby eliminating identification failures caused by playback speed variations.
2Productivity
If frequency-domain representations are used for audio identification, then matching can be performed, but time-scale modifications cause skewed representations that fail to match reference signals
Solution Approach 1:
The patent applies parameter changes by modifying the frequency-domain representation through normalization. The system changes the frequency parameters of the query signal by applying a normalization factor derived from frequency peak analysis. This transformation adjusts the skewed frequency-domain representation to match the reference representation, thereby maintaining both identification efficiency and matching accuracy despite time-scale modifications in the audio signal.
Data Source
AI summary
A method includes receiving, by a computing system, an audio signal, where the audio signal defines a segment of media content over time. The method also includes establishing by the computing system, based on the received audio signal, a normalized query frequency-domain representation of the received audio signal. The method further includes matching, by the computing system, the normalized query frequency-domain representation of the received audio signal with a correspondingly normalized reference frequency-domain representation of a reference audio signal having an associated identity. The method additionally includes based on the matching, determining by the computing system that an identity of the received audio signal is the associated identity of the reference audio signal.


