Pitch-Resistant Audio Matching via Magnitude Ratio Descriptors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio matching systems face challenges in accurately matching audio files when the probe audio clip has been subjected to pitch shifting, time stretching, or other transformations, as these modifications alter the audio characteristics, leading to unreliable matching results.
Innovation Solution
The generation of audio file descriptors based on the time-frequency spectrogram's stable and invariant characteristics, using techniques such as magnitude ordering, anchor point comparisons, and magnitude ratios to create a composite identifier that is resistant to pitch shifting and other transformations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio matching uses standard audio characteristics for comparison, then matching is simple and fast, but matching accuracy deteriorates when audio clips undergo pitch shifting or time stretching transformations
Solution Approach 1:
The patent transforms the audio signal from time-domain to frequency-domain representation using spectrogram analysis, changing the parameter space from raw audio waveforms to frequency magnitude distributions. This parameter transformation enables the system to capture pitch-related information in a form that remains comparable even when pitch shifting occurs, as the relative frequency relationships are preserved in the spectrogram representation
Solution Approach 2:
The patent divides the audio signal into overlapping short-time frames and computes spectrograms for each frame, segmenting the continuous audio stream into discrete temporal segments. This segmentation allows the system to analyze local frequency patterns that are invariant to global pitch shifts, as each segment's relative frequency structure remains consistent despite transformations
2Ease of operation
If audio characteristics are extracted from transformed probe clips (pitch shifted, time stretched), then the probe can be processed flexibly, but the descriptor no longer matches the stored audio file characteristics
Solution Approach 1:
The patent normalizes the spectrogram magnitude values across different frames and segments, creating an equipotential reference framework where magnitude relationships are comparable across transformations. By establishing consistent normalization procedures, the system ensures that pitch-shifted or time-stretched probes operate from the same reference potential as stored files, enabling reliable matching despite operational flexibility
Solution Approach 2:
The patent designs the spectrogram-based descriptor to serve multiple functions: it captures temporal frequency patterns, pitch relationships, and timbral characteristics simultaneously. This universal descriptor can match audio files regardless of whether they have undergone pitch shifting, time stretching, or other transformations, making the matching system universally applicable to various probe processing scenarios
3Measurement precision
If the system uses detailed audio characteristics for matching, then matching precision is high for original clips, but the system becomes sensitive to transformations and produces different descriptors
Solution Approach 1:
The patent extracts only the magnitude information from the spectrogram, separating it from phase information and other transient characteristics that are highly sensitive to pitch shifting and time stretching. By taking out and utilizing only the magnitude component, the system achieves a balance between discrimination precision and transformation robustness, as magnitude relationships in the spectrogram are more stable under these transformations
Data Source
AI summary
Systems and methods for generating unique pitch-resistant descriptors for audio clips are provided. In one or more embodiments, a descriptor for an audio clip is generated as a function of relative magnitudes between interest points within the audio clip's time-frequency representation. A number of techniques for leveraging the relative magnitudes to generate descriptors are considered. These techniques include ordering of interest points as a function of ascending or descending magnitude, creation of binary vectors based on magnitude comparisons between pairs of points, and calculation of quantized magnitude ratios between pairs of points. Descriptors generated based on relative magnitudes according to the techniques disclosed herein are relatively invariant to common transformations to the original audio clip, such as pitch shifting, time stretching, global volume changes, equalization, and/or dynamic range compression.


