Neural Network Audio Fingerprint Extraction for Obfuscation Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing content-based audio recognition methods perform poorly in the presence of audio obfuscations such as changes in pitch, tempo, filtering, or the addition of background noise, and struggle to identify similarities between derivative works and their parent works.
Innovation Solution
A method using a trained neural network to generate fingerprints from frequency representations of audio content, which are robust to audio obfuscations, by converting selected areas of data points into vectors in a metric space and comparing them to reference fingerprints using a specified distance metric.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If pre-specified recipes are used to extract fingerprints from audio content, then the extraction process is simple and fast, but the fingerprints are highly sensitive to audio obfuscations such as pitch changes, tempo changes, filtering, and background noise
Solution Approach 1:
The patent transforms the audio fingerprinting approach by changing from fixed pre-specified extraction recipes to dynamic neural network-based feature extraction. The neural network learns optimal frequency representations and temporal patterns that are invariant to common audio obfuscations, thereby maintaining extraction efficiency while dramatically improving robustness to pitch shifts, tempo changes, filtering, and background noise.
Solution Approach 2:
The patent replaces traditional mechanical signal processing methods (filter banks, onset detection, inter-onset interval calculation) with a data-driven neural network system. This substitution allows the system to automatically learn robust acoustic features that are inherently more resistant to obfuscation, while maintaining computational efficiency through optimized neural network architectures.
2Use of energy by moving object
If traditional fingerprint extraction methods are used, then the method is computationally efficient, but it fails to identify similarities between derivative works and parent works when partial audio content is shared
Solution Approach 1:
The patent segments the audio signal into multiple frequency bins and extracts features from different temporal scales simultaneously. This multi-resolution analysis enables the system to detect partial similarities in derivative works by identifying matching patterns at various granularities, from individual onsets to broader temporal structures, thereby improving detection accuracy for partially shared content.
Solution Approach 2:
The patent extends the fingerprint representation by incorporating additional dimensional information, including multi-scale temporal features and frequency distribution patterns. This dimensional enrichment allows the system to capture subtle similarities in derivative works that traditional methods miss, while maintaining computational efficiency through compact feature representations.
Data Source
AI summary
Methods and a computer-readable storage device are disclosed for generating a frequency representation of a query audio file. The frequency representation represents information about at least a number of frequencies within a time range containing a number of time frames of the audio content information and a level associated with each of said frequencies. At least one of area of data points in the frequency representation is selected. A fingerprint for each selected area of data points is generated by applying a trained neural network onto said selected area of data points thereby generating a vector in a metric space. A distance between at least one of the generated query fingerprints and at least one reference fingerprint is calculated using a specified distance metric. A reference audio file having associated reference fingerprints which have produced at least one associated distance satisfying a predetermined threshold is identified.


