Neural Network Audio Fingerprint Extraction for Obfuscation Robustness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content-based audio recognition methods perform poorly in the presence of audio obfuscations such as changes in pitch, tempo, filtering, or the addition of background noise, and struggle to identify similarities between derivative works and their parent works.

Innovation Solution

A method using a trained neural network to generate fingerprints from frequency representations of audio content, which are robust to audio obfuscations, by converting selected areas of data points into vectors in a metric space and comparing them to reference fingerprints using a specified distance metric.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If pre-specified recipes are used to extract fingerprints from audio content, then the extraction process is simple and fast, but the fingerprints are highly sensitive to audio obfuscations such as pitch changes, tempo changes, filtering, and background noise

Engineering Contradiction:
Improvefingerprint extraction speedVSAvoidfingerprint matching accuracy under obfuscation
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transforms the audio fingerprinting approach by changing from fixed pre-specified extraction recipes to dynamic neural network-based feature extraction. The neural network learns optimal frequency representations and temporal patterns that are invariant to common audio obfuscations, thereby maintaining extraction efficiency while dramatically improving robustness to pitch shifts, tempo changes, filtering, and background noise.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical signal processing methods (filter banks, onset detection, inter-onset interval calculation) with a data-driven neural network system. This substitution allows the system to automatically learn robust acoustic features that are inherently more resistant to obfuscation, while maintaining computational efficiency through optimized neural network architectures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Use of energy by moving object

If traditional fingerprint extraction methods are used, then the method is computationally efficient, but it fails to identify similarities between derivative works and parent works when partial audio content is shared

Engineering Contradiction:
Improvecomputational energy consumptionVSAvoidsimilarity detection accuracy for derivative works
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent segments the audio signal into multiple frequency bins and extracts features from different temporal scales simultaneously. This multi-resolution analysis enables the system to detect partial similarities in derivative works by identifying matching patterns at various granularities, from individual onsets to broader temporal structures, thereby improving detection accuracy for partially shared content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extends the fingerprint representation by incorporating additional dimensional information, including multi-scale temporal features and frequency distribution patterns. This dimensional enrichment allows the system to capture subtle similarities in derivative works that traditional methods miss, while maintaining computational efficiency through compact feature representations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10657175B2Audio fingerprint extraction and audio recognition using said fingerprints
Publication Date: 2020.05.19 SPOTIFY
  • US10657175B2 patent drawing
  • US10657175B2 patent drawing
  • US10657175B2 patent drawing

AI summary

Methods and a computer-readable storage device are disclosed for generating a frequency representation of a query audio file. The frequency representation represents information about at least a number of frequencies within a time range containing a number of time frames of the audio content information and a level associated with each of said frequencies. At least one of area of data points in the frequency representation is selected. A fingerprint for each selected area of data points is generated by applying a trained neural network onto said selected area of data points thereby generating a vector in a metric space. A distance between at least one of the generated query fingerprints and at least one reference fingerprint is calculated using a specified distance metric. A reference audio file having associated reference fingerprints which have produced at least one associated distance satisfying a predetermined threshold is identified.