Audio Copy Detection via Two-Phase Fingerprint Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio copy detection methods are inefficient for monitoring peer-to-peer music sharing and advertisement campaigns, as they are either expensive or limited in their ability to track competitors' ads, and existing solutions prioritize speed over accuracy in large repositories with various distortions.

Innovation Solution

A method and apparatus for audio copy detection that generate and match fingerprint sets between query audio data and test audio data units, using energy-difference fingerprints and nearest-neighbor fingerprints, with a two-phase search process to efficiently identify matching segments, leveraging CPU and GPU processing for speed and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If exhaustive search is performed to ensure all copies are detected, then recall rate is improved, but computational cost and processing time increase significantly

Engineering Contradiction:
Improverecall rateVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the exhaustive search process into two distinct phases: a fast filtering phase that quickly identifies potential matches using simplified criteria, and a slower verification phase that performs detailed comparison only on candidates from the first phase. This segmentation allows the system to achieve high recall rates while maintaining practical processing speeds by avoiding unnecessary detailed comparisons.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions in the first phase by pre-processing query audio and test audio to extract features and generate candidate matches before the main verification step. This preliminary filtering reduces the search space significantly, allowing the system to maintain high recall while reducing overall computational cost.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If fast search algorithms are used to improve processing speed, then productivity is improved, but measurement precision and false alarm rates worsen

Engineering Contradiction:
Improveprocessing speedVSAvoidfalse alarm rate
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary verification step between the fast search phase and the final match declaration. Candidates identified by the fast search algorithm are subjected to additional verification checks that filter out false positives before confirming a match. This intermediary action maintains high processing speed while improving measurement precision by eliminating false alarms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies partial verification to all candidates and excessive verification (detailed comparison) only to promising candidates. This selective application of verification effort allows the system to maintain high speed while ensuring precision for the most likely matches, reducing false alarm rates without sacrificing overall processing efficiency.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If watermarking is used for ad monitoring, then ad detection capability is improved, but system complexity and cost increase

Engineering Contradiction:
Improvead detection capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses content-based copying and comparison of audio features instead of watermarking. By extracting and comparing intrinsic features of the audio content itself, the system achieves ad detection capability without requiring complex watermark embedding and detection infrastructure, thereby reducing system complexity while maintaining reliability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent enables the audio content itself to serve as the detection key through its inherent acoustic features, rather than requiring an external watermarking system. The audio's own characteristics are used for identification and matching, eliminating the need for separate watermarking hardware or software components.

Inventive Principle:
Principle #25Self-service

4Measurement precision

If bit matching and threshold computation are performed for complete search, then measurement precision is improved, but computational expense increases

Engineering Contradiction:
Improvematch accuracyVSAvoidcomputational expense
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the comparison process into coarse matching (using simplified features) and fine matching (using detailed bit-level comparison). By dividing the search into these stages, the system achieves high measurement precision for final matches while reducing overall computational expense by avoiding detailed comparison for all candidate pairs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial bit matching and threshold computation only to candidate matches identified in the first phase, rather than performing complete bit-level comparison for all possible audio pairs. This selective application of computationally intensive operations maintains high match accuracy while significantly reducing computational expense.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8831760B2Content based audio copy detection
Publication Date: 2014.09.09 CENT DE RECH INFORMATIQUE DE MONTREAL
  • US8831760B2 patent drawing
  • US8831760B2 patent drawing
  • US8831760B2 patent drawing

AI summary

A method for performing audio copy detection, comprising, providing a query audio data, the query audio data having a succession of frames and also providing a plurality of test audio data units, each test audio data unit including a succession of frames. For each test audio data unit the method generates a test fingerprint set. The generation of the test fingerprint test including computing similarity measurements between at least one frame of the test audio data and a plurality of frames of the query audio data. A test audio data unit is then selected as a match for the query audio data at least in part on the basis of the fingerprint sets.