Audio Copy Detection via Two-Phase Fingerprint Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio copy detection methods are inefficient for monitoring peer-to-peer music sharing and advertisement campaigns, as they are either expensive or limited in their ability to track competitors' ads, and existing solutions prioritize speed over accuracy in large repositories with various distortions.
Innovation Solution
A method and apparatus for audio copy detection that generate and match fingerprint sets between query audio data and test audio data units, using energy-difference fingerprints and nearest-neighbor fingerprints, with a two-phase search process to efficiently identify matching segments, leveraging CPU and GPU processing for speed and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If exhaustive search is performed to ensure all copies are detected, then recall rate is improved, but computational cost and processing time increase significantly
Solution Approach 1:
The patent segments the exhaustive search process into two distinct phases: a fast filtering phase that quickly identifies potential matches using simplified criteria, and a slower verification phase that performs detailed comparison only on candidates from the first phase. This segmentation allows the system to achieve high recall rates while maintaining practical processing speeds by avoiding unnecessary detailed comparisons.
Solution Approach 2:
The patent performs preliminary actions in the first phase by pre-processing query audio and test audio to extract features and generate candidate matches before the main verification step. This preliminary filtering reduces the search space significantly, allowing the system to maintain high recall while reducing overall computational cost.
2Productivity
If fast search algorithms are used to improve processing speed, then productivity is improved, but measurement precision and false alarm rates worsen
Solution Approach 1:
The patent introduces an intermediary verification step between the fast search phase and the final match declaration. Candidates identified by the fast search algorithm are subjected to additional verification checks that filter out false positives before confirming a match. This intermediary action maintains high processing speed while improving measurement precision by eliminating false alarms.
Solution Approach 2:
The patent applies partial verification to all candidates and excessive verification (detailed comparison) only to promising candidates. This selective application of verification effort allows the system to maintain high speed while ensuring precision for the most likely matches, reducing false alarm rates without sacrificing overall processing efficiency.
3Reliability
If watermarking is used for ad monitoring, then ad detection capability is improved, but system complexity and cost increase
Solution Approach 1:
The patent uses content-based copying and comparison of audio features instead of watermarking. By extracting and comparing intrinsic features of the audio content itself, the system achieves ad detection capability without requiring complex watermark embedding and detection infrastructure, thereby reducing system complexity while maintaining reliability.
Solution Approach 2:
The patent enables the audio content itself to serve as the detection key through its inherent acoustic features, rather than requiring an external watermarking system. The audio's own characteristics are used for identification and matching, eliminating the need for separate watermarking hardware or software components.
4Measurement precision
If bit matching and threshold computation are performed for complete search, then measurement precision is improved, but computational expense increases
Solution Approach 1:
The patent segments the comparison process into coarse matching (using simplified features) and fine matching (using detailed bit-level comparison). By dividing the search into these stages, the system achieves high measurement precision for final matches while reducing overall computational expense by avoiding detailed comparison for all candidate pairs.
Solution Approach 2:
The patent applies partial bit matching and threshold computation only to candidate matches identified in the first phase, rather than performing complete bit-level comparison for all possible audio pairs. This selective application of computationally intensive operations maintains high match accuracy while significantly reducing computational expense.
Data Source
AI summary
A method for performing audio copy detection, comprising, providing a query audio data, the query audio data having a succession of frames and also providing a plurality of test audio data units, each test audio data unit including a succession of frames. For each test audio data unit the method generates a test fingerprint set. The generation of the test fingerprint test including computing similarity measurements between at least one frame of the test audio data and a plurality of frames of the query audio data. A test audio data unit is then selected as a match for the query audio data at least in part on the basis of the fingerprint sets.


