Audio Fingerprint Peak Matching for Noisy Music Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems fail to accurately identify music or sounds in public or private spaces where there is little to no identifying information available, due to challenges such as background noise, hardware distortions, noise cancellation algorithms, codec distortions, and transmission errors, making it difficult to reliably match query audio with reference audio.

Innovation Solution

A method and system that use a characteristic matrix representation of audio signals, processed through filterbanks and perceptual encoding, to robustly identify music by capturing resilient features that survive additive noise and nonlinear distortions, allowing for alignment and matching of query and reference sounds despite environmental and transmission challenges.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional audio matching methods are used, then the system can operate with simple processing, but the identification accuracy deteriorates in noisy environments with background noise, hardware distortions, and transmission errors

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple frames, with each frame further segmented into frequency bins. This hierarchical segmentation allows the system to process and compare specific portions of the audio signal independently, improving robustness against noise and distortions while maintaining manageable computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the audio signal into a different representation domain (frequency domain via FFT), creating a 'color change' in the signal representation. This transformation reveals features that are more robust to noise and distortions, enabling accurate identification even in adverse conditions.

Inventive Principle:
Principle #32Color changes

2Productivity

If the system processes audio through multiple distortion stages (noise cancellation, codecs, transmission), then the audio can be transmitted and stored efficiently, but the signal fidelity deteriorates making matching difficult

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidsignal fidelity
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system extracts specific features (frequency bin magnitudes) from the audio signal that are most resilient to distortions introduced by noise cancellation, codecs, and transmission. By focusing on these extracted features rather than the complete signal, the system maintains identification accuracy despite efficiency-optimizing processing stages.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent acknowledges that distortions from noise cancellation and codecs are inevitable but transforms this challenge into an opportunity by designing a matching system that is specifically tuned to recognize patterns that survive these processing stages. The characteristic matrix representation is designed to capture features that persist through these distortion stages.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Measurement precision

If the system uses detailed frequency analysis to improve matching accuracy, then the identification precision improves, but the computational requirements and processing time increase

Engineering Contradiction:
Improvematching precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs frequency analysis on only the most relevant portions of the audio signal (selected frequency bins within specific ranges) rather than analyzing the entire spectrum in detail. This partial analysis approach achieves sufficient matching precision while significantly reducing computational requirements and processing time.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8452586B2Identifying music from peaks of a reference sound fingerprint
Publication Date: 2013.05.28 SOUNDHOUND AI IP LLC
  • US8452586B2 patent drawing
  • US8452586B2 patent drawing
  • US8452586B2 patent drawing

AI summary

Components of a method and system that allow identification of music from the song or sound using only the sound of the audio being played. A system built using the method and device components disclosed processes inputs sent from a mobile phone over a telephone or data connection, though inputs might be sent through any variety of computers, communications equipment, or consumer audio devices over any of their associated audio or data networks.