Audio Fingerprint Extraction Using Spectrogram Masking and Weight Bits

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio fingerprint extraction methods suffer from poor accuracy due to their lack of robustness against noise and complexity, leading to ineffective audio fingerprint retrieval in multimedia applications.

Innovation Solution

The method involves converting an audio signal into a spectrogram through fast Fourier transformation, determining characteristic points, creating masks around these points, calculating mean energy of spectrum regions, and combining audio fingerprint bits with weight bits to enhance extraction accuracy, using techniques like MEL transformation and human auditory system filtering to improve robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional audio fingerprint extraction methods are used, then the extraction process is simple, but the accuracy and robustness against noise are poor

Engineering Contradiction:
Improveaudio fingerprint accuracyVSAvoidextraction process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the spectrogram into multiple mask regions around characteristic points, extracting multiple audio fingerprint bits from different regions. This segmentation allows the system to capture more comprehensive acoustic information and improve accuracy by considering multiple local characteristics rather than a single global feature.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces the dimension of mask region selection around characteristic points in the spectrogram. By extracting fingerprint bits from multiple spatial regions (masks) surrounding characteristic points rather than using single-point features, the system adds a spatial dimension to the extraction process, thereby improving robustness and accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If conventional audio fingerprint extraction methods are used, then the processing is fast, but the robustness with respect to noises is poor

Engineering Contradiction:
Improvenoise robustnessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-defining mask regions around characteristic points and pre-processing the spectrogram. This allows the extraction process to efficiently query pre-organized regions rather than performing complex analysis in real-time, thereby improving noise robustness while maintaining acceptable processing speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality analysis by examining the spectrogram locally around characteristic points using masks. Instead of analyzing the entire spectrogram uniformly, the system focuses on local regions with specific acoustic characteristics, which improves noise robustness by capturing local patterns that are more resilient to noise while reducing overall processing complexity.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If more characteristic points and masks are used in extraction, then the accuracy improves, but the computational complexity increases

Engineering Contradiction:
Improvefingerprint extraction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes parameters such as the number of characteristic points, mask radii, and spectral bin configurations to optimize the balance between accuracy and computational complexity. By adjusting these parameters, the system can adapt to different audio signals and application requirements, achieving high accuracy without excessive computational burden.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly enhances the accuracy and robustness of audio fingerprint extraction, leading to improved performance in audio comparison, search, deduplication, and surveillance applications by effectively addressing noise-related errors.

Implementation Method 1

converting the audio signal to a two-dimensional time-frequency spectrogram by fast Fourier transformation

Methodology Applied
Scientific EffectFast Fourier Transformation:

Implementation Method 2

processing the spectrogram by MEL transformation

Methodology Applied
Scientific EffectMEL transformation:

Implementation Method 3

processing the spectrogram by human auditory system filtering

Methodology Applied
Scientific EffectHuman auditory system filtering:

Data Source

PatentUS10950255B2Audio fingerprint extraction method and device
Publication Date: 2021.03.16 DOUYIN VISION CO LTD
  • US10950255B2 patent drawing
  • US10950255B2 patent drawing
  • US10950255B2 patent drawing

AI summary

An audio fingerprint extraction method and device are provided. The method includes: converting an audio signal to a spectrogram; determining one or more characteristic points in the spectrogram; in the spectrogram, determining one or more masks for the characteristic points; determining mean energy of each of the spectrum regions; determining one or more audio fingerprint bits according to mean energy of the plurality of spectrum regions in the one or more masks; judging credibility of the audio fingerprint bits to determine one or more weight bits; and combining the audio fingerprint bits and the weight bits to obtain an audio fingerprint. Each of the one or more masks includes a plurality of spectrum regions.