Audio Fingerprint Extraction Using Spectrogram Masking and Weight Bits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio fingerprint extraction methods suffer from poor accuracy due to their lack of robustness against noise and complexity, leading to ineffective audio fingerprint retrieval in multimedia applications.
Innovation Solution
The method involves converting an audio signal into a spectrogram through fast Fourier transformation, determining characteristic points, creating masks around these points, calculating mean energy of spectrum regions, and combining audio fingerprint bits with weight bits to enhance extraction accuracy, using techniques like MEL transformation and human auditory system filtering to improve robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio fingerprint extraction methods are used, then the extraction process is simple, but the accuracy and robustness against noise are poor
Solution Approach 1:
The patent segments the spectrogram into multiple mask regions around characteristic points, extracting multiple audio fingerprint bits from different regions. This segmentation allows the system to capture more comprehensive acoustic information and improve accuracy by considering multiple local characteristics rather than a single global feature.
Solution Approach 2:
The patent introduces the dimension of mask region selection around characteristic points in the spectrogram. By extracting fingerprint bits from multiple spatial regions (masks) surrounding characteristic points rather than using single-point features, the system adds a spatial dimension to the extraction process, thereby improving robustness and accuracy.
2Reliability
If conventional audio fingerprint extraction methods are used, then the processing is fast, but the robustness with respect to noises is poor
Solution Approach 1:
The patent performs preliminary actions by pre-defining mask regions around characteristic points and pre-processing the spectrogram. This allows the extraction process to efficiently query pre-organized regions rather than performing complex analysis in real-time, thereby improving noise robustness while maintaining acceptable processing speed.
Solution Approach 2:
The patent applies local quality analysis by examining the spectrogram locally around characteristic points using masks. Instead of analyzing the entire spectrogram uniformly, the system focuses on local regions with specific acoustic characteristics, which improves noise robustness by capturing local patterns that are more resilient to noise while reducing overall processing complexity.
3Measurement precision
If more characteristic points and masks are used in extraction, then the accuracy improves, but the computational complexity increases
Solution Approach 1:
The patent changes parameters such as the number of characteristic points, mask radii, and spectral bin configurations to optimize the balance between accuracy and computational complexity. By adjusting these parameters, the system can adapt to different audio signals and application requirements, achieving high accuracy without excessive computational burden.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly enhances the accuracy and robustness of audio fingerprint extraction, leading to improved performance in audio comparison, search, deduplication, and surveillance applications by effectively addressing noise-related errors.
Implementation Method 1
converting the audio signal to a two-dimensional time-frequency spectrogram by fast Fourier transformation
Implementation Method 2
processing the spectrogram by MEL transformation
Implementation Method 3
processing the spectrogram by human auditory system filtering
Data Source
AI summary
An audio fingerprint extraction method and device are provided. The method includes: converting an audio signal to a spectrogram; determining one or more characteristic points in the spectrogram; in the spectrogram, determining one or more masks for the characteristic points; determining mean energy of each of the spectrum regions; determining one or more audio fingerprint bits according to mean energy of the plurality of spectrum regions in the one or more masks; judging credibility of the audio fingerprint bits to determine one or more weight bits; and combining the audio fingerprint bits and the weight bits to obtain an audio fingerprint. Each of the one or more masks includes a plurality of spectrum regions.


