Compressed Domain Audio Fingerprint Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio fingerprinting techniques require fully decoding compressed audio files to PCM format, which is time-consuming and inefficient, especially for files in formats like MP3, and often necessitates transcoding and spectral calculations, increasing processing load and latency.
Innovation Solution
A method that extracts and compresses frequency-domain data directly from encoded audio files, such as MP3 files, using a modified decoding process that avoids full decoding and spectral image computation, generating a compressed frequency domain value file for identification, which can be transmitted to a server for audio file verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio fingerprinting techniques are used that fully decode compressed audio files to PCM format, then accurate audio identification can be achieved, but processing time and computational load increase significantly
Solution Approach 1:
The patent extracts only the essential frequency-domain data needed for fingerprinting from the compressed audio bitstream, rather than fully decoding to PCM. This selective extraction of relevant information (MDCT coefficients and spectral data) maintains identification accuracy while avoiding unnecessary computational steps.
Solution Approach 2:
The system performs partial decoding by extracting frequency-domain representations directly from the compressed domain without completing the full decoding process to PCM format. This partial action provides sufficient information for accurate fingerprinting while significantly reducing processing time and computational resources.
2Reliability
If full decoding and spectral calculations are performed on compressed audio files, then comprehensive audio analysis is achieved, but device power consumption and processing load increase
Solution Approach 1:
The patent extracts only the necessary frequency-domain components from the compressed audio stream that are sufficient for reliable fingerprinting, avoiding the energy-intensive full decoding process while maintaining analysis reliability.
Solution Approach 2:
The system changes the processing domain from time-domain PCM decoding to direct frequency-domain extraction from compressed bitstreams, altering the computational parameters to reduce power consumption while preserving analytical reliability.
3Ease of operation
If uncompressed PCM format is used for audio fingerprinting, then processing is straightforward, but file size and bandwidth requirements increase
Solution Approach 1:
The system uses partial decoding to extract only the frequency-domain data necessary for fingerprinting from compressed files, avoiding full decompression to PCM. This produces compact fingerprint data with reduced bandwidth requirements while maintaining processing simplicity.
Solution Approach 2:
The patent creates a compressed frequency-domain representation (a copy of the essential spectral information) that serves the fingerprinting function without requiring the full uncompressed PCM data, thereby reducing data size while preserving processing ease.
4Adaptability or versatility
If conventional fingerprinting methods are used that work with fully decoded audio, then broad compatibility is achieved, but processing speed decreases
Solution Approach 1:
The patent performs preliminary extraction of frequency-domain data directly from the compressed bitstream before full decoding would be required. This preliminary action prepares the necessary fingerprint data in advance, significantly increasing processing speed while maintaining compatibility with standard audio formats.
Solution Approach 2:
The system performs partial decoding to extract frequency-domain representations sufficient for fingerprinting across multiple audio formats, achieving broad compatibility through standardized frequency-domain processing while avoiding the speed penalty of full decoding.
Data Source
AI summary
A computer-implemented method performed by a data processing apparatus includes receiving an audio signal that includes a frequency-domain representation of an audio file, extracting, from the audio signal, a plurality of frequency-domain data values that correspond to at least a portion of the audio file, compressing the plurality of data values to form a compressed frequency domain value file, and transmitting the compressed frequency domain value file to a server to identify the audio file.


