Audio Content Recognition via DCT Hash Codes in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional multimedia content recognition technologies have a low recognition rate in high noise or asynchronous environments, making it difficult to provide additional services when multimedia content information is not available.
Innovation Solution
An audio content recognition method that generates hash codes based on the spectral shape of audio signals using discrete cosine transform (DCT) coefficient differences, with a frame interval delta_F, and applies weights based on frequency domain energy, to improve recognition accuracy even in noisy and asynchronous conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional content recognition technology is used, then the system can identify multimedia contents, but the recognition rate becomes remarkably low in high noise or asynchronous environments
Solution Approach 1:
The audio signal is divided into multiple frames, and audio fingerprints are extracted from each frame independently. This segmentation allows the system to process the audio signal in manageable portions, making it more robust to noise and asynchronous conditions by analyzing local characteristics rather than relying on the entire signal.
Solution Approach 2:
The patent transforms audio fingerprints into hash codes by changing the representation parameter. Instead of directly comparing audio fingerprints, the system applies a hash function to generate compact hash code representations, which are more resistant to noise and asynchronous variations while maintaining recognition accuracy.
2Measurement precision
If conventional content recognition technology is used, then the system can determine content ID and frame number, but the recognition rate becomes remarkably low in asynchronous environments
Solution Approach 1:
The system pre-generates hash codes for audio fingerprints from multiple frames before performing content recognition. This preliminary processing creates a robust representation that can tolerate asynchronous conditions and signal delays, allowing accurate content identification even when timing information is lost or distorted.
Solution Approach 2:
The patent creates hash code copies of audio fingerprints that serve as resilient representations. These hash code copies can be stored and compared without requiring precise timing information, enabling the system to overcome asynchronous environments and signal delays while maintaining measurement precision.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An audio content recognition method according to an embodiment of the present invention to solve a technical problem comprises the steps of: receiving an audio signal; obtaining an audio fmger-print (AFP) of the received audio signal; generating a hash code for the obtained audio fmger-print; transmitting a matching query for the generated hash code and a hash code stored in a database; and receiving a recognition result for a content of the audio signal as a response to transmission, wherein the step of generating the hash code comprises a step of determining a frame interval delta_F of the audio finger-print of which a hash code is to be generated among the obtained audio finger-prints.