Audio Content Recognition Using Variable Frame Shifts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio fingerprinting methods face high computational complexity and long processing times, leading to increased network resource loads and reduced matching accuracy.
Innovation Solution
An apparatus and method that vary the frame shift size between frames in a section with significant information, allowing for overlapping frames and adaptive signal-to-noise ratio-based section determination, coupled with a two-stage matching process for efficient and accurate content recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional fingerprint generation and matching methods are used, then matching accuracy can be maintained, but computational complexity is high and processing time is long
Solution Approach 1:
The patent segments the audio signal processing into distinct stages: audio signal input, feature extraction, fingerprint generation, and matching. By dividing the complex processing task into manageable segments, the system reduces overall computational complexity while maintaining matching accuracy through specialized processing at each stage.
Solution Approach 2:
The patent applies parameter changes by adjusting fingerprint generation parameters such as frame size, hop size, and frequency resolution based on the characteristics of the audio signal. This allows the system to optimize computational complexity for different signal types while preserving matching accuracy through adaptive parameter selection.
2Measurement precision
If conventional fingerprint generation and matching methods are used, then matching accuracy can be maintained, but processing time is excessively long
Solution Approach 1:
The patent implements preliminary action by pre-processing the audio signal to extract features and generate fingerprints before the actual matching operation. This preliminary fingerprint generation stores essential characteristics in advance, allowing rapid matching operations without re-processing the entire audio signal, thus reducing processing time while maintaining accuracy.
Solution Approach 2:
The patent extracts only the essential features from the audio signal to create compact fingerprints, taking out only the most discriminative characteristics needed for matching. This extraction process eliminates redundant information, reducing processing time for both fingerprint generation and matching while preserving matching accuracy through selective feature extraction.
3Productivity
If the number of fingerprints is decreased to reduce computational load, then processing time is reduced, but matching accuracy deteriorates
Solution Approach 1:
The patent applies local quality by generating variable numbers of fingerprints at different stages of the matching process. Initial broad matching uses fewer fingerprints for fast screening, while subsequent refined matching uses more fingerprints for accurate verification. This localized adjustment of fingerprint quantity optimizes both processing speed and matching accuracy at different stages.
4Device complexity
If the matching process is simplified to reduce computational complexity, then processing time is reduced, but matching accuracy deteriorates
Solution Approach 1:
The patent implements dynamics in the matching process by adaptively adjusting the matching algorithm's complexity based on the characteristics of the audio signals being compared. For highly similar signals, more complex matching is applied to ensure accuracy, while for dissimilar signals, simpler matching suffices, reducing overall processing time while maintaining accuracy when needed.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to an apparatus and method for recognizing content using an audio signal. The content recognition apparatus includes a query fingerprint extraction unit for forming frames having a preset frame length for an audio signal, and generating frame-based feature vectors for respective frames, thus extracting a query fingerprint. A reference fingerprint DB stores reference fingerprints to be compared with the query fingerprint and pieces of content information corresponding to the reference fingerprints. A fingerprint matching unit determines a reference fingerprint matching the query fingerprint. In this case, the query fingerprint extraction unit forms the frames while varying a frame shift size that is an interval between start points of neighboring frames in a partial section. According to the present invention, there can be provided a content recognition apparatus and method which can maintain the accuracy and reliability of matching while promptly providing results.