Multimedia Signal Recognition Using Time-Frequency Hashing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multimedia signal recognition systems face accuracy issues due to distortion in fingerprints generated from audio signals, particularly caused by noise components, which affects the retrieval of multimedia files.
Innovation Solution
An electronic apparatus and method that segment a detection signal into frames, then into blocks, representing each block as a hash word based on both time and frequency features to minimize distortion and enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a fingerprint is generated based on only energy distribution according to frequency direction, then the generation process is simple, but distortion occurs due to noise frequency components leading to degraded retrieval accuracy
Solution Approach 1:
The audio signal is segmented into multiple frames, and each frame is further divided into multiple blocks. This segmentation allows the system to process different portions of the signal independently, reducing the impact of noise in any single segment on the overall fingerprint accuracy.
Solution Approach 2:
The patent transitions from using only frequency-direction energy distribution to incorporating both time and frequency features. By adding the time dimension through temporal feature extraction from multiple blocks, the system creates a more robust fingerprint that can distinguish between actual signal content and noise components.
2Device complexity
If only frequency feature is used for fingerprint generation, then processing is simpler, but errors in frequency detection cannot be compensated leading to fingerprint distortion
Solution Approach 1:
The patent merges frequency features and time features to create a composite fingerprint representation. By combining these two types of features from multiple blocks, the system achieves mutual compensation where errors in one feature type can be offset by the other, improving overall reliability.
Solution Approach 2:
The system changes from using a single parameter (frequency energy distribution) to using multiple parameters (time features and frequency features from multiple blocks). This parameter expansion allows for more accurate signal characterization and error compensation.
3Ease of operation
If a fingerprint is generated from the entire audio signal as a single unit, then the process is straightforward, but noise components cause significant distortion in the fingerprint
Solution Approach 1:
The audio signal is divided into multiple frames and blocks to isolate noise components. By processing smaller segments independently and combining their features, the system reduces the impact of noise on any single segment while maintaining overall fingerprint quality.
Solution Approach 2:
The patent uses the temporal distribution of features across multiple blocks to identify and compensate for noise-induced errors. Errors detected in frequency features can be corrected using time feature information, effectively converting the potential harm of noise into an opportunity for error correction.
Data Source
AI summary
Disclosed are an electronic apparatus for recognizing a multimedia signal and an operating method of the electronic apparatus, including segmenting a detection signal into a plurality of frames; segmenting each of the frames into a plurality of blocks; and representing each of the blocks as a hash word based on a time feature and a frequency feature for each of the blocks.


