Audio Fingerprinting With Weak-Feature Replacement for Noise Resistance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Historical audio fingerprinting technologies often generate fingerprints based on loudest parts of an audio signal, which can be associated with noise and do not utilize the entire audio spectrum effectively, leading to reduced usefulness and inaccuracies, especially in multi-source audio environments.
Innovation Solution
The method involves determining the relative strength of subfingerprints by evaluating how dependent each portion is on surrounding audio characteristics, labeling portions as weak or strong, and modifying or excluding weak portions to generate query fingerprints, while generating alternative fingerprints based on probability during reference fingerprinting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio fingerprints are generated based on loudest parts of the audio signal, then the fingerprinting process is simple and fast, but the accuracy and reliability of audio identification deteriorates due to noise interference and inability to utilize the entire audio spectrum
Solution Approach 1:
The audio signal is divided into multiple frequency bands (e.g., low, mid, high frequencies) and time segments. Instead of using the loudest parts only, the system analyzes multiple segments across different frequency ranges to create a comprehensive fingerprint that balances speed and accuracy.
Solution Approach 2:
Different portions of the audio signal are evaluated based on their local characteristics. The system identifies and weights specific frequency bands and time segments according to their relevance to the overall audio content, rather than uniformly treating all parts equally or focusing only on the loudest sections.
2Device complexity
If audio fingerprints focus only on loudest parts, then the processing complexity is reduced, but the reliability of audio verification deteriorates in multi-source audio environments
Solution Approach 1:
The audio signal is segmented into multiple frequency bands and time intervals. The system processes each segment separately to identify dominant sources, then integrates the results. This segmentation allows the system to handle complex multi-source environments without overwhelming processing complexity.
Solution Approach 2:
The system transitions from analyzing only the loudest part (one-dimensional approach) to analyzing multiple frequency bands and time segments simultaneously (multi-dimensional approach). This dimensional expansion enables reliable verification in multi-source environments while maintaining manageable processing complexity through structured analysis.
3Loss of information
If the entire audio spectrum is analyzed, then the completeness of audio information is improved, but the fingerprint generation time increases
Solution Approach 1:
The audio spectrum is divided into multiple frequency bands and time segments. The system analyzes each segment to identify dominant characteristics, then combines these into a comprehensive fingerprint. This segmentation enables complete spectral analysis while controlling generation time through efficient processing of individual segments.
Solution Approach 2:
The system performs preliminary analysis of audio segments to identify dominant frequency bands and time intervals before generating the final fingerprint. This preliminary action filters and prioritizes the most relevant information, reducing the time needed for complete spectral analysis while maintaining information completeness.
Data Source
AI summary
Methods, apparatus, systems, and articles of manufacture to fingerprint an audio signal. An example apparatus disclosed herein includes an audio segmenter to divide an audio signal into a plurality of audio segments, a bin normalizer to normalize the second audio segment to thereby create a first normalized audio segment, a subfingerprint generator to generate a first subfingerprint from the first normalized audio segment, the first subfingerprint including a first portion corresponding to a location of an energy extremum in the normalized second audio segment, a portion strength evaluator to determine a likelihood of the first portion to change, and a portion replacer to, in response to determining the likelihood does not satisfy a threshold, replace the first portion with a second portion to thereby generate a second subfingerprint.


