Harmonic Source Enhancement Using Time-Frequency Audio Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies struggle to effectively separate and enhance harmonic sources from audio signals, which is crucial for improved audio identification, authentication, and media classification.
Innovation Solution
The development of an audio analyzer and determiner system that processes audio signals to enhance harmonic sources by decomposing them into harmonic and percussive components, using techniques such as magnitude spectrogram analysis and time-frequency masking.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio signals are processed to separate harmonic and percussive components, then audio identification accuracy is improved, but processing complexity increases
Solution Approach 1:
The audio signal is segmented into harmonic and percussive components using spectral decomposition. The magnitude spectrogram is divided into harmonic regions (repeating patterns) and percussive regions (transient events), allowing separate processing and enhancement of each component to improve identification accuracy.
Solution Approach 2:
A time-frequency mask is introduced as an intermediary tool to selectively enhance harmonic components while suppressing percussive components. The mask is generated by analyzing the spectrogram and applying thresholding operations to create a binary mask that guides the enhancement process.
2Measurement precision
If harmonic sources are enhanced by separating from percussive components, then audio fingerprinting precision is improved, but computational requirements increase
Solution Approach 1:
Instead of processing the entire audio signal uniformly, the method applies partial action by focusing computational resources only on regions identified as harmonic through spectral analysis. The time-frequency mask selectively processes only the harmonic portions of the signal, reducing overall computational requirements while maintaining fingerprinting precision.
Solution Approach 2:
The spectrogram is pre-processed and analyzed to identify harmonic regions before the main enhancement operation. This preliminary action allows the system to prepare a time-frequency mask in advance, which then guides the enhancement process and reduces the computational burden during actual fingerprinting.
3Measurement precision
If audio signals are decomposed into harmonic and percussive components, then media classification accuracy is improved, but processing time increases
Solution Approach 1:
The method leverages the periodic nature of harmonic sounds (musical tones, vowels) to identify and enhance them. By detecting the periodic repetitions in the spectrogram, the system can quickly distinguish harmonic from percussive components without requiring extensive processing, as the periodic pattern provides a clear temporal signature.
Solution Approach 2:
The patent replaces complex mechanical signal processing with spectral domain analysis. By transforming the audio signal into the frequency domain using Fast Fourier Transform (FFT), the system can efficiently separate harmonic and percussive components through spectral characteristics rather than time-domain convolution, reducing processing time.
Data Source
AI summary
Methods and apparatus for harmonic source enhancement are disclosed herein. An example apparatus includes an interface to receive a media signal. The example apparatus also includes a harmonic source enhancer to determine a magnitude spectrogram of audio corresponding to the media signal; generate a time-frequency mask based on the magnitude spectrogram; and apply the time-frequency mask to the magnitude spectrogram to enhance a harmonic source of the media signal.


