Audio Similarity Evaluator for Parametric Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding technologies face challenges in efficiently encoding audio signals using parametric techniques while maintaining perceptual quality, as they rely on heuristic methods rather than psychoacoustic models, especially when dealing with phase-insensitive human auditory processing and temporal envelope modulation.
Innovation Solution
An audio similarity evaluator is developed to obtain envelope signals and modulation information for frequency ranges, comparing them with reference signals to determine similarity, allowing for the adjustment of coding parameters based on perceptual relevance and computational complexity, using neural networks and psychoacoustic modeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If parametric techniques like SBR or IGF are used for bandwidth extension, then bitrate is reduced, but the error signal becomes large even when artifacts are hardly audible
Solution Approach 1:
The patent transforms the audio signal representation from time-domain waveform to frequency-domain spectral parameters. By encoding spectral envelopes, modulation information, and perceptual features instead of raw waveform data, the system achieves efficient compression while maintaining perceptual quality. The decoder reconstructs the audio signal from these parameters rather than adding noise to the waveform.
Solution Approach 2:
The patent replaces traditional waveform-based quantization with a psychoacoustic model-driven approach. Instead of quantizing time-domain samples and shaping quantization noise, the system uses spectral analysis, envelope extraction, and modulation detection to create a parametric representation that aligns with human auditory processing, substituting mechanical waveform manipulation with perceptual modeling.
2Measurement precision
If traditional perceptual masking models are used to evaluate audio quality, then they work well for waveform encoders, but they fail to accurately assess quality for parametric techniques like SBR or IGF
Solution Approach 1:
The patent introduces dynamic adaptation of the quality assessment model based on the encoding method used. The system switches between different evaluation approaches: traditional perceptual masking models for waveform encoders, and a new modulation-based assessment model for parametric techniques. This dynamic adaptation allows accurate quality measurement across diverse encoding methodologies.
Solution Approach 2:
The patent creates a universal quality assessment framework that can evaluate both waveform and parametric encoding methods. The new model extracts modulation information and spectral envelope characteristics that are relevant to human perception regardless of the encoding approach, making the assessment system versatile across different audio coding technologies.
3Reliability
If the human auditory system's phase insensitivity is considered, then temporal envelope becomes the main auditory information, but current coding methods do not adequately preserve envelope modulation
Solution Approach 1:
The patent extracts the temporal envelope modulation information from the audio signal by analyzing the amplitude variations of spectral components over time. This extracted envelope information is then encoded separately from the spectral data, ensuring that the perceptually critical modulation characteristics are preserved even when phase information is discarded or quantized coarsely.
Data Source
AI summary
An audio similarity evaluator obtains envelope signals for a plurality of frequency ranges on the basis of an input audio signal. The audio similarity evaluator is configured to obtain a modulation information associated with the envelope signals for a plurality of modulation frequency ranges, wherein the modulation information describes the modulation of the envelope signals. The audio similarity evaluator is configured to compare the obtained modulation information with a reference modulation information associated with a reference audio signal, in order to obtain an information about a similarity between the input audio signal and the reference audio signal. An audio encoder uses such an audio similarity evaluator. Another audio similarity evaluator uses a neural net trained using the audio similarity evaluator.


