Neural Formant Level Adjustment for Consistent Audio Mixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems lack effective methods for automatically adjusting formant levels based on target monitoring loudness, leading to inconsistent audio quality and mixing results.
Innovation Solution
A method and electronic device that utilize a neural network trained to determine optimal formant attenuation or amplification coefficients based on equal-loudness-level contours, ensuring the audio's overall spectral profile is preserved.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual formant adjustment methods are used in audio processing, then audio quality can be improved, but the complexity of operation increases and consistency deteriorates
Solution Approach 1:
The system automatically detects formants in the audio signal and adjusts their levels without requiring manual user intervention. The neural network model performs self-service by autonomously analyzing the spectral profile, identifying formant frequencies, and applying appropriate attenuation or amplification coefficients, thereby eliminating complex manual operations while maintaining high audio quality
Solution Approach 2:
The system dynamically changes the attenuation/amplification parameters of formants based on the detected spectral characteristics and target monitoring loudness. By automatically adjusting these parameters according to the audio content and playback conditions, the system achieves consistent audio quality without requiring manual parameter tuning
2Productivity
If automatic formant detection is implemented, then productivity is improved, but measurement precision of formant levels deteriorates
Solution Approach 1:
The patent introduces an equal-loudness-level contour as an intermediary between the automatic detection system and the final formant adjustment. This contour serves as a reference that compensates for playback system characteristics, allowing the neural network to accurately predict formant levels that will sound correct on various monitoring systems, thereby maintaining measurement precision while enabling automatic processing
Solution Approach 2:
The system replaces manual formant measurement and adjustment mechanisms with a neural network-based automatic detection system. The neural network is trained to accurately identify formant frequencies and levels by learning from labeled audio data, substituting manual spectral analysis with an automated intelligent system that maintains high precision while dramatically improving processing efficiency
3Device complexity
If formant levels are adjusted without considering monitoring loudness, then device complexity is reduced, but audio quality consistency worsens
Solution Approach 1:
The system dynamically adapts formant adjustment parameters based on the target monitoring loudness level. The neural network takes loudness information as input and adjusts attenuation/amplification coefficients accordingly, allowing the same audio content to sound consistent across different playback volumes and monitoring conditions. This dynamic adaptation maintains audio quality consistency without requiring multiple fixed systems for different loudness scenarios
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method comprising determining feature values of an input audio window and determining a formant attenuation/amplification coefficient for the input audio window based on the processing of the feature values by a neural network.