Frame-Level Audio Loudness Control Using Signal Analysis and Deep Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio loudness control systems fail to accurately analyze unusual audio characteristics, leading to inappropriate loudness control.
Innovation Solution
A method and system that combine signal analysis and deep learning to analyze audio characteristics at a frame level, determining the importance of frames and adjusting loudness levels based on the results from both analysis steps, using a layered structure of algorithms inspired by the human brain's neural network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If signal analysis is used to analyze audio characteristics, then the analysis process is simple and fast, but the reliability of analysis for unusual audio characteristics deteriorates
Solution Approach 1:
The patent combines signal analysis and deep learning analysis into a hybrid system. The signal analysis unit provides fast processing for typical audio, while the deep learning analysis unit handles unusual audio characteristics. The controller integrates both analysis results to determine final audio characteristics, achieving both speed and reliability.
Solution Approach 2:
The controller acts as an intermediary that receives analysis results from both the signal analysis unit and deep learning analysis unit. It combines these results to determine audio characteristics, allowing the system to leverage the strengths of both methods while mitigating their individual weaknesses.
2Device complexity
If only signal analysis is used, then the system complexity is low, but the ability to accurately grasp unusual audio characteristics deteriorates
Solution Approach 1:
The patent merges signal analysis and deep learning analysis into a unified system. The signal analysis unit handles simple, fast analysis while the deep learning analysis unit provides accurate analysis for complex cases. This combination improves measurement precision without excessive complexity increase.
Solution Approach 2:
The system applies different analysis methods to different audio characteristics. Signal analysis is used for typical audio patterns where simplicity is sufficient, while deep learning analysis is applied to unusual characteristics requiring higher precision. This localized approach optimizes the balance between complexity and precision.
3Reliability
If deep learning analysis is added to the system, then the reliability of audio characteristic analysis is improved, but the device complexity increases
Solution Approach 1:
The patent combines signal analysis and deep learning analysis in a hybrid architecture. The controller integrates results from both analysis units, allowing the system to achieve high reliability through deep learning while maintaining efficiency through signal analysis for routine cases.
Solution Approach 2:
The system applies deep learning analysis selectively rather than to all audio processing. The controller determines when deep learning analysis is needed based on the audio characteristics, applying it only when necessary to improve reliability, thus avoiding unnecessary complexity for all processing paths.
Data Source
AI summary
The present disclosure relates to a method and system for controlling loudness of an audio based on signal analysis and deep learning. The method includes analyzing an audio characteristic in a frame level based on signal analysis, analyzing the audio characteristic in the frame level based on learning, and controlling loudness of the audio in the frame level, by combining the analysis results. Accordingly, reliability of audio characteristic analysis can be enhanced and audio loudness can be optimally controlled.


