Audio Loudness Estimation for Consistent Speech and Music Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio playback devices, especially mobile devices, face challenges in maintaining consistent perceived loudness across different audio sources, such as speech and music, leading to user discomfort due to frequent volume adjustments.
Innovation Solution
A method and apparatus that determine a loudness estimate of an audio signal using loudness models, including digital and parametric filters, and adjust digital signal processing parameters to control the audio signal, ensuring consistent perceived loudness by differentiating between speech and music, and considering environmental audio filtering models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Illumination intensity
If volume level is increased to compensate for low perceived loudness of speech audio, then speech loudness is improved, but music loudness becomes excessively high causing user discomfort
Solution Approach 1:
The patent applies different digital signal processing parameters specifically tailored to speech audio signals. When speech is detected, the system adjusts parameters such as gain, dynamic range compression, and equalization to optimize perceived loudness without affecting music playback parameters, thereby resolving the contradiction between speech intelligibility and music quality
Solution Approach 2:
The system automatically detects the type of audio signal (speech or music) and self-adjusts the processing parameters without requiring user intervention. The speech detector and parameter selector work autonomously to maintain optimal playback conditions for different signal types, eliminating the need for manual volume adjustments
2Illumination intensity
If user manually adjusts volume for each audio source type, then perceived loudness consistency is improved, but ease of operation deteriorates
Solution Approach 1:
The system automatically detects the type of audio signal (speech or music) and self-adjusts the processing parameters without requiring user intervention. The speech detector and parameter selector work autonomously to maintain optimal playback conditions for different signal types, eliminating the need for manual volume adjustments
3Illumination intensity
If different digital signal processing parameters are applied for speech and music, then perceived loudness consistency is improved, but device complexity increases
Solution Approach 1:
The patent segments the audio processing into distinct paths: one for speech signals and one for music signals. The speech detector divides the input stream, and separate parameter sets are applied to each segment. This segmentation allows complex processing to be applied only where needed (speech) while keeping music processing simpler
Solution Approach 2:
The system dynamically selects processing parameters based on real-time detection of signal type. The parameter selector continuously adapts the processing chain by switching between speech-optimized and music-optimized parameter sets, allowing the system to maintain simplicity for music while providing enhanced processing only when speech is detected
Data Source
AI summary
An apparatus comprising at least one processor and at least one memory including computer program code. The at least one memory and the computer program code is configured to, with the at least one processor, cause the apparatus at least to determine a loudness estimate of a first audio signal, generate a parameter dependent on the loudness estimate; and control the first audio signal dependent on the parameter.


