Audio Loudness Estimation for Consistent Speech and Music Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Audio playback devices, especially mobile devices, face challenges in maintaining consistent perceived loudness across different audio sources, such as speech and music, leading to user discomfort due to frequent volume adjustments.

Innovation Solution

A method and apparatus that determine a loudness estimate of an audio signal using loudness models, including digital and parametric filters, and adjust digital signal processing parameters to control the audio signal, ensuring consistent perceived loudness by differentiating between speech and music, and considering environmental audio filtering models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If volume level is increased to compensate for low perceived loudness of speech audio, then speech loudness is improved, but music loudness becomes excessively high causing user discomfort

Engineering Contradiction:
Improveperceived loudnessVSAvoiduser discomfort
Core Design Contradiction:
Illumination intensityVSObject-affected harmful factors

Solution Approach 1:

The patent applies different digital signal processing parameters specifically tailored to speech audio signals. When speech is detected, the system adjusts parameters such as gain, dynamic range compression, and equalization to optimize perceived loudness without affecting music playback parameters, thereby resolving the contradiction between speech intelligibility and music quality

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system automatically detects the type of audio signal (speech or music) and self-adjusts the processing parameters without requiring user intervention. The speech detector and parameter selector work autonomously to maintain optimal playback conditions for different signal types, eliminating the need for manual volume adjustments

Inventive Principle:
Principle #25Self-service

2Illumination intensity

If user manually adjusts volume for each audio source type, then perceived loudness consistency is improved, but ease of operation deteriorates

Engineering Contradiction:
Improveperceived loudness consistencyVSAvoiduser convenience
Core Design Contradiction:
Illumination intensityVSEase of operation

Solution Approach 1:

The system automatically detects the type of audio signal (speech or music) and self-adjusts the processing parameters without requiring user intervention. The speech detector and parameter selector work autonomously to maintain optimal playback conditions for different signal types, eliminating the need for manual volume adjustments

Inventive Principle:
Principle #25Self-service

3Illumination intensity

If different digital signal processing parameters are applied for speech and music, then perceived loudness consistency is improved, but device complexity increases

Engineering Contradiction:
Improveperceived loudness consistencyVSAvoidsignal processing complexity
Core Design Contradiction:
Illumination intensityVSDevice complexity

Solution Approach 1:

The patent segments the audio processing into distinct paths: one for speech signals and one for music signals. The speech detector divides the input stream, and separate parameter sets are applied to each segment. This segmentation allows complex processing to be applied only where needed (speech) while keeping music processing simpler

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects processing parameters based on real-time detection of signal type. The parameter selector continuously adapts the processing chain by switching between speech-optimized and music-optimized parameter sets, allowing the system to maintain simplicity for music while providing enhanced processing only when speech is detected

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9998081B2Method and apparatus for processing an audio signal based on an estimated loudness
Publication Date: 2018.06.12 NOKIA TECHNOLOGIES OY
  • US9998081B2 patent drawing
  • US9998081B2 patent drawing
  • US9998081B2 patent drawing

AI summary

An apparatus comprising at least one processor and at least one memory including computer program code. The at least one memory and the computer program code is configured to, with the at least one processor, cause the apparatus at least to determine a loudness estimate of a first audio signal, generate a parameter dependent on the loudness estimate; and control the first audio signal dependent on the parameter.