Adaptive Speech Intelligibility Compensation for Noisy Audio Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio and video devices face challenges in effectively compensating for noise and improving speech intelligibility, particularly in environments where conventional noise compensation algorithms are insufficient due to limited capabilities of audio reproduction transducers.

Innovation Solution

Implementing methods that determine noise metrics and speech intelligibility metrics to adjust audio and non-audio features, such as altering audio processing, controlling closed captioning systems, and applying non-audio-based compensation techniques, without using broadband gain increases, to enhance user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional noise compensation algorithms are used, then basic noise reduction is achieved, but speech intelligibility remains insufficient in noisy environments

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidnoise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system dynamically adjusts audio processing parameters including noise reduction strength, equalization curves, and gain settings based on measured noise metrics and speech intelligibility metrics. This allows optimization of speech clarity across varying noise conditions without relying solely on conventional algorithms.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system continuously measures noise metrics and speech intelligibility metrics from the audio environment and uses this feedback to automatically adjust audio processing parameters. This closed-loop approach enables real-time optimization of speech intelligibility in response to changing noise conditions.

Inventive Principle:
Principle #23Feedback

2Object-affected harmful factors

If broadband gain increase is applied to compensate for noise, then overall audio volume increases, but distortion and loss of audio quality occur

Engineering Contradiction:
Improvenoise compensationVSAvoidaudio quality
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

Instead of applying broadband gain increase, the system applies targeted noise reduction and speech enhancement to specific frequency bands where speech components are present. This preserves audio quality by avoiding unnecessary amplification across the entire frequency spectrum while still compensating for noise in critical speech ranges.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The audio spectrum is divided into multiple frequency bands, and different processing strategies are applied to each band based on the presence of speech and noise characteristics. This allows selective noise compensation in speech-relevant bands without affecting other frequency ranges, thereby maintaining overall audio quality.

Inventive Principle:
Principle #1Segmentation

3Reliability

If audio processing is intensified to improve speech intelligibility, then speech clarity improves, but complexity of the audio processing system increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically measures noise metrics and speech intelligibility metrics and adjusts processing parameters without requiring manual intervention or complex user configuration. This self-adjusting capability reduces the perceived complexity for users while maintaining sophisticated processing for optimal speech intelligibility.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-calculates and stores optimal processing parameters for various noise conditions and speech intelligibility levels. During operation, it quickly selects and applies appropriate pre-computed parameters based on current measurements, avoiding the need for complex real-time optimization calculations.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If closed captioning is enabled to compensate for poor speech intelligibility, then accessibility improves, but user experience is degraded due to reliance on text instead of audio

Engineering Contradiction:
ImproveaccessibilityVSAvoiduser experience
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The system dynamically adjusts audio processing parameters including noise reduction strength, equalization, and gain settings based on measured noise metrics and speech intelligibility metrics. This allows optimization of speech clarity across varying noise conditions, reducing the need for closed captioning and improving overall user experience while maintaining accessibility.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250342851A1Adjusting audio and non-audio features based on noise metrics and speech intelligibility metrics
Publication Date: 2025.11.06 DOLBY LABORATORIES LICENSING CORP
  • US20250342851A1 patent drawing
  • US20250342851A1 patent drawing
  • US20250342851A1 patent drawing

AI summary

Some implementations involve determining a noise metric and/or a speech intelligibility metric and determining a compensation process corresponding to the noise metric and/or the speech intelligibility metric. The compensation process may involve altering a processing of audio data and/or applying a non-audio-based compensation method. In some examples, altering the processing of the audio data does not involve applying a broadband gain increase to the audio signals. Some examples involve applying the compensation process in an audio environment. Other examples involve determining compensation metadata corresponding to the compensation process and transmitting an encoded content stream that includes encoded compensation metadata, encoded video data and encoded audio data from a first device to one or more other devices.