Voice Intelligibility Processor With Multiband Noise Compensation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice playback devices struggle with voice intelligibility in noisy environments due to practical challenges such as physical limitations of playback and noise capture devices, signal headroom, and long-term voice characteristics, leading to degraded voice quality and inaccurate intelligibility analysis.

Innovation Solution

A voice intelligibility processor (VIP) that employs digital-to-acoustic level conversion, multiband noise and voice correction, short segment analysis, and long-term profiling, along with global and per-band gain analysis, to enhance voice intelligibility by adjusting gain parameters based on device characteristics and environmental noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If voice playback is performed in noisy environments using conventional techniques, then voice playback functionality is provided, but voice intelligibility is degraded due to background noise masking

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidbackground noise masking
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The voice signal is divided into multiple frequency bands for separate processing. Each band is analyzed and enhanced independently based on its specific intelligibility requirements, allowing targeted noise suppression in critical frequency ranges while preserving natural sound in less critical bands.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different processing strategies are applied to different frequency bands based on their importance for voice intelligibility. Critical bands receive aggressive noise suppression and enhancement, while non-critical bands receive minimal processing to maintain natural sound quality.

Inventive Principle:
Principle #3Local quality

2Reliability

If noise capture devices and processing techniques are used to enhance voice intelligibility, then voice clarity in noise is improved, but practical implementation challenges arise including physical limitations of devices and signal headroom constraints

Engineering Contradiction:
Improvevoice intelligibility enhancementVSAvoidphysical limitations of playback and noise capture devices
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system dynamically adjusts processing parameters including gain values, frequency band boundaries, and enhancement strength based on real-time analysis of noise characteristics, voice signal properties, and device capabilities. This allows optimization of intelligibility enhancement while respecting physical device limitations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The processing system continuously adapts to changing acoustic environments and signal conditions by dynamically adjusting enhancement parameters. The system monitors signal headroom, noise levels, and device performance characteristics to modulate processing intensity in real-time.

Inventive Principle:
Principle #15Dynamics

3Reliability

If aggressive noise suppression and voice enhancement processing is applied, then voice intelligibility is improved, but signal headroom is reduced and natural sound transitions are compromised

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidsignal headroom
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system applies enhancement processing selectively to only those frequency bands and time periods where it is most needed for intelligibility, rather than uniformly across the entire signal. This partial action approach maintains signal headroom in regions where enhancement is not critical.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system continuously monitors the processed output signal levels and intelligibility improvements, using this feedback to adjust enhancement strength and prevent excessive processing that would consume signal headroom. The feedback loop ensures natural transitions by detecting when enhancement should be reduced or discontinued.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If conventional voice processing is used without considering long-term voice characteristics, then processing simplicity is maintained, but accurate intelligibility analysis and natural sound quality are compromised

Engineering Contradiction:
Improveintelligibility analysis accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of long-term voice characteristics, noise profiles, and device response functions before applying enhancement processing. This preliminary characterization enables more accurate real-time intelligibility analysis and informs adaptive processing parameter selection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically characterizes the acoustic environment, noise sources, and device properties through self-monitoring and adaptive analysis, eliminating the need for manual configuration. The processing system serves itself by learning optimal parameters from the operating conditions.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12354617B2Context-aware voice intelligibility enhancement
Publication Date: 2025.07.08 DTS INC(US)
  • US12354617B2 patent drawing
  • US12354617B2 patent drawing
  • US12354617B2 patent drawing

AI summary

A method comprises: detecting noise in an environment with a microphone to produce a noise signal; receiving a voice signal to be played into the environment through a loudspeaker; performing multiband correction of the noise signal based on a microphone transfer function of the microphone, to produce a corrected noise signal; performing multiband correction of the voice signal based on a loudspeaker transfer function of the loudspeaker to produce a corrected voice signal; and computing multiband voice intelligibility results based on the corrected noise signal and the corrected voice signal.