Audio Signal Processing for Speech Intelligibility in Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In noisy environments, speech intelligibility is severely impaired due to background noise, making it difficult for individuals to understand conversations or announcements, especially when the noise level is unpredictable and varies significantly.

Innovation Solution

A method and system that adapt audio signal processing to enhance speech intelligibility by approximating noise-free spectral features, adjusting frequency band gains, and imposing constraints to minimize distortion while maintaining power levels, using processors and computer-readable media to execute these operations in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio signal processing is applied to enhance speech intelligibility in noisy environments, then speech recognition accuracy is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is divided into multiple frequency bands, and each band is processed independently through spectral flattening and gain adjustment. This segmentation allows the system to handle complex computations in manageable portions, improving speech recognition accuracy while controlling overall computational complexity through parallel processing of frequency components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts parameters such as spectral flattening degree, frequency band gains, and noise floor thresholds based on the acoustic environment. By optimizing these parameters in real-time, the system achieves high speech recognition accuracy without requiring excessive computational resources, as the parameter adjustments are made based on simple noise floor measurements rather than complex signal analysis.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If spectral features are adapted to approximate noise-free environment, then speech intelligibility is enhanced, but processing delay increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidprocessing delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary spectral flattening and gain adjustments on frequency bands based on estimated noise floors before final speech recognition. This preliminary action prepares the audio signal in advance, reducing the need for complex real-time processing and minimizing overall processing delay while maintaining high speech intelligibility through pre-computed spectral corrections.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a simplified representation of the audio signal by copying and processing only the essential spectral features rather than the entire signal. By focusing on key frequency bands and their spectral characteristics, the system achieves accurate speech intelligibility enhancement with significantly reduced processing time, as only the most important signal components are analyzed and corrected.

Inventive Principle:
Principle #26Copying

3Measurement precision

If frequency band gains are adjusted to maximize intelligibility, then speech recognition improves, but power constraint violations may occur

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidpower constraint compliance
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system continuously monitors the output power levels of frequency band adjustments and feeds this information back to the processing algorithm. Based on this feedback, the system dynamically modifies gain parameters to maximize speech recognition accuracy while ensuring that power constraints are not violated. This closed-loop control allows the system to achieve optimal intelligibility without exceeding acceptable power levels.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system adjusts gain parameters for individual frequency bands while maintaining an overall power balance. By changing gain values in a controlled manner and compensating for power increases in certain bands with corresponding adjustments in other bands, the system achieves improved speech recognition accuracy while maintaining compliance with power constraints through coordinated parameter optimization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9866955B2Enhancement of intelligibility in noisy environment
Publication Date: 2018.01.09 GOOGLE LLC
  • US9866955B2 patent drawing
  • US9866955B2 patent drawing
  • US9866955B2 patent drawing

AI summary

Provided are methods and systems for enhancing the intelligibility of an audio (e.g., speech) signal rendered in a noisy environment, subject to a constraint on the power of the rendered signal. A quantitative measure of intelligibility is the mean probability of decoding of the message correctly. The methods and systems simplify the procedure by approximating the maximization of the decoding probability with the maximization of the similarity of the spectral dynamics of the noisy speech to the spectral dynamics of the corresponding noise-free speech. The intelligibility enhancement procedures provided are based on this principle, and all have low computational cost and require little delay, thus facilitating real-time implementation.