Adaptive Latency Speech Enhancement for Wireless Handsets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High ambient noise levels in wireless communication systems reduce the signal-to-noise ratio, leading to lower speech coding performance and inefficient bandwidth usage, while existing speech enhancement techniques are limited by high latency requirements, making it difficult to effectively cancel noise without compromising voice quality.

Innovation Solution

An adaptive latency system that dynamically adjusts processing time for speech enhancement modules based on ambient noise levels, allowing for increased latency in high noise conditions to improve signal analysis and noise reduction while maintaining low latency in low noise conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If speech enhancement techniques are applied to cancel background noise, then speech quality is improved, but latency increases beyond acceptable limits

Engineering Contradiction:
Improvebackground noiseVSAvoidlatency
Core Design Contradiction:
Object-affected harmful factorsVSLoss of time

Solution Approach 1:

The patent implements dynamic latency adjustment where the speech enhancement module adapts its processing time based on ambient noise levels. In high noise conditions, the module accepts increased latency to perform thorough signal analysis and noise cancellation. In low noise conditions, it reduces latency to meet real-time communication requirements. This dynamic adaptation resolves the contradiction by making latency flexible rather than fixed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the latency parameter dynamically based on noise conditions. A latency management module monitors ambient noise levels and adjusts the latency allocation for the speech enhancement module accordingly. This parameter change allows the system to optimize between noise cancellation effectiveness and real-time performance requirements.

Inventive Principle:
Principle #35Parameter changes

2Speed

If latency is reduced to meet real-time requirements, then voice communication responsiveness is improved, but signal analysis accuracy deteriorates in high noise conditions

Engineering Contradiction:
Improveprocessing speedVSAvoidsignal analysis accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent makes signal analysis accuracy dynamic by adjusting processing time based on noise conditions. In high noise environments, the system allocates more processing time to achieve accurate signal analysis and noise cancellation. In low noise environments, it uses minimal processing time while maintaining adequate accuracy. This dynamic approach resolves the contradiction between speed and accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the processing time parameter based on ambient noise levels and speech activity detection. The latency management module adjusts the observation time for signal analysis dynamically, allowing sufficient time for accurate analysis when needed while maintaining fast response when conditions permit.

Inventive Principle:
Principle #35Parameter changes

3Loss of substance

If bandwidth saving techniques like VAD and DTX are used, then bandwidth consumption is reduced, but performance deteriorates in high ambient noise levels

Engineering Contradiction:
Improvebandwidth consumptionVSAvoiddetection reliability
Core Design Contradiction:
Loss of substanceVSReliability

Solution Approach 1:

The patent introduces a speech enhancement module as an intermediary between the microphone and the bandwidth saving techniques. This module pre-processes the signal by canceling background noise before VAD/DTX detection. By improving the signal-to-noise ratio beforehand, the detection reliability of VAD/DTX is enhanced while still maintaining bandwidth saving benefits. The speech enhancement module acts as a mediator that enables reliable detection even in high noise conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If observation time is increased to improve signal detection, then detection accuracy is improved, but end-to-end latency increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidend-to-end latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements dynamic observation time adjustment where the speech enhancement module adapts its observation window based on ambient noise levels and speech activity. In high noise conditions, it increases observation time to improve detection accuracy. In low noise conditions or during active speech, it reduces observation time to minimize latency. This dynamic adaptation resolves the contradiction between detection accuracy and latency.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9437211B1Adaptive delay for enhanced speech processing
Publication Date: 2016.09.06 QOSOUND IP INNOVATIONS LLC
  • US9437211B1 patent drawing
  • US9437211B1 patent drawing
  • US9437211B1 patent drawing

AI summary

Provided is a system, method, and computer program product for improving the quality of voice communications on a mobile handset device by dynamically and adaptively selecting adjusting the latency of a voice call to accommodate an optimal speech enhancement technique in accordance with the current ambient noise level. The system, method and computer program product improves the quality of a voice call transmitted over a wireless link to a communication device dynamically increasing the latency of the voice call when the ambient noise level is above a predetermined threshold in order to use a more robust high-latency voice enhancement technique and by dynamically decreasing the latency of the voice call when the ambient noise level is below a predetermined threshold to use the low-latency voice enhancement techniques. The latency periods are adjusted by adding or deleting voice samples during periods of unvoiced activity.