Camera-Assisted Audio Reception for Low-SNR Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In noisy environments, the signal-to-noise ratio (SNR) of audio signals in mobile wireless communications can be low, making it difficult for listeners to understand speakers, especially when the speaker is at a distance or turned away.

Innovation Solution

Leveraging visual data from a camera to enhance audio reception by indexing audio artifacts using lip shapes and speech patterns, injecting the correct library indices into the audio stream to improve SNR.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If visual data is used to enhance audio reception, then audio quality is improved, but device complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines visual data from a camera with audio data from a microphone to enhance audio reception. The processing system integrates both modalities to infer sounds that cannot be reliably detected through audio alone, thereby improving audio quality while using existing device components.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing system performs multiple functions: it processes both audio and visual data, calculates signal-to-noise ratios, determines when visual enhancement is needed, and transfers enhanced audio to receiving devices. This multi-functionality improves audio quality across various scenarios without requiring dedicated hardware for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If visual data processing is performed, then speech deciphering accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeech deciphering accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system preliminarily calculates the signal-to-noise ratio of the audio stream to determine whether visual data processing is necessary. This preliminary assessment avoids unnecessary visual processing when the audio signal is already sufficient, thereby reducing overall processing time while maintaining high accuracy when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The processing system dynamically adjusts its operation based on real-time conditions. It alternates between audio-only processing and combined audio-visual processing based on the calculated signal-to-noise ratio, optimizing processing time while maintaining speech deciphering accuracy according to environmental conditions.

Inventive Principle:
Principle #15Dynamics

3Loss of energy

If audio stream transmission is reduced, then bandwidth consumption is reduced, but audio reception quality may deteriorate

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidaudio reception quality
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

Instead of transmitting raw audio streams, the system creates a compressed representation by indexing sounds to a library index based on visual and audio analysis. This copying approach reduces bandwidth consumption while maintaining audio reception quality through the indexed representation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameter of audio transmission from raw audio streams to indexed representations. By transforming the audio data into a compressed format based on visual inference and audio analysis, it reduces bandwidth consumption while preserving the essential information needed for audio reception quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250291543A1Leveraging visual data to enhance audio reception
Publication Date: 2025.09.18 AT&T MOBILITY II LLC
  • US20250291543A1 patent drawing
  • US20250291543A1 patent drawing
  • US20250291543A1 patent drawing

AI summary

In one example, a method includes calculating a signal to noise ratio of a captured audio stream, determining that the signal to noise ratio of the captured audio stream is lower than a predefined threshold, acquiring visual data of a source of the captured audio stream in response to the determining that the signal to noise ratio of the captured audio stream is lower than the predefined threshold, using the visual data to infer a sound that is being made by the source of the captured audio stream, indexing the sound that is being made by the source of the captured audio stream to a library index, and transferring the library index to a receiving user endpoint device.