Camera-Assisted Audio Reception for Low-SNR Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In noisy environments, the signal-to-noise ratio (SNR) of audio signals in mobile wireless communications can be low, making it difficult for listeners to understand speakers, especially when the speaker is at a distance or turned away.
Innovation Solution
Leveraging visual data from a camera to enhance audio reception by indexing audio artifacts using lip shapes and speech patterns, injecting the correct library indices into the audio stream to improve SNR.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If visual data is used to enhance audio reception, then audio quality is improved, but device complexity increases
Solution Approach 1:
The patent combines visual data from a camera with audio data from a microphone to enhance audio reception. The processing system integrates both modalities to infer sounds that cannot be reliably detected through audio alone, thereby improving audio quality while using existing device components.
Solution Approach 2:
The processing system performs multiple functions: it processes both audio and visual data, calculates signal-to-noise ratios, determines when visual enhancement is needed, and transfers enhanced audio to receiving devices. This multi-functionality improves audio quality across various scenarios without requiring dedicated hardware for each function.
2Measurement precision
If visual data processing is performed, then speech deciphering accuracy is improved, but processing time increases
Solution Approach 1:
The system preliminarily calculates the signal-to-noise ratio of the audio stream to determine whether visual data processing is necessary. This preliminary assessment avoids unnecessary visual processing when the audio signal is already sufficient, thereby reducing overall processing time while maintaining high accuracy when needed.
Solution Approach 2:
The processing system dynamically adjusts its operation based on real-time conditions. It alternates between audio-only processing and combined audio-visual processing based on the calculated signal-to-noise ratio, optimizing processing time while maintaining speech deciphering accuracy according to environmental conditions.
3Loss of energy
If audio stream transmission is reduced, then bandwidth consumption is reduced, but audio reception quality may deteriorate
Solution Approach 1:
Instead of transmitting raw audio streams, the system creates a compressed representation by indexing sounds to a library index based on visual and audio analysis. This copying approach reduces bandwidth consumption while maintaining audio reception quality through the indexed representation.
Solution Approach 2:
The system changes the parameter of audio transmission from raw audio streams to indexed representations. By transforming the audio data into a compressed format based on visual inference and audio analysis, it reduces bandwidth consumption while preserving the essential information needed for audio reception quality.
Data Source
AI summary
In one example, a method includes calculating a signal to noise ratio of a captured audio stream, determining that the signal to noise ratio of the captured audio stream is lower than a predefined threshold, acquiring visual data of a source of the captured audio stream in response to the determining that the signal to noise ratio of the captured audio stream is lower than the predefined threshold, using the visual data to infer a sound that is being made by the source of the captured audio stream, indexing the sound that is being made by the source of the captured audio stream to a library index, and transferring the library index to a receiving user endpoint device.


