Audio Signal Merging for Speech Recognition in Noisy Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of speech recognition in automated assistants degrades when audio signals have low signal quality or high noise, such as background noise and reverberation, leading to a trade-off between noise reduction and loss of useful speech signal.
Innovation Solution
A method is introduced to generate a merged audio signal by combining multiple audio signals from client devices using weight values determined by signal-to-noise ratios, processed through a trained neural network, to improve speech recognition accuracy without sacrificing useful speech information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio denoising is applied to improve speech recognition accuracy, then noise is reduced, but useful speech signal is lost
Solution Approach 1:
The patent combines multiple audio signals from different client devices (smartphone, smart speaker, earbuds) that are positioned at different locations and orientations relative to the user. By merging these signals with different noise profiles, the system achieves noise cancellation while preserving speech content, resolving the contradiction between denoising effectiveness and speech preservation.
Solution Approach 2:
The patent applies different processing weights to different audio signals based on their individual signal-to-noise ratios. Each audio signal is evaluated locally and assigned a weight that reflects its quality, allowing the system to optimize the contribution of each signal source to the final merged output, thereby preserving speech while reducing noise.
2Measurement precision
If multiple audio signals are combined to reduce noise, then speech recognition accuracy improves, but system complexity increases
Solution Approach 1:
The patent transforms audio signals into the frequency domain using short-time Fourier transform (STFT) to generate spectrograms. This parameter transformation simplifies the noise estimation and signal separation process by operating on frequency-domain representations rather than time-domain waveforms, reducing computational complexity while maintaining accuracy.
Solution Approach 2:
The patent introduces a trained neural network model as an intermediary that automatically estimates noise spectrograms from the audio signals. This neural network mediator handles the complex task of noise characterization and signal separation, reducing the overall system complexity by encapsulating complex processing in a trained model that can be efficiently deployed.
Data Source
AI summary
Merging first and second audio data to generate merged audio data, where the first audio data captures a spoken utterance of a user and is collected by a first computing device within an environment, and the second audio data captures the spoken utterance and is collected by a distinct second computing device that is within the environment. In some implementations, the merging includes merging the first audio data using a first weight value and merging the second audio data using a second weight value. The first and second weight values can be based on predicted signal-to-noise ratios (SNRs) for respective of the first audio data and the second audio data, such as a first SNR predicted by processing the first audio data using a neural network model and a second SNR predicted by processing the second audio data using the neural network model.


