Binaural Speech Signal Generation Using Deep Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise reduction methods in speech enhancement often distort speech and fail to utilize the human binaural hearing system, resulting in limited improvement in speech intelligibility, as they produce a single output and do not effectively separate speech and noise components in a listener's perceptual space.
Innovation Solution
A deep neural network (DNN) based on a temporal convolutional network (TCN) architecture is trained to generate binaural signals, rendering speech and noise components as if they come from different directions, utilizing a single-input/binaural-output (SIBO) method to improve speech intelligibility by transforming single-channel noisy observations into two waveform-domain signals for the left and right ears.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If optimal filtering techniques, spectral estimation procedures, statistical approaches, subspace methods, or deep learning based methods are used to reduce noise, then signal-to-noise ratio and speech quality are improved, but speech intelligibility is limited due to speech distortion and failure to utilize binaural hearing
Solution Approach 1:
The patent segments the single-channel noisy speech signal into separate speech and noise components by training a deep neural network to predict binaural room impulse responses for speech and noise independently, then combines them in the binaural domain to preserve speech intelligibility while reducing noise
Solution Approach 2:
The patent transforms the problem from single-channel monophonic processing to multi-channel binaural processing by predicting different binaural room impulse responses for speech and noise, utilizing the spatial dimension to separate components and improve intelligibility
2Device complexity
If single-channel noise reduction methods are used, then processing complexity is reduced, but speech and noise components cannot be effectively separated in the listener's perceptual space
Solution Approach 1:
The patent introduces binaural room impulse responses as an intermediary that enables effective separation of speech and noise components in the listener's perceptual space, allowing the system to achieve both low complexity and effective separation by operating in the binaural domain
Data Source
AI summary
A system and method of generating binaural signals includes receiving, by a processing device, a sound signal including speech and noise components, and transforming, by the processing device using a deep neural network (DNN), the sound signal into a first signal and a second signal. The transforming further includes encoding, by an encoding layer of the DNN, the sound signal into a sound signal representation in a latent space, rendering, by a rendering layer of the DNN, the sound signal representation into a first signal representation and a second signal representation in the latent space, and decoding, by a decoding layer of the DNN, the first signal representation into the first signal and the second signal representation into the second signal.


