Context-Based Speech Enhancement With Multi-Encoder Transformers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech enhancement systems using deep neural networks fail to effectively suppress abrupt and stationary noise, such as clapping, leading to suboptimal performance in speech-related applications like ASR, speaker recognition, and emotion recognition.
Innovation Solution
A context-based speech enhancement system utilizing a multi-encoder transformer that processes both input spectral data and context data, including speaker information, text information, video information, and emotion, to generate improved speech enhancement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional deep neural network-based speech enhancement is used, then the system can process speech signals, but it fails to effectively suppress abrupt and stationary noise such as clapping
Solution Approach 1:
The system dynamically adapts to different noise conditions by using a transformer model that can process variable-length context data and adjust its attention mechanisms based on the specific characteristics of the input signal and context, enabling effective handling of both abrupt and stationary noise types
Solution Approach 2:
The system changes the parameter representation by transforming speech signals into spectral domain representations and using multiple encoders to extract different feature parameters (audio, visual, text) that can be combined to enhance speech under various noise conditions
2Reliability
If multi-modal context data is incorporated, then speech enhancement quality improves significantly, but the device complexity increases
Solution Approach 1:
The system segments the complex processing task into multiple independent encoder modules (audio encoder, visual encoder, text encoder) that each handle specific types of data, allowing the complex multi-modal processing to be broken down into manageable, specialized components
Solution Approach 2:
The transformer model acts as an intermediary that receives encoded representations from multiple specialized encoders and integrates them into a unified speech enhancement output, managing the complexity of combining multiple data modalities through a centralized processing mechanism
Data Source
AI summary
A device to perform speech enhancement includes one or more processors configured to process image data to detect at least one of an emotion, a speaker characteristic, or a noise type. The one or more processors are also configured to generate context data based at least in part on the at least one of the emotion, the speaker characteristic, or the noise type. The one or more processors are further configured to obtain input spectral data based on an input signal. The input signal represents sound that includes speech. The one or more processors are also configured to process, using a multi-encoder transformer, the input spectral data and the context data to generate output spectral data that represents a speech enhanced version of the input signal.


