Context-Based Speech Enhancement With Multi-Encoder Transformers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech enhancement systems using deep neural networks fail to effectively suppress abrupt and stationary noise, such as clapping, leading to suboptimal performance in speech-related applications like ASR, speaker recognition, and emotion recognition.

Innovation Solution

A context-based speech enhancement system utilizing a multi-encoder transformer that processes both input spectral data and context data, including speaker information, text information, video information, and emotion, to generate improved speech enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional deep neural network-based speech enhancement is used, then the system can process speech signals, but it fails to effectively suppress abrupt and stationary noise such as clapping

Engineering Contradiction:
Improvenoise suppression effectivenessVSAvoidhandling of different noise types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts to different noise conditions by using a transformer model that can process variable-length context data and adjust its attention mechanisms based on the specific characteristics of the input signal and context, enabling effective handling of both abrupt and stationary noise types

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter representation by transforming speech signals into spectral domain representations and using multiple encoders to extract different feature parameters (audio, visual, text) that can be combined to enhance speech under various noise conditions

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multi-modal context data is incorporated, then speech enhancement quality improves significantly, but the device complexity increases

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the complex processing task into multiple independent encoder modules (audio encoder, visual encoder, text encoder) that each handle specific types of data, allowing the complex multi-modal processing to be broken down into manageable, specialized components

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The transformer model acts as an intermediary that receives encoded representations from multiple specialized encoders and integrates them into a unified speech enhancement output, managing the complexity of combining multiple data modalities through a centralized processing mechanism

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12380909B2Context-data based speech enhancement
Publication Date: 2025.08.05 QUALCOMM INC
  • US12380909B2 patent drawing
  • US12380909B2 patent drawing
  • US12380909B2 patent drawing

AI summary

A device to perform speech enhancement includes one or more processors configured to process image data to detect at least one of an emotion, a speaker characteristic, or a noise type. The one or more processors are also configured to generate context data based at least in part on the at least one of the emotion, the speaker characteristic, or the noise type. The one or more processors are further configured to obtain input spectral data based on an input signal. The input signal represents sound that includes speech. The one or more processors are also configured to process, using a multi-encoder transformer, the input spectral data and the context data to generate output spectral data that represents a speech enhanced version of the input signal.