Spatial Audio Rendering for Speech Noise Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech enhancement algorithms fail to effectively separate speech and noise in non-stationary environments, leading to imperfect noise suppression and reduced user experience due to the assumption of stationary noise, which is not always valid in real-world scenarios.

Innovation Solution

An apparatus and method that separate sound signals into speech and noise components and use spatial rendering to distribute these components differently in three-dimensional space, allowing the human auditory system to exploit spatial localization cues for improved separation, rather than relying on conventional noise suppression techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If conventional noise suppression techniques are used, then noise level is reduced, but speech intelligibility and quality deteriorate due to noise suppression artifacts

Engineering Contradiction:
Improvenoise levelVSAvoidspeech intelligibility
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent transforms the noise suppression problem from a one-dimensional amplitude reduction task into a three-dimensional spatial distribution task. By distributing speech and noise components to different spatial locations, the system exploits the additional spatial dimension to achieve separation without degrading speech quality, thus resolving the contradiction between noise reduction and speech intelligibility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent segments the mixed audio signal into distinct speech and noise components through spatial separation. By assigning different spatial distributions to speech (directional) and noise (diffuse), the system enables the auditory system to naturally separate the components, improving speech intelligibility while maintaining noise suppression effectiveness.

Inventive Principle:
Principle #1Segmentation

2Reliability

If spatial rendering is used to distribute speech and noise components, then speech intelligibility is improved, but device complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidspatial rendering complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent leverages the human auditory system's inherent ability to perform spatial localization and source separation as a free resource. By distributing speech and noise to different spatial locations, the system allows the listener's brain to automatically separate the components without requiring additional complex processing, thus improving speech intelligibility while avoiding excessive device complexity.

Inventive Principle:
Principle #25Self-service

3Device complexity

If stationary noise assumption is used, then processing is simplified, but performance deteriorates in non-stationary environments

Engineering Contradiction:
Improveprocessing complexityVSAvoidnoise suppression performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent bypasses the need for complex temporal analysis of noise stationarity by introducing spatial distribution as an additional dimension. Instead of trying to accurately model and track non-stationary noise over time, the system distributes noise components to diffuse spatial locations, allowing effective noise suppression even in non-stationary environments without significantly increasing processing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3005362B1Apparatus and method for improving a perception of a sound signal
Publication Date: 2021.09.22 HUAWEI TECH CO LTD
  • EP3005362B1 patent drawingFigure 1~3
  • EP3005362B1 patent drawingFigure 4~5

AI summary

The present invention relates to an apparatus (100) for improving a perception of a sound signal (S), the apparatus comprising: a separation unit (10) configured to separate the sound signal (S) into at least one speech component (SC) and at least one noise component (NC); and a spatial rendering unit (20) configured to generate an auditory impression of the at least one speech component (SC) at a first virtual position (VP1) with respect to a user, when output via a transducer unit (30),and of the at least one noise component (NC) at a second virtual position (VP2) with respect to the user, when output via the transducer unit (30).