Audio Signal Separation via Mask-Based Blind Source Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing technologies using multi-MIC beamforming are sensitive to microphone position errors and increase costs, while blind source separation methods for two MICs struggle to enhance voice signal quality effectively.

Innovation Solution

The method involves acquiring original noisy signals from multiple microphones, performing time-frequency estimation, determining mask values based on these signals, and updating the signals to separate audio sources accurately, reducing voice damage and improving quality without requiring precise microphone positioning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multi-MIC beamforming technology is used to improve voice signal processing quality, then voice recognition rate increases, but the system becomes sensitive to microphone position errors and product cost increases

Engineering Contradiction:
Improvevoice recognition rateVSAvoidmicrophone position accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical/physical beamforming approach that relies on precise microphone positioning with a signal processing-based blind source separation method. Instead of using spatial filtering that requires accurate knowledge of microphone positions and sound source directions, the invention uses statistical independence properties of sound sources to separate them in the time-frequency domain, thereby eliminating sensitivity to position errors

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the problem from spatial domain to time-frequency domain by applying short-time Fourier transform. This parameter transformation allows the system to work with spectral characteristics rather than spatial positions, changing the basis of separation from geometric relationships to statistical relationships in the frequency domain, which are invariant to microphone positioning errors

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the number of microphones is increased to improve separation performance, then audio signal quality improves, but product cost increases

Engineering Contradiction:
Improveaudio signal qualityVSAvoidnumber of microphones
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent replaces the hardware-based solution of using more microphones with a software-based blind source separation algorithm. The invention demonstrates that with only two microphones, by exploiting the statistical independence of sound sources and using iterative optimization in the time-frequency domain, high-quality separation can be achieved without increasing the number of physical sensors

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates virtual microphones through signal processing by combining signals from the two physical microphones in different ways across different frequency bins. This allows the system to effectively simulate the response of multiple physical microphones positioned at different locations, achieving the separation performance of a larger array using fewer physical elements

Inventive Principle:
Principle #26Copying

3Ease of manufacture

If blind source separation technology is used with two microphones to reduce hardware cost, then product cost decreases, but voice signal quality enhancement becomes difficult

Engineering Contradiction:
Improveproduct costVSAvoidvoice signal quality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments the audio signal into multiple frequency bins using short-time Fourier transform, and processes each frequency bin separately. This segmentation allows the application of different separation strategies to different frequency regions, improving overall voice quality by addressing the specific characteristics of each frequency band rather than treating the entire spectrum uniformly

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements an iterative optimization process where the separation results from one iteration are used to improve the separation in the next iteration. The algorithm continuously refines the separation by comparing the estimated sources with the observed mixture and adjusting the separation parameters accordingly, thereby enhancing voice signal quality through progressive improvement rather than a single-pass processing

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3839950B1Audio signal processing method, audio signal processing device and storage medium
Publication Date: 2024.10.09 BEIJING XIAOMI INTELLIGENT TECH CO LTD
  • EP3839950B1 patent drawingFigure 1
  • EP3839950B1 patent drawingFigure 2
  • EP3839950B1 patent drawingFigure 3

AI summary

A method for processing audio signal includes that: audio signals emitted respectively from at least two sound sources are acquired through at least two microphones to obtain respective original noisy signals of the at least two microphones; sound source separation is performed on the respective original noisy signals of the at least two microphones to obtain respective time-frequency estimated signals of the at least two sound sources; a mask value of the time-frequency estimated signal of each sound source in the original noisy signal of each microphone is determined based on the respective time-frequency estimated signals; the respective time-frequency estimated signals of the at least two sound sources are updated based on the respective original noisy signals of the at least two microphones and the mask values; and the audio signals emitted respectively from the at least two sound sources are determined.