Neural Speech Separation via Attentional Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing systems struggle to effectively separate speech signals from multiple speakers in crowded environments, particularly for individuals with auditory disorders, as existing methods require spatial separation and multiple microphones, which are not always feasible.
Innovation Solution
The implementation of neural-network-based speech separation processing using deep learning frameworks like deep attractor networks, Time-domain Audio Separation Networks (TasNet), and convolutional TasNet, which utilize neural signals to selectively amplify the attended speaker and attenuate others using a single microphone, allowing for real-time processing and low-power applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple microphones and spatial separation are used for speech separation, then speech separation performance is improved, but device complexity and cost increase
Solution Approach 1:
The patent replaces the mechanical approach of using multiple microphones and spatial processing with a neural network-based signal processing system. The deep neural network learns to separate speech sources directly from the mixed signal captured by a single microphone, substituting complex hardware configurations with intelligent software processing that achieves comparable or superior separation performance.
Solution Approach 2:
The patent transforms the speech separation problem from a spatial-domain problem (requiring multiple microphones) to a feature-space problem solved by neural networks. By changing the representation parameters from raw audio signals to learned feature embeddings, the system achieves effective separation using minimal hardware.
2Measurement precision
If traditional speech separation methods are used, then speech signals can be separated, but computational cost and processing latency increase
Solution Approach 1:
The patent performs preliminary feature extraction and representation learning from the raw mixed speech signal before applying the separation logic. The neural network pre-processes the input signal into meaningful feature embeddings that capture the essential characteristics of different speech sources, enabling faster and more accurate separation in subsequent processing stages.
Solution Approach 2:
The patent replaces traditional signal processing methods (such as spectral subtraction, Wiener filtering, or independent component analysis) with a deep neural network approach. This substitution enables end-to-end learning of separation patterns from data, achieving superior performance with optimized computational efficiency and reduced latency compared to conventional methods.
3Measurement precision
If neural network-based speech separation is implemented, then speech separation accuracy is improved, but energy consumption increases
Solution Approach 1:
The patent implements a neural network architecture that processes only the essential features needed for speech separation rather than analyzing the entire frequency-time spectrum in detail. By focusing computational resources on the most discriminative features for source separation, the system achieves high accuracy while minimizing unnecessary computations and energy consumption.
Data Source
AI summary
Disclosed are devices, systems, apparatus, methods, products, and other implementations, including a method comprising obtaining, by a device, a combined sound signal for signals combined from multiple sound sources in an area in which a person is located, and applying, by the device, speech-separation processing (e.g., deep attractor network (DAN) processing, online DAN processing, LSTM-TasNet processing, Conv-TasNet processing), to the combined sound signal from the multiple sound sources to derive a plurality of separated signals that each contains signals corresponding to different groups of the multiple sound sources. The method further includes obtaining, by the device, neural signals for the person, the neural signals being indicative of one or more of the multiple sound sources the person is attentive to, and selecting one of the plurality of separated signals based on the obtained neural signals. The selected signal may then be processed (amplified, attenuated).


