Auditory Selection Using Dynamic Memory for Speech Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech separation technologies face challenges with unsatisfactory performance due to unknown numbers of speakers and fixed memory dimensions, leading to inaccurate results in real-world applications, especially in environments with interference.
Innovation Solution
An auditory selection method based on a memory and attention model that encodes speech signals into time-frequency matrices, uses a bi-directional long short-term memory (BiLSTM) network to transform these matrices into speech vectors, and employs a long-term memory unit to store and update speaker vectors for effective separation of target speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised learning method is used for speech separation, then speech recognition can be performed, but the arrangement of supervised labels is uncertain when the number of speakers is unknown, leading to inaccurate results
Solution Approach 1:
The patent transforms the fixed supervised learning approach into a dynamic unsupervised learning system using neural networks that can automatically adapt to varying numbers of speakers. The system dynamically adjusts to the actual number of speakers in the environment without requiring pre-defined labels, resolving the contradiction between measurement precision and adaptability to unknown speaker counts.
Solution Approach 2:
The patent changes the fundamental parameter of speaker count from a fixed input requirement to a variable that the system automatically determines. By using neural network-based unsupervised learning, the system can handle any number of speakers without requiring the number to be specified in advance, thereby improving both accuracy and adaptability.
2Reliability
If fixed dimension memory unit is used, then memory structure is simple, but it is difficult to effectively store voiceprint information of different speakers who are unregistered or infrequently appear
Solution Approach 1:
The patent replaces the fixed-dimension memory unit with a dynamic neural network-based memory structure that can adaptively store and retrieve voiceprint information for any number of speakers. This dynamic structure allows the system to effectively store information about unregistered or infrequently appearing speakers, improving reliability while maintaining adaptability.
3Reliability
If supervised learning method is used, then speech signal can be processed, but the method produces unreliable results in real world application due to interference
Solution Approach 1:
The patent replaces the traditional supervised learning mechanical system with a neural network-based unsupervised learning system. This substitution enables the system to automatically learn and adapt to various environmental interference conditions without requiring manually labeled training data, thereby improving reliability in real-world applications with interference.
Data Source
AI summary
An auditory selection method based on a memory and attention model, including: step S1, encoding an original speech signal into a time-frequency matrix; step S2, encoding and transforming the time-frequency matrix to convert the matrix into a speech vector; step S3, using a long-term memory unit to store a speaker and a speech vector corresponding to the speaker; step S4, obtaining a speech vector corresponding to a target speaker, and separating a target speech from the original speech signal through an attention selection model. A storage device includes a plurality of programs stored in the storage device. The plurality of programs are configured to be loaded by a processor and execute the auditory selection method based on the memory and attention model. A processing unit includes the processor and the storage device.


