The invention provides a voice
processing method and
system for mixed voice separation,
noise reduction and voice recognition, and relates to the technical field of voice
processing. The invention discloses the voice
processing method and
system for mixed voice separation,
noise reduction and voice recognition. The
system is mainly composed of a voiceprint
database construction module, a mixed voice separation module, a self-adaptive
noise suppression module and a voice recognition module. A CNN-RNN model is adopted to extract voiceprint features. Separating the mixed voice through a
deep learning model based on an
encoder-decoder structure and an attention mechanism; a
noise estimation algorithm based on minimum statistics and a multi-
modal noise reduction method are adopted to suppress noise; mFCC, LPCC and depth features are extracted from the noise-reduced voice and fused, and after optimization of an auto-
encoder, an improved
support vector machine (SVM) or a deep
neural network classifier is used for recognition. According to the method and the device, the defects in the aspects of mixed voice separation,
noise reduction, voiceprint recognition and processing efficiency in the prior art are overcome, the known voiceprint voice can be accurately separated, the noise is effectively reduced, high-precision voice recognition is realized, and the real-
time processing requirement is met.