The invention discloses a real-time
dysarthria speech
restoration method based on progressive model
distillation, and belongs to the technical field of speech
signal processing. The invention provides a combined solution for the problem of double interference of voice degradation and
environmental noise faced by disabled old people in a complex
nursing environment. Firstly, a pseudo-parallel corpus is constructed through a self-supervised repair strategy, an ideal target speech is generated by using a high-precision speech conversion model, and the problem of
truth value missing in a
pathological speech enhancement task is solved. Secondly, constructing a complex convolutional
loop network based on voiceprint embedding, and introducing a speaker embedding vector to suppress non-target human voice interference; meanwhile, a gradual
receptive field perception distillation strategy is adopted, a non-causal wide
convolution teacher model is compressed into a causal narrow
convolution student model through micro sparsification, and real-
time processing of an embedded terminal is realized. And finally, phoneme
perception loss is introduced, and semantic level supervision is performed by using the pre-trained ASR model, so that the intelligibility of the repaired voice is remarkably improved.