The invention discloses a training method of a voice recognition model based on
noise deconstruction, a voice recognition method, a device, equipment and a medium. According to the method, a staged training strategy that a
noise unwrapping module is firstly isolated and trained and then a Conformer-
Transducer architecture is finely adjusted and trained is adopted, so that high calculation complexity and training difficulty caused by simultaneous training of a plurality of complex modules are avoided. In the isolation training stage, the performance of the
noise unwrapping module can be quickly optimized; in the
fine tuning training stage, the trained noise unwrapping module is utilized to concentrate on optimizing the Conformer-
Transducer architecture, so that the training efficiency is improved, and the
training time and the consumption of computing resources are reduced. In the isolation training and
fine tuning training process, the parameters of part of modules are frozen, the number of parameters needing to be optimized is reduced, and therefore the calculation complexity is reduced.
Noise and pure voice in a voice
signal are deconstructed through the noise unwrapping module, and accurate semantic understanding is carried out in combination with a Conformer-
Transducer architecture, so that the whole voice recognition model has higher robustness to noise.