The invention relates to an
automatic speech recognition method in a complex environment, which comprises the following steps of: firstly, carrying out availability labeling classification on acquired speech data, and firstly excluding unavailable speech data; and then searching a comprehensive optimal VAD model to filter
voice data, filtering out a non-voice part, and leaving effective
voice data. Through the basic
speech recognition model, preliminary automatic recognition and speaker labeling are carried out, dialects,
environmental noise and background music are further recognized, and according to professional terms possibly appearing in different occasions, the actual application scene of the recognition model is generalized. Finally, secondary correction is carried out manually for model optimization, in the model optimization, through speaker voiceprint recognition, the voice of a speaker can be positioned more accurately, and environment
noise and background music are regarded as individual speakers for
identity recognition, so that the recognition precision and performance of the training model are remarkably improved. The method has more accurate
speech recognition performance, and has better generalization at the same time.