The invention discloses a voice authentic identification method based on dual-track difference modeling, and belongs to the technical field of voice authentic identification. The method comprises the following steps: converting
monaural audio into
stereophonic sound through a pre-trained
monaural and dual-channel voice conversion model, extracting left and right channel Mel spectrum features, and performing
texture enhancement optimization
processing on an absolute value of a difference between the left and right channel Mel spectrum features and an original
monaural Mel spectrum feature. The method comprises the following steps: firstly, optimizing the voice authenticity, respectively inputting an optimization result into a left double-
branch feature extractor and a right double-
branch feature extractor as input of double-
branch feature extractors, carrying out extraction,
processing and
information fusion on feature information to obtain a final attention map, and inputting the final attention map into an attention
pooling layer and a final dichotomy layer to obtain a voice authenticity discrimination result. According to the invention, two-track differential modeling, a fine-grained
texture enhancement method and a multi-head attention
feature fusion method are adopted to carry out authentic identification on the voice, so that the method has higher accuracy, mobility and generalization.