The invention discloses a multi-language deep synthesis voice adaptive detection method in cross-border communication, and the method comprises the steps: S100, receiving a cross-border voice
stream and associated
metadata in real time through a communication protocol interface, the
metadata comprising equipment information, a network protocol packet header and session
time sequence information; s200, constructing a multi-
modal feature extraction pipeline, wherein multi-
modal features comprise an audio mode, a
text mode and a behavior mode; s300, inputting the features into a multi-
modal adaptive fusion engine: outputting a language tag and confidence c through a
language recognition module, and dynamically adjusting modal weight;
feature fusion is realized by adopting a gating multi-mode unit; s400, executing by a detection decision-making layer, and S500, when the detection confidence coefficient is gt; and when 90%,
incremental learning is triggered. According to the method, weight dynamic allocation driven by languages is adopted, and cross-language BERT semantic
verification is combined, so that the low-resource language detection accuracy is greatly improved, VPN disguise is effectively recognized, the generative
attack recognition rate is improved, and the success rate of confronting sample attacks is greatly reduced.