The invention provides a multi-path voice
stream real-time separation and
content retrieval method, which comprises the following steps of: constructing a local voiceprint
library by using a registered voice sample, and extracting and normalizing voiceprint characteristics of a target speaker through a deep voiceprint
encoder; performing frame-level feature analysis on the mixed voice
signal, calculating a semantic distance between a current voice frame and a target voiceprint, and generating a dynamic
confidence score; feature conflict detection is carried out by combining a confidence coefficient threshold value and a de-
jitter mechanism, and a conflict
perception attention gating module is triggered to realize
frequency band enhancement and suppression; a continuous feature
confusion region is dynamically identified based on a sliding window, and the separation precision of the
confusion region is improved through local reconstruction and time-frequency
mask optimization iteration. According to the method, high-precision recognition and separation of the voice of the target speaker in the mixed voice are realized, non-target components are effectively inhibited, and the separation effect and robustness are improved.