The invention discloses a target voice extraction method and
system and a medium, and belongs to the technical field of multi-mode voice
signal processing. Comprising the following steps: giving a mixed voice and a lip video of a target speaker, and extracting a voice logarithmic power spectrum, a cross-channel
phase difference and a visual
time sequence feature; splicing multi-
modal features, inputting the spliced multi-
modal features into an improved DPCRN network, estimating a voice
mask and deducing a
noise mask; calculating a beam forming weight through the
covariance matrix and generalized eigenvalue
decomposition to obtain beam features of voice and
noise; the logarithmic power spectrum, the visual features and the beam features are fused, and high-dimensional representation of voice and
noise is constructed; feature
mutual exclusion enhancement is realized by using a frame-level cross attention mechanism,
residual noise is removed from voice features, and leaked voice is removed from noise features; and finally, an optimization
mask is generated through a decoder, and pure target voice is output through spectrum reconstruction. According to the invention, complex noise environment target voice can be accurately separated.