This application relates to a method,
system, device, and medium for separating overlapping speech in video conferencing based on AI
noise suppression, belonging to the field of
online video conferencing service technology. The method for separating overlapping speech in video conferencing includes: acquiring and cleaning the mixed audio
stream of participants, segmenting and converting it into a short-
time frame spectrogram; extracting the voiceprint features from the short-
time frame spectrogram to obtain a general acoustic
feature vector and a participant representation vector; separating and mapping the general acoustic
feature vector according to a preset
sound source separation network to generate a sound source speech
mask and
record potential sound source labels; aggregating the participant representation vector and the sound source speech
mask according to the potential sound source labels to obtain a participant speech
mask; reconstructing the short-
time frame spectrogram based on the participant speech mask, and adaptively applying
gain and
noise suppression to obtain the participant speech
stream; using a
deep learning model to distinguish speech,
noise, and multiple
sound sources, achieving the separation of overlapping speech and improving speech intelligibility in noisy environments.