The present disclosure relates to an audio-visual speech
separation method and device,
electronic equipment and storage medium, video information including target object sound and at least one reference object sound is acquired, and an
image frame sequence composed of target object lip image frames in the video information and mixed audio are extracted, and the video feature and the audio feature corresponding to the target object are obtained by encoding respectively. The video feature and the audio feature are input into a trained multi-
modal separation network, and the sound
mask corresponding to the target object is obtained after multiple feature fusions. The multi-
modal separation network includes a top module, a middle module and a bottom module for three times of
feature fusion. The target audio recording the target object sound is determined according to the sound
mask and the audio feature. The present disclosure fuses the information of two levels of vision and hearing multiple times through three modules, enhances the intra-
modal context information, improves the audio-visual separation performance, and obtains an accurate audio
separation result.