The invention relates to a training method, an audio-visual segmentation method, an electronic device and a storage medium, and the method comprises the steps: obtaining a training sample which comprises an
audio signal, an image, a target semantic tag and a segmentation tag; extracting audio features and image features, and performing self-attention enhancement on the audio features based on the target semantic tag to apply target
semantic consistency constraint to obtain enhanced audio features; performing cross attention reweighting on the unenhanced audio features based on the guidance of the enhanced audio features to obtain target pointing audio features; bidirectional interactive fusion is executed based on the image features and the target pointing audio features, and a sparse self-
attention network is adopted in the image
feature fusion process to suppress unmatched image region response; generating a segmentation prediction result based on the fused image features and audio features, and training to obtain an audiovisual segmentation model; and inputting a to-be-processed video into the trained audio-visual segmentation model, and outputting a segmentation result. The
image segmentation accuracy of the sounding target can be improved.