This invention discloses an audio-video segmentation method,
system, and medium based on global temporal mixing and multi-scale audio input, belonging to the field of
multimedia signal processing and
information fusion technology. First, video data is divided into audio data and a continuous sequence of video frames, and acoustic and video features are extracted. Then, visual and acoustic features are input into a global temporal audio-video mixer to perform cross-
modal global temporal fusion, obtaining non-homogeneous acoustic query features and visual features. Finally, a
mask prediction decoder is used to achieve accurate matching between acoustic query features and visual features. This invention breaks through the dimensionality limitations of traditional fusion methods by designing a GTAVM module, significantly improving the accuracy of target segmentation; simultaneously, by introducing the MSAI mechanism, the model's adaptability to complex sound environments is enhanced.