Audio-visual saliency prediction method and system based on collaborative attention mechanism
By adopting collaborative attention mechanism and space-time adversarial learning in audio-visual significance prediction, the problem of insufficient space-time alignment of audio-visual features is solved, and more efficient audio-visual information fusion and model performance improvement are achieved.
Patent Information
- Application Number
- CN202510470452.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art has shortcomings in the space-time alignment of audio-visual features, resulting in the fusion of audio-visual information that may reduce model performance in the case of space-time inconsistency.
The audio-visual significance prediction method based on the collaborative attention mechanism is adopted, and visual and audio features are processed through the visual perception guide and the audio-aware guide, combining space-time adversarial learning and feature fusion to obtain the aligned and fused audio-visual features.
Effectively capture the long-term dependence between video frames and the potential semantic relationship between audio-visual features, solving the problem of audio-visual features not aligned in space-time and space, and improving the training efficiency and performance of the model.
Smart Images

Figure CN119988894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and multimedia digital image processing, and in particular to a method and system for predicting audiovisual saliency based on a collaborative attention mechanism. Background Art
[0002] Visual technology combined with deep learning methods has greatly promoted the development of many applications, such as video object tracking, object detection, 3D vision and 3D reconstruction. As one of the basic research directions of computer vision, visual saliency detection (VSOD) focuses on identifying areas in videos that attract visual attention and has attracted the attention of many researchers in recent years. Humans rely on two main perceptual methods, vision and hearing, to receive external signals. The brain can not only process this information separately, but also integrate audiovisual signals to achieve the synergy of multiple perceptual channels, so as to better understand the surrounding environment and improve the ability to perceive and respond to the environment. Although significant progress has been made in image and video saliency prediction, the research on audiovisual saliency prediction is still in its infancy, mainly divided into artificial fusion methods and deep learning methods.
[0003] Artificial fusion methods usually integrate audio and visual features based on multiplication operations. Some researchers have achieved effective fusion of multimodal information by using positioning technology, while others have explored how to predict the observer's gaze movement through audio-visual fusion in natural dialogue scenarios, thereby gaining a deeper understanding of the distribution of human attention. After entering the era of deep learning, researchers began to explore how to simulate human audio-visual mechanisms through deep learning models. Some researchers manually annotated the audio-visual consistency labels of videos and designed a classifier to regulate the fusion process of audio-visual information so that it can learn clear correspondences, thereby enhancing the performance of the model.
[0004] Although the above methods have made significant progress in the effective fusion of audiovisual features and multimodal information processing, their premise is that the audiovisual features have temporal and spatial consistency, thus ignoring the inconsistency of time and space. However, in reality, the misalignment of audiovisual features in time and space is common. For example, the sound in a video may come from the background rather than the target object. In this case, blindly fusing audiovisual information may reduce model performance. Summary of the invention
[0005] In view of the above situation, the main purpose of the present invention is to propose an audiovisual saliency prediction method and system based on a collaborative attention mechanism to solve the above technical problems.
[0006] The present invention proposes an audiovisual saliency prediction method based on a collaborative attention mechanism, the method comprising the following steps: Step 1: Obtain each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Step 2: Extract features of the preprocessed frame image through visual coding to obtain visual features of four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Step 3: Process the high-level visual features and audio salient features of the four-scale visual features through a visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Step 4: Perform spatiotemporal adversarial learning and feature fusion on the visual-audio fusion features and the audio-visual fusion features in sequence to obtain aligned and fused audio-visual features; Step 5: The aligned and fused audio-visual features are processed by a multi-layer decoder to obtain a salient prediction map.
[0007] The present invention also proposes an audiovisual saliency prediction system based on a collaborative attention mechanism, the system comprising: Preprocessing module for: Acquire each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Encoder module for: The preprocessed frame images are subjected to feature extraction through visual coding to obtain visual features at four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Feature fusion module, used for: The high-level visual features and audio salient features of the four-scale visual features are processed by the visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Learning Fusion Module for: The visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features; Decoder module for: The aligned and fused audio-visual features are processed by multiple layers of decoders to obtain a saliency prediction map.
[0008] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention designs a feature alignment fusion module, obtains the audio-visual fusion features and visual-audio fusion features of each frame, and passes the visual-audio fusion features of the current frame as weights to the next frame. After multiple iterations, this method not only effectively captures the long-term dependencies between video frames, but also fully explores the potential semantic relationships between audio-visual features, thereby solving the problem of audio-visual features not being aligned in time and space in the video saliency prediction task; 2. The present invention designs an adversarial network module based on spatiotemporal adversarial loss. This module uses multi-scale spatiotemporal aligned audiovisual features as input, and conducts adversarial learning of visual features and audio features frame by frame in a self-supervised manner, effectively suppressing the interference of background noise disturbance on model performance, while getting rid of the limitation of pre-training on video data sets, thereby significantly improving the training efficiency of the model and helping to promote the development of deep learning VSOD models.
[0009] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A flowchart of the audiovisual saliency prediction method based on the collaborative attention mechanism proposed by the present invention; Figure 2 Schematic diagram of the overall framework of the audiovisual saliency prediction system based on the collaborative attention mechanism proposed in the present invention. DETAILED DESCRIPTION
[0011] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0012] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0013] See also Figure 1 The embodiment of the present invention proposes an audiovisual saliency prediction method based on a collaborative attention mechanism, the method comprising the following steps: Step 1: Obtain each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Furthermore, in this step, the method for preprocessing the image and audio signals specifically includes the following steps: Set the image sampling size, the image sampling extraction size is 224×348; Perform random rotation, size transformation and regularization on each image to obtain a preprocessed frame image; The audio signal is subjected to short-time Fourier transform processing, and the audio stream is converted into a 257×111 logarithmic Mel spectrum to obtain a processed audio signal.
[0014] Step 2: Extract features of the preprocessed frame image through visual coding to obtain visual features of four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; The preliminary audio features are processed by an audio time series extractor to obtain audio salient features.
[0015] Furthermore, in order to prevent overfitting, the present invention uses the training set and test set of the existing VSOD data set for training and testing respectively.
[0016] In step 2, the preliminary audio features are processed by the audio time series extractor to obtain audio salient features. The corresponding relationship in the process is: ; in, Indicates the audio salient features, represents the audio query vector obtained after the initial audio features are linearly transformed. Represents the audio key vector obtained after linear transformation of the preliminary audio features. represents the transpose symbol, Represents the audio value vector obtained after the initial audio features are linearly transformed. Indicates the dimension size of each attention head.
[0017] Furthermore, in this step, the preprocessed frame image is processed by the MViTV2 network to obtain visual features of four scales, which are recorded as , underlying visual features and Only global features are considered, while high-level visual features and It pays more attention to the areas with rich semantics, so the alignment and fusion operations mainly target high-level visual features; At the same time, the preprocessed audio signal is used to extract preliminary audio features through ResNet-18 pre-trained on VGGSound, and then input into the time series extractor to obtain audio salient features, which facilitates the subsequent frame-by-frame alignment and fusion of audio-visual features.
[0018] Step 3: Process the high-level visual features and audio salient features of the four-scale visual features through a visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; In step 3, the high-level visual features and audio salient features of the four-scale visual features are processed by the visual perception guide to obtain visual-audio fusion features. The specific steps are as follows: The multimodal learning weights are obtained by using high-level visual features and audio salient features through attention calculation, and then the high-level visual features after the attention mechanism are obtained through the multimodal learning weights and high-level visual features. The relationship between the corresponding process is: ; in, It means the first Frame high-level visual features, Indicates Frame high-level visual features, represents the multimodal learning weights, Represents the audio key vector obtained by linear transformation of audio salient features, Represents the audio value vector obtained by linear transformation of audio salient features, Indicates the index of the frame, Represents the index of the multi-scale feature level; Based on the high-level visual features after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion features are obtained. The relationship between the corresponding process is: ; in, Indicates High-level visual features of the frame, Represents the high-level visual features after the attention mechanism, Indicates that after the connection operation, represents the weight parameter, Represents visual-audio fusion features; The high-level visual features and audio salient features are processed by the audio perception guide to obtain the audio-visual fusion features. The relationship between the corresponding process is: ; in, represents the audio-visual fusion feature, represents the multimodal learning weights, Represents the video key vector obtained by linear transformation of high-level visual features, Represents the video key vector obtained by linear transformation of high-level visual features.
[0019] Step 4: Perform spatiotemporal adversarial learning and feature fusion on the visual-audio fusion features and the audio-visual fusion features in sequence to obtain aligned and fused audio-visual features; In step 4, the visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features, wherein the audio-visual symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features, and the relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function, represents the similarity between calculated features. Indicates stopping the gradient operation. represents L2 regularization, Indicates Audio-visual fusion features of frames, Indicates Visual-audio fusion features of frames; The visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain the aligned and fused audio-visual features. The visual-audio symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features. The relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function.
[0020] Furthermore, the calculation relationship of the total optimized symmetric loss function in this step is: ; in, represents the total optimized symmetric loss function, Indicates the total number of frames.
[0021] Step 5: Process the aligned and fused audio-visual features through a multi-layer decoder to obtain a salient prediction map; In step 5, the aligned and fused audiovisual features are processed by multiple layers of decoders to obtain a significant prediction map, wherein the aligned and fused audiovisual features are processed by multiple layers of decoders to obtain a significant prediction map under the guidance of a loss function, and the loss function includes KL divergence and linear correlation coefficient, and the relationship between KL divergence and linear correlation coefficient is: ; in, represents the KL divergence, represents the predicted saliency map, Represents the real picture, represents the category label, represents the probability value of the saliency map at the spatial position, represents the probability value of the saliency map at the spatial position, represents the linear correlation coefficient, represents the covariance, Represents standard deviation.
[0022] Furthermore, the total loss function of the present invention is: ; in, Represents the weight parameter, set to 1; Represents the weight parameter, set to -1; Represents the weight parameter, set to 1.
[0023] See also Figure 2 , an embodiment of the present invention further provides an audiovisual saliency prediction system based on a collaborative attention mechanism, the system comprising: Preprocessing module for: Acquire each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Encoder module for: The preprocessed frame images are subjected to feature extraction through visual coding to obtain visual features at four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Feature fusion module, used for: The high-level visual features and audio salient features of the four-scale visual features are processed by the visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Learning Fusion Module for: The visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features; Decoder module for: The aligned and fused audio-visual features are processed by multiple layers of decoders to obtain a saliency prediction map.
[0024] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0025] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0026] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for audiovisual saliency prediction based on a collaborative attention mechanism, characterized in that: The method comprises the following steps: Step 1, obtaining each frame image in the video, preprocessing each frame image, and obtaining a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Step 2: Extract features of the preprocessed frame image through visual coding to obtain visual features of four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Step 3: Process the high-level visual features and audio salient features of the four-scale visual features through a visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Step 4: Perform spatiotemporal adversarial learning and feature fusion on the visual-audio fusion features and the audio-visual fusion features in sequence to obtain aligned and fused audio-visual features; Step 5: The aligned and fused audio-visual features are processed by a multi-layer decoder to obtain a salient prediction map.
2. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 1, characterized in that: In step 2, the preliminary audio features are processed by an audio time series extractor to obtain audio salient features. The relationship between the corresponding process is: ; in, Indicates the audio salient features, represents the audio query vector obtained after the initial audio features are linearly transformed. Represents the audio key vector obtained after linear transformation of the preliminary audio features. represents the transpose symbol, Represents the audio value vector obtained after the initial audio features are linearly transformed. Indicates the dimension size of each attention head.
3. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 2, characterized in that: In step 3, the high-level visual features and audio salient features in the visual features of the four scales are processed by a visual perception guide to obtain visual-audio fusion features, which specifically includes the following sub-steps: The high-level visual features and audio salient features are used to obtain multimodal learning weights through attention calculation, and then the high-level visual features after the attention mechanism are obtained through the multimodal learning weights and high-level visual features; Based on the high-level visual features after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion features are obtained.
4. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 3, characterized in that: The multimodal learning weights are obtained by using high-level visual features and audio salient features through attention calculation, and then the high-level visual features after the attention mechanism are obtained through the multimodal learning weights and high-level visual features. The relationship between the corresponding process is: ; in, It means the first Frame high-level visual features, Indicates Frame high-level visual features, represents the multimodal learning weights, Represents the audio key vector obtained by linear transformation of audio salient features, Represents the audio value vector obtained by linear transformation of audio salient features, Indicates the index of the frame, Represents the index of the multi-scale feature level.
5. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 3, characterized in that: Based on the high-level visual features after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion features are obtained. The relationship between the corresponding process is: ; in, Indicates High-level visual features of the frame, Represents the high-level visual features after the attention mechanism, Indicates that after the connection operation, represents the weight parameter, Represents visual-audio fusion features.
6. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 5, characterized in that: In step 3, the high-level visual features and audio salient features are processed by the audio perception guide to obtain the audio-visual fusion features. The relationship between the corresponding process is: ; in, represents the audio-visual fusion feature, represents the multimodal learning weights, Represents the video key vector obtained by linear transformation of high-level visual features, Represents the video key vector obtained by linear transformation of high-level visual features.
7. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 6, characterized in that: In step 4, the visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features, wherein the audio-visual symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features, and the relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function, Indicates the similarity between calculated features. Indicates stopping the gradient operation. represents L2 regularization, Indicates Audio-visual fusion features of frames, Indicates Visual-audio fusion features of frames.
8. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 7, characterized in that: In step 4, the visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features, wherein the visual-audio symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features, and the relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function.
9. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 8, characterized in that: In step 5, the aligned and fused audiovisual features are processed by a multi-layer decoder to obtain a significant prediction map, wherein the aligned and fused audiovisual features are processed by a multi-layer decoder to obtain a significant prediction map under the guidance of a loss function, and the loss function includes a KL divergence and a linear correlation coefficient, and the relationship between the KL divergence and the linear correlation coefficient is: ; in, represents the KL divergence, represents the predicted saliency map, Represents the real picture, represents the category label, represents the probability value of the saliency map at the spatial position, represents the probability value of the saliency map at the spatial position, represents the linear correlation coefficient, represents the covariance, Represents standard deviation.
10. A system for predicting audiovisual saliency based on a collaborative attention mechanism, characterized in that: The system applies any one of the audiovisual saliency prediction methods based on the collaborative attention mechanism of claims 1 to 9, and the system comprises: Preprocessing module for: Acquire each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Encoder module for: The preprocessed frame images are subjected to feature extraction through visual coding to obtain visual features at four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Feature fusion module, used for: The high-level visual features and audio salient features of the four-scale visual features are processed by the visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Learning Fusion Module for: The visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features; Decoder module for: The aligned and fused audio-visual features are processed by multiple layers of decoders to obtain a saliency prediction map.
Citation Information
Patent Citations
Video quality evaluation method, system and terminal based on audio-visual joint attention
CN111479109A
Weak supervision time sequence behavior positioning method based on adversarial learning
CN114842402A
Video saliency region detection method based on multi-data-set collaborative learning
CN116030077A
Regional dynamic dimming method based on audiovisual fusion saliency
CN118737069A
Audiovisual secondary haptic signal reconstruction method based on cloud-edge collaboration
US20230290234A1
Cited By
Cross-modal saliency prediction method and device for audio visual touch
CN121278634A