Audiovisual Salience Prediction Method and System Based on Collaborative Attention Mechanism

By adopting collaborative attention mechanism and space-time adversarial learning in audio-visual significance prediction, the alignment and fusion of audio-visual features is achieved, the problem of space-time inconsistency is solved, and the model performance and training efficiency are improved.

CN119988894BActive Publication Date: 2025-06-13JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510470452.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has shortcomings in the space-time alignment of audio-visual features, especially in the case of space-time inconsistent audio-visual features, resulting in a degradation of model performance.

Method used

The audio-visual significance prediction method based on the collaborative attention mechanism is adopted to fusion of features through visual perception guide and audio-aware guide, and the spatial and temporal adversarial learning and feature fusion modules are used to realize the alignment and fusion of audio-visual features.

Benefits of technology

It effectively captures the long-term dependence relationship between video frames and the potential semantic relationship between audio-visual features, solves the problem of space-time aberration of audio-visual features, and significantly improves the performance and training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988894B_ABST
    Figure CN119988894B_ABST
Patent Text Reader

Abstract

The present invention proposes an audio-visual saliency prediction method and system based on a collaborative attention mechanism. The method includes: obtaining preprocessed frame images and processed audio signals; extracting features from the preprocessed frame images through visual coding to obtain high-level visual features; obtaining preliminary audio features based on the processed audio signals; processing the preliminary audio features through an audio temporal extractor to obtain audio saliency features; obtaining visual-audio fusion features and audio-visual fusion features through the high-level visual features and the audio saliency features; obtaining aligned and fused audio-visual features based on the visual-audio fusion features and the audio-visual fusion features; and obtaining a saliency prediction map based on the aligned and fused audio-visual features. The present invention adopts a frame-by-frame strategy to fuse audio-visual features, accurately aligns audio-visual features in space and time, and no longer relies on pre-training of video datasets, and finally accurately locates salient targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and multimedia digital image processing, and particularly relates to an audiovisual saliency prediction method and system based on a collaborative attention mechanism. Background Art

[0002] Visual technologies combined with deep learning methods have greatly promoted the development of many applications, such as video object tracking, object detection, 3D vision, and three-dimensional reconstruction. As one of the basic research directions in computer vision, visual saliency detection (VSOD) focuses on identifying regions in a video that attract visual attention and has attracted the attention of many researchers in recent years. Humans rely on two main perceptual modalities, vision and audition, to receive external signals. The brain can not only process this information separately but also integrate audiovisual signals to achieve the synergistic effect of multiple perceptual channels, thereby better understanding the surrounding environment and enhancing the cognitive and response abilities to the environment. Although significant progress has been made in image and video saliency prediction, the research on audiovisual saliency prediction is still in its infancy and is mainly divided into artificial fusion methods and deep learning methods.

[0003] Artificial fusion methods usually integrate audio and visual features based on multiplication operations. Some researchers have achieved effective fusion of multimodal information by using localization techniques, while others have explored how to predict the gaze movement of observers through audiovisual fusion in natural dialogue scenarios, and further deeply understand the distribution of human attention. After entering the era of deep learning, researchers began to explore how to simulate the human audiovisual mechanism through deep learning models. Some researchers designed a classifier by manually annotating the audiovisual consistency labels of videos to regulate the fusion process of audiovisual information, enabling it to learn clear corresponding relationships, thereby enhancing the performance of the model.

[0004] Although the above methods have made significant progress in the effective fusion of audiovisual features and multimodal information processing, their premise is that the audiovisual features have spatio-temporal consistency, thus ignoring the situation of spatio-temporal inconsistency. However, in reality, the spatio-temporal misalignment of audiovisual features is widespread. For example, the sound in a video may come from the background rather than the target object. In this case, blindly fusing audiovisual information may instead reduce the performance of the model. Summary of the Invention

[0005] In view of the above situation, the main purpose of the present invention is to propose an audiovisual saliency prediction method and system based on a collaborative attention mechanism to solve the above technical problems.

[0006] The present invention proposes an audiovisual saliency prediction method based on a collaborative attention mechanism, and the method includes the following steps:

[0007] Step 1: Obtain each frame image in the video, preprocess each frame image to obtain the preprocessed frame image;

[0008] Preprocess the audio signal to obtain the processed audio signal;

[0009] Step 2: Extract features from the preprocessed frame image through visual coding to obtain visual features at four scales;

[0010] Extract audio features from the processed audio signal through audio coding to obtain preliminary audio features;

[0011] Process the preliminary audio features through an audio temporal extractor to obtain audio salient features;

[0012] Step 3: Process the high-level visual features and audio salient features among the four-scale visual features through a visual perception guide to obtain visual-audio fusion features;

[0013] Process the high-level visual features and audio salient features through an audio perception guide to obtain audio-visual fusion features;

[0014] Step 4: Perform spatio-temporal adversarial learning and feature fusion on the visual-audio fusion features and audio-visual fusion features in sequence to obtain the aligned and fused audio-visual features;

[0015] Step 5: Process the aligned and fused audio-visual features through a multi-layer decoder to obtain a salient prediction map.

[0016] The present invention also proposes an audio-visual saliency prediction system based on a collaborative attention mechanism, and the system includes:

[0017] A preprocessing module, used for:

[0018] Obtain each frame image in the video, preprocess each frame image to obtain the preprocessed frame image;

[0019] Preprocess the audio signal to obtain the processed audio signal;

[0020] An encoder extraction module, used for:

[0021] Extract features from the preprocessed frame image through visual coding to obtain visual features at four scales;

[0022] Extract audio features from the processed audio signal through audio coding to obtain preliminary audio features;

[0023] Process the preliminary audio features through an audio temporal extractor to obtain audio salient features;

[0024] A feature fusion module, configured to:

[0025] Process the high-level visual features and audio salient features among the visual features at four scales through a visual perception guide to obtain visual-audio fusion features;

[0026] Process the high-level visual features and audio salient features through an audio perception guide to obtain audio-visual fusion features;

[0027] A learning fusion module, configured to:

[0028] Perform spatio-temporal adversarial learning and feature fusion on the visual-audio fusion features and audio-visual fusion features in sequence to obtain aligned and fused audio-visual features;

[0029] A decoder module, configured to:

[0030] Process the aligned and fused audio-visual features through a multi-layer decoder to obtain a salient prediction map.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] 1. The present invention designs a feature alignment and fusion module, obtains the audio-visual fusion features and visual-audio fusion features of each frame, and transfers the visual-audio fusion features of the current frame as weights to the next frame. After multiple iterations, this method not only effectively captures the long-term dependencies between video frames, but also fully explores the potential semantic relationships between audio-visual features, thereby solving the problem of spatio-temporal misalignment of audio-visual features in the video saliency prediction task;

[0033] 2. The present invention designs an adversarial network module based on spatio-temporal adversarial loss. This module takes the multi-scale spatio-temporally aligned audio-visual features as input, and in a self-supervised manner, performs adversarial learning on visual features and audio features frame by frame, effectively suppressing the interference of background noise perturbations on the model performance, and at the same time getting rid of the limitation of pre-training on video datasets, thereby significantly improving the training efficiency of the model and contributing to the development of deep learning VSOD models.

[0034] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0035] Figure 1 It is a flowchart of the audio-visual saliency prediction method based on the collaborative attention mechanism proposed by the present invention;

[0036] Figure 2 It is a schematic diagram of the overall framework of the audio-visual saliency prediction system based on the collaborative attention mechanism proposed by the present invention. Detailed implementation manners

[0037] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0038] Referring to the following description and the accompanying drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0039] Please refer to Figure 1 , the embodiments of the present invention propose an audiovisual saliency prediction method based on a collaborative attention mechanism. The method includes the following steps:

[0040] Step 1: Obtain each frame image in the video, perform preprocessing on each frame image to obtain a preprocessed frame image;

[0041] Perform preprocessing on the audio signal to obtain a processed audio signal;

[0042] Furthermore, in this step, the method for preprocessing the image and the audio signal specifically includes the following steps:

[0043] Set the image sampling size, and the image sampling extraction size is 224×348;

[0044] Perform random rotation, size transformation, and regularization processing on each image to obtain a preprocessed frame image;

[0045] Perform short-time Fourier transform processing on the audio signal, convert the audio stream into a logarithmic Mel spectrogram of 257×111 to obtain a processed audio signal.

[0046] Step 2: Extract features from the preprocessed frame image through visual coding to obtain visual features at four scales;

[0047] Extract audio features from the processed audio signal through audio coding to obtain preliminary audio features;

[0048] Process the preliminary audio features through an audio temporal extractor to obtain audio saliency features.

[0049] Furthermore, in order to prevent overfitting, the present invention uses the training set and test machine of the existing VSOD dataset for training and testing respectively.

[0050] In step 2, the preliminary audio features are processed by an audio temporal extractor to obtain audio salient features. The relational expressions in the corresponding process are as follows:

[0051] ;

[0052] Among them, represents the audio salient features, represents the audio query vector obtained by linearly transforming the preliminary audio features, represents the audio key vector obtained by linearly transforming the preliminary audio features, represents the transpose symbol, represents the audio value vector obtained by linearly transforming the preliminary audio features, represents the dimensionality size of each attention head.

[0053] Furthermore, in this step, the preprocessed frame images are processed by the MViT V2 network to obtain visual features at four scales, denoted as , the low-level visual features and only focus on the global features, while the high-level visual features and pay more attention to the regions with rich semantics. Therefore, the alignment and fusion operations mainly target the high-level visual features;

[0054] Meanwhile, the preprocessed audio signal extracts the preliminary audio features through ResNet-18 pre-trained on VGGSound, and then inputs them into the temporal extractor to obtain the audio salient features, which facilitates the frame-by-frame alignment and fusion of subsequent audiovisual features.

[0055] Step 3: Process the high-level visual features and the audio salient features among the visual features at four scales through a visual perception guide to obtain visual-audio fusion features;

[0056] Process the high-level visual features and the audio salient features through an audio perception guide to obtain audio-visual fusion features;

[0057] In step 3, the high-level visual features and the audio salient features among the visual features at four scales are processed through a visual perception guide to obtain visual-audio fusion features. The specific steps are as follows:

[0058] Use the high-level visual features and the audio salient features to calculate the multi-modal learning weights through attention, and then obtain the high-level visual features after the attention mechanism through the multi-modal learning weights and the high-level visual features. The relational expressions in the corresponding process are as follows:

[0059] ;

[0060] Among them, represents the high-level visual feature of the -th frame after the attention mechanism, represents the -th frame of high-level visual feature, represents the multi-modal learning weight, represents the audio key vector obtained by linearly transforming the audio salient feature, represents the audio value vector obtained by linearly transforming the audio salient feature, represents the index of the frame, represents the index of the multi-scale feature hierarchy;

[0061] Based on the high-level visual feature after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion feature is obtained. The relational expression existing in the corresponding process is:

[0062] ;

[0063] Among them, represents the high-level visual feature of the -th frame, represents each high-level visual feature after the attention mechanism, represents the operation after connection, represents the weight parameter, represents the visual-audio fusion feature;

[0064] The high-level visual feature and the audio salient feature are processed by the audio perception guide to obtain the audio-visual fusion feature. The relational expression existing in the corresponding process is:

[0065] ;

[0066] Among them, represents the audio-visual fusion feature, represents the multi-modal learning weight, represents the video key vector obtained by linearly transforming the high-level visual feature, represents the video key vector obtained by linearly transforming the high-level visual feature.

[0067] Step 4: Perform spatio-temporal adversarial learning and feature fusion on the visual-audio fusion feature and the audio-visual fusion feature in sequence to obtain the aligned and fused audiovisual feature;

[0068] In step 4, the visual-audio fusion feature and the audio-visual fusion feature are successively subjected to spatio-temporal adversarial learning and feature fusion to obtain the aligned and fused audio-visual feature. Among them, an audio-visual symmetry loss is constructed through the visual-audio fusion feature and the audio-visual fusion feature, and the relational expression existing in the corresponding process is:

[0069] ;

[0070] Among them, represents the visual-audio symmetry loss function, represents calculating the similarity between features, represents stopping the gradient operation, represents L2 regularization, represents the th frame of the audio-visual fusion feature, represents the th frame of the visual-audio fusion feature;

[0071] The visual-audio fusion feature and the audio-visual fusion feature are successively subjected to spatio-temporal adversarial learning and feature fusion to obtain the aligned and fused audio-visual feature. Among them, a visual-audio symmetry loss is constructed through the visual-audio fusion feature and the audio-visual fusion feature, and the relational expression existing in the corresponding process is:

[0072] ;

[0073] Among them, represents the visual-audio symmetry loss function.

[0074] Furthermore, the calculation relational expression of the total optimization symmetry loss function in this step is:

[0075] ;

[0076] Among them, represents the total optimization symmetry loss function, represents the total number of frames.

[0077] Step 5: Process the aligned and fused audio-visual feature through multiple layers of decoders to obtain a saliency prediction map;

[0078] In step 5, the aligned and fused audio-visual feature is processed through multiple layers of decoders to obtain a saliency prediction map. Among them, the aligned and fused audio-visual feature passes through multiple layers of decoders and obtains a saliency prediction map under the guidance of a loss function. The loss function includes KL divergence and linear correlation coefficient, and the relational expression of KL divergence and linear correlation coefficient is:

[0079] ;

[0080] Among them, represents the KL divergence, represents the predicted saliency map, represents the ground truth map, represents the class label, represents the probability value of the saliency map at the spatial position, represents the probability value of the saliency map at the spatial position, represents the linear correlation coefficient, represents the covariance, represents the standard deviation.

[0081] Furthermore, the total loss function of the present invention is:

[0082] ;

[0083] Among them, represents the weight parameter, set to 1; represents the weight parameter, set to -1; represents the weight parameter, set to 1.

[0084] Please refer to Figure 2 , the embodiment of the present invention also provides an audiovisual saliency prediction system based on a collaborative attention mechanism. The system includes:

[0085] A preprocessing module, used for:

[0086] Obtain each frame image in the video, preprocess each frame image, and obtain the preprocessed frame image;

[0087] Preprocess the audio signal to obtain the processed audio signal;

[0088] An encoder acquisition module, used for:

[0089] Extract features of the preprocessed frame image through visual coding to obtain visual features at four scales;

[0090] Extract audio features of the processed audio signal through audio coding to obtain preliminary audio features;

[0091] Process the preliminary audio features through an audio temporal extractor to obtain audio saliency features;

[0092] A feature fusion module, used for:

[0093] Process the high-level visual features and audio saliency features among the visual features at four scales through a visual perception guide to obtain visual-audio fusion features;

[0094] Process the high-level visual features and audio salient features through an audio perception guide to obtain audio-visual fusion features;

[0095] A learning fusion module, configured to:

[0096] Perform spatio-temporal adversarial learning and feature fusion on the visual-audio fusion features and the audio-visual fusion features in sequence to obtain aligned and fused audio-visual features;

[0097] A decoder module, configured to:

[0098] Process the aligned and fused audio-visual features through a multi-layer decoder to obtain a salient prediction map.

[0099] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0100] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0101] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A method for audiovisual saliency prediction based on a collaborative attention mechanism, characterized in that: The method comprises the following steps: Step 1, obtaining each frame image in the video, preprocessing each frame image, and obtaining a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Step 2: Extract features of the preprocessed frame image through visual coding to obtain visual features of four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Step 3: Process the high-level visual features and audio salient features of the four-scale visual features through a visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Step 4: Perform spatiotemporal adversarial learning and feature fusion on the visual-audio fusion features and the audio-visual fusion features in sequence to obtain aligned and fused audio-visual features; Step 5: Process the aligned and fused audio-visual features through a multi-layer decoder to obtain a salient prediction map; In step 4, the visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features, wherein the audio-visual symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features, and the relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function, Indicates the similarity between calculated features. Indicates stopping the gradient operation. represents L2 regularization, Indicates Audio-visual fusion features of frames, Indicates Visual-audio fusion features of frames.

2. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 1, characterized in that: In step 2, the preliminary audio features are processed by an audio time series extractor to obtain audio salient features. The relationship between the corresponding process is: ; in, Indicates the audio salient features, represents the audio query vector obtained after the initial audio features are linearly transformed. Represents the audio key vector obtained after linear transformation of the preliminary audio features. represents the transpose symbol, Represents the audio value vector obtained after the initial audio features are linearly transformed. Indicates the dimension size of each attention head.

3. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 2, characterized in that: In step 3, the high-level visual features and audio salient features in the visual features of the four scales are processed by a visual perception guide to obtain visual-audio fusion features, which specifically includes the following sub-steps: The high-level visual features and audio salient features are used to obtain multimodal learning weights through attention calculation, and then the high-level visual features after the attention mechanism are obtained through the multimodal learning weights and high-level visual features; Based on the high-level visual features after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion features are obtained.

4. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 3, characterized in that: The multimodal learning weights are obtained by using high-level visual features and audio salient features through attention calculation, and then the high-level visual features after the attention mechanism are obtained through the multimodal learning weights and high-level visual features. The relationship between the corresponding process is: ; in, It means the first Frame high-level visual features, Indicates Frame high-level visual features, represents the multimodal learning weights, Represents the audio key vector obtained by linear transformation of audio salient features, Represents the audio value vector obtained by linear transformation of audio salient features, Indicates the index of the frame, Represents the index of the multi-scale feature level.

5. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 3, characterized in that: Based on the high-level visual features after the attention mechanism, the output of the current frame is used as the weight value and added to the input of the next frame. After a certain number of iterations, the visual-audio fusion features are obtained. The relationship between the corresponding process is: ; in, Indicates High-level visual features of the frame, Represents the high-level visual features after the attention mechanism, Indicates that after the connection operation, represents the weight parameter, Represents visual-audio fusion features.

6. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 5, characterized in that: In step 3, the high-level visual features and audio salient features are processed by the audio perception guide to obtain the audio-visual fusion features. The relationship between the corresponding process is: ; in, represents the audio-visual fusion feature, represents the multimodal learning weights, Represents the video key vector obtained by linear transformation of high-level visual features, Represents the video key vector obtained by linear transformation of high-level visual features.

7. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 1, characterized in that: In step 4, the visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features, wherein the visual-audio symmetric loss is constructed by the visual-audio fusion features and the audio-visual fusion features, and the relationship between the corresponding processes is: ; in, represents the visual-audio symmetric loss function.

8. The audiovisual saliency prediction method based on the collaborative attention mechanism according to claim 7, characterized in that: In step 5, the aligned and fused audiovisual features are processed by a multi-layer decoder to obtain a significant prediction map, wherein the aligned and fused audiovisual features are processed by a multi-layer decoder to obtain a significant prediction map under the guidance of a loss function, and the loss function includes a KL divergence and a linear correlation coefficient, and the relationship between the KL divergence and the linear correlation coefficient is: ; in, represents the KL divergence, represents the predicted saliency map, Represents the real picture, represents the category label, represents the probability value of the saliency map at the spatial position, represents the probability value of the saliency map at the spatial position, represents the linear correlation coefficient, represents the covariance, Represents standard deviation.

9. A system for predicting audiovisual saliency based on a collaborative attention mechanism, characterized in that: The system applies any one of the audiovisual saliency prediction methods based on the collaborative attention mechanism of claims 1 to 8, and the system comprises: Preprocessing module for: Acquire each frame image in the video, preprocess each frame image, and obtain a preprocessed frame image; Preprocessing the audio signal to obtain a processed audio signal; Encoder module for: The preprocessed frame images are subjected to feature extraction through visual coding to obtain visual features at four scales; The processed audio signal is subjected to audio feature extraction through audio coding to obtain preliminary audio features; Process the preliminary audio features through the audio time series extractor to obtain audio salient features; Feature fusion module, used for: The high-level visual features and audio salient features of the four-scale visual features are processed by the visual perception guide to obtain visual-audio fusion features; Process high-level visual features and audio salient features through an audio-aware guide to obtain audio-visual fusion features; Learning Fusion Module for: The visual-audio fusion features and the audio-visual fusion features are sequentially subjected to spatiotemporal adversarial learning and feature fusion to obtain aligned and fused audio-visual features; Decoder module for: The aligned and fused audio-visual features are processed by multiple layers of decoders to obtain a saliency prediction map.

Citation Information

Patent Citations

  • Video saliency region detection method based on multi-data-set collaborative learning

    CN116030077A

  • Regional dynamic dimming method based on audiovisual fusion saliency

    CN118737069A