Audio and visual fusion method for panoramic video saliency prediction
Through a multi-stage fusion strategy and cross-attention mechanism, the problems of spherical distortion and modal fusion in panoramic videos are solved, more accurate saliency prediction is achieved, and the ability to align audio and visual features in virtual reality environments is improved.
Patent Information
- Application Number
- CN202510595411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing panoramic video saliency prediction models have shortcomings in handling spherical distortion, maintaining global consistency, fusing audio and visual modalities, and integrating audio temporal semantic depth, resulting in insufficient prediction accuracy and robustness in virtual reality environments.
A multi-stage fusion strategy and cross-attention mechanism are adopted to obtain visual modal features through tangent image processing, and combined with the audio semantic temporal attention mechanism, the audio and visual modal features are gradually integrated, and the multi-head cross-attention mechanism and dynamic convolutional decoding are used to predict salient regions.
The alignment capability of audio and visual features in panoramic videos is improved, and the prediction accuracy and robustness of the model in immersive environments are enhanced, especially showing higher accuracy in real-time saliency prediction tasks.
Smart Images

Figure CN120656099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an audio-visual fusion method for panoramic video saliency prediction. Background Art
[0002] With the rapid development of virtual reality (VR) technology, immersive experiences are becoming increasingly important in various application areas, including entertainment, education, and healthcare. Human attention plays a key role in the perception of these immersive environments. Neuroscience research has shown that identifying regions within a scene that attract user attention (visual saliency) can effectively represent the dynamics of human attention. Traditionally, most saliency prediction models for panoramic videos have focused on the visual modality, focusing on compressing, transmitting, and rendering visual information to enhance the user experience. However, auditory experience, a crucial component of human perception in VR environments, is often overlooked in these saliency prediction models. In reality, auditory cues—whether background sounds, dialogue, or music—interact with visual stimuli to guide attentional shifts in VR environments. In panoramic videos, the relationship between audio and visual modalities becomes even more crucial, as sound sources in a 360-degree environment encourage users to focus their attention on specific sound-producing areas. Therefore, integrating audio and visual modalities into saliency prediction models is crucial for modeling more accurate and natural human attentional shifts.
[0003] Furthermore, challenges such as spherical distortion, sudden scene changes, and the alignment between audio temporal semantics and visual spatial perception hinder the effective performance of existing audiovisual saliency prediction methods in this context. Ensuring that an audiovisual multimodal panoramic video saliency prediction model can adapt to a variety of real-world scenarios and possesses good generalization capabilities to further enhance the user's immersive experience is key.
[0004] The existing technology has the following defects:
[0005] After an in-depth exploration of the research field of panoramic video saliency prediction technology, it is not difficult to find that although many researchers have invested a lot of effort in improving the performance of prediction models and attempting to incorporate audio cues, existing research results still expose some limitations that cannot be ignored. First, traditional projection-based models, such as equirectangular projection (ERP) and cubic projection techniques, while to some extent alleviate the distortion problem when converting spherical videos to flat videos, these methods increase computational complexity and easily cause global information to be broken at the junction of video frames, failing to fully utilize spatial continuity. This results in a decrease in prediction accuracy in practical applications.
[0006] Furthermore, the development of audiovisual saliency prediction technology also faces challenges. While advanced techniques such as the tangent image method effectively reduce distortion, they fail to break through the limitations of the visual modality and neglect the effective integration of audio features. This limits the model's predictive capabilities when dealing with complex audiovisual environments. In terms of audio feature extraction, although techniques such as Log-Mel spectrograms and Mel-frequency cepstral coefficients (MFCCs) have been widely used in 2D models, they fail to fully consider the spatial relationship between sound sources and visual attention in the field of panoramic video, making the alignment of spatial and temporal features an insurmountable obstacle in immersive environments.
[0007] Finally, a method for predicting panoramic video saliency by combining audio and visual information has also been proposed. Although existing models such as AVS360 and SVGC-AVA have made attempts at multimodal fusion, they often rely too much on early fusion or fail to fully explore the depth of audio temporal semantics. This processing method results in the ineffective utilization of temporal dynamic information in the video and semantic information in the audio, making the model lack of robustness in capturing attention shifts in dynamic scenes. Therefore, the shortcomings of current technology point out that a method is needed that can cleverly balance spherical distortion processing, global consistency maintenance, the fusion of audio and visual modalities, and the deep integration of audio temporal semantic information at different fusion stages, so as to achieve certain progress and breakthroughs in the field of panoramic video saliency prediction.
[0008] Given these limitations, integrating advanced multi-stage fusion strategies and cross-attention mechanisms is crucial for achieving more accurate and efficient audiovisual saliency prediction models for panoramic videos. Developing a solution that gradually fuses audio and visual modalities in multiple stages, while simultaneously addressing spherical distortion and aligning audiovisual spatiotemporal features, can significantly improve the model's ability to predict changes in human attention in panoramic videos. Summary of the Invention
[0009] In order to overcome the above-mentioned shortcomings and deficiencies of the prior art, the object of the present invention is to provide an audio-visual fusion method for saliency prediction of panoramic videos.
[0010] The purpose of the present invention is achieved through the following technical solutions:
[0011] An audio-visual fusion method for panoramic video saliency prediction includes the following:
[0012] Preprocessing stage;
[0013] Input a panoramic video and obtain video frames and audio data constituting the panoramic video;
[0014] The video frame processing stages are:
[0015] Calculate the tangent image for each set of video frames to obtain the K sets of tangent image sequences x of the panoramic video tan ∈R F×T×3×p×p ;
[0016] Obtaining the global features of the section based on the tangent image sequence;
[0017] Obtain spherical geometry-aware embedding features x based on the global features of the cut surface Tan ∈R F×T×D ;
[0018] By embedding features based on spherical geometry perception, we can obtain the spatiotemporal features of the cross-section and further obtain the visual modality features.
[0019] The audio data processing stages are:
[0020] Perform Mel frequency conversion on the audio data to obtain a logarithmic Mel spectrum, and further obtain K groups of audio data samples of the panoramic video;
[0021] Encode the audio data samples through the VGGish network to obtain global audio information features
[0022] The global audio information features The audio features are encoded through the audio semantic time attention mechanism to generate enhanced global temporal semantic information f a ;
[0023] Modal fusion stage:
[0024] A two-stage fusion strategy is used to obtain the final modal features
[0025] The final modal features Decode and extract salient areas in panoramic videos.
[0026] Furthermore, the two-stage fusion strategy is adopted to obtain the final modal features Specifically include:
[0027] Primary integration, including:
[0028] Global audio information features are aligned with visual modality features;
[0029] Slicing operation is used to reduce the computational complexity of modal fusion;
[0030] Recursive thinking completes primary fusion to obtain primary features
[0031] The final integration includes:
[0032] Use updated tokens to integrate visual modality features and global temporal semantic information f a indirect interaction;
[0033] The final modal features are obtained by using the multi-head cross-attention mechanism to fuse information after cross-modal indirect interaction.
[0034] Furthermore, the global audio information features are aligned with the visual modal features. Specifically, the number of channels of the global audio information features is changed through a linear layer, and the number of tangent images repeated in space is multiplied by the number of times their corresponding length, width and height are multiplied, so that the dimensions of the global audio information features are consistent with those of the visual modal features.
[0035] Furthermore, update tokens to integrate visual modality features with global temporal semantic information f a The indirect interaction is:
[0036] Update tokens: Randomly initialize to generate initial values, use Represents the initialized tokens, using Indicates tokens updated from the previous indirect interaction layer, and f l fusion It is a progressive relationship, and fusion tokens are updated through the back propagation of gradients;
[0037] f a and Convert them into vector form respectively, specifically: re represents the tensor rearrangement operation,
[0038] Use a linear layer to reduce visual Z v The total length of the
[0039] use and and Z a Perform calculations on the cross attention mechanism, and after internal iterative updates and fusion, obtain the visual and audio Use standard cross-attention sharing To produce the final updated visual and audio characteristics
[0040] The final audio semantic information is fused with the projected tangent image through the residual connection transformer, that is, As a query, and As key and value, and then calculated Before decoding output, it is necessary to pass the primary fusion feature Interaction features with primary Combine multiple cross-attention mechanisms to obtain the final modal features
[0041] Furthermore, dynamic convolutional decoding is adopted and ERP back-projection is used to obtain the predicted salient regions in the panoramic video.
[0042] Furthermore, the audio data is subjected to Mel frequency conversion to obtain a logarithmic Mel spectrum, and K groups of audio data samples of the panoramic video are further obtained, specifically:
[0043] The audio data is sampled at a 16KHz sampling rate to obtain audio samples;
[0044] Divide the audio samples into several frames;
[0045] Use Hann window function to weight each frame of data;
[0046] Short-time Fourier transform is used on each frame of weighted data to convert from time domain to frequency domain, and Mel frequency conversion is performed to obtain the logarithmic Mel spectrum;
[0047] Repeat the above steps to obtain K groups of audio data samples.
[0048] Furthermore, the tangent image is processed to obtain the spherical geometric perception embedding feature x Tan ∈R F×T×D , specifically:
[0049] The global features are obtained by passing the tangent image sequence through the ResNet-18 encoder, which is downsampled and flattened to obtain a tangent feature vector with dimension D = 512. At this time, the ResNet-18 encoder remains frozen during the training process.
[0050] A fully connected layer is used to map the angular coordinates (φ, θ) of each pixel of the tangent viewport to the same feature dimension and add it to the spherical position encoding (grid sampling points of the icosahedron projection) to obtain a spherical geometry-aware embedding feature.
[0051] Furthermore, the spherical geometry perception embedding feature is used to obtain the spatiotemporal features of the section, and further obtain the visual modality features. Specifically:
[0052] The spherical geometric perception embedding features are calculated in the spherical ViT respectively using the temporal and spatial attention mechanisms, and feature enhancement is performed to obtain the spatiotemporal features of the cross-section;
[0053] Concatenate the slice spatiotemporal features with the global features obtained by the ResNet-18 encoder to obtain visual modality features.
[0054] Furthermore, the audio data also includes the following processing:
[0055] The fffmpeg processing method is used to distinguish the channels of the audio data. If there are four channels, W of the four channels W, X, Y, and Z is treated as one channel, and X, Y, and Z are concatenated as another channel. Finally, they are merged into one channel and Mel frequency conversion is performed;
[0056] If there is only one channel, the Mel frequency conversion is performed directly.
[0057] Furthermore, the audio semantic time attention mechanism includes:
[0058] In the global audio information feature The time dimension increases the position encoding to learn the transformation process of time series;
[0059] The audio features are encoded through two consecutive Transformer layers to obtain the enhanced global temporal semantic information f a .
[0060] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0061] 1) This method adopts a multi-level fusion strategy. In the two-stage fusion process, the audio information is encoded in different ways to obtain global audio information and enhanced global temporal semantic information, which are then fused with different levels of visual modality features. This helps capture the association between audio and video at different levels, thereby achieving coarse-grained to fine-grained saliency prediction of attention regions guided by audio.
[0062] 2) This method uses a temporal semantic enhancement method based on a self-attention mechanism. By introducing an audio semantic temporal attention mechanism, it refines audio features to better utilize their temporal semantic information. This method enables audio features to be more closely integrated with video frame features in the final fusion stage, enhancing the model's cross-modal understanding capabilities.
[0063] 3) This method features a more efficient fusion strategy, consisting of two stages: preliminary fusion and final fusion. Preliminary fusion achieves preliminary spatial localization in panoramic videos through multi-level interaction, integrating features from different fields of view and global audio information. Final fusion, by introducing semantic cross-modal perceptual fusion, enhances the model's focus on semantic temporal information through indirect interaction. Deep attention fusion enables precise spatial perception of the audio modality in panoramic videos, thereby improving fusion accuracy.
[0064] 4) The prediction results and performance on several existing datasets demonstrate significant advantages. In particular, the model achieves higher accuracy in real-time saliency prediction tasks, demonstrating its powerful ability to simultaneously capture audio and visual features. These achievements not only provide new perspectives for research in audio and video fusion, but also offer strong technical support for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a structural schematic diagram of the present invention;
[0066] Figure 2 It is a workflow diagram of the present invention. DETAILED DESCRIPTION
[0067] The present invention will be further described in detail below with reference to the examples, but the embodiments of the present invention are not limited thereto.
[0068] Example
[0069] like Figure 1 and Figure 2 As shown, an audio-visual fusion method for panoramic video saliency prediction includes the following steps:
[0070] Preprocessing stage:
[0071] Input a panoramic video and obtain the video frames (ERP format) and audio data that constitute the panoramic video.
[0072] Before saliency prediction, we need to use an isotropic Gaussian filter on the human eye's annotation point map to generate a saliency grayscale map as the label. Specifically, we first need to resample the video to 25 frames per second to ensure consistent alignment of the comparison data labels. Following the 68-95-99.7 rule of the Gaussian distribution, σ is set to 5° to generate the corresponding saliency grayscale map.
[0073] The model will predict the last keyframe. For example, if our model inputs 16 consecutive frames, we need to predict the last frame. Since a video is composed of many frames, for example, a video with 250 frames, such as 000.jpg, 001.jpg, ..., 249.jpg, the first input is 000.jpg-015.jpg, the next is 001.jpg-016.jpg, and so on. The model then predicts the saliency map of the last frame.
[0074] Video frame processing stages
[0075] A,tangent image of a video frame is calculated.,The tangent image maps the spherical data onto an oriented square plane centered on the icosahedron subdivision facets,through Gnomonic projection. The number of square planes is determined by the number of,the basic icosahedron faces and the subdivision level.
[0076] Their spatial extent is the vertex resolution R of the icosahedron at level b-1. v (b-1) and the resolution of the image grid is given by the following equation. Let (φ f λ f ) is the coordinate of the centroid of the triangular face of the icosahedron in the spherical coordinate system, and then the boundary of the plane is calculated as the central latitude and longitude in the spherical coordinate system by the following formula (1) f λ f ) is the inverse Gnomonic projection of the point.
[0077]
[0078] The vertex resolution Rv for a level b icosahedron is obtained by calculating the average angle between all vertices and their adjacent vertices adj(v), as shown in formula (2):
[0079]
[0080] Using Rv(b_1) ensures that the tangent images fully cover the triangles they correspond to. Because vertex resolution is roughly halved at each subsequent subdivision level, define Rv(-1) = 2Rv(0).
[0081] Repeat the tangent image calculation method to obtain K groups of tangent image sequences x tang ∈R F×T×3×p×p , F is the time length of the video frame, T is the length of the tangent image sequence, and p is the size of a single tangent image. Each set of tangent image sequences is formed by projecting a video frame in ERP format to form a subsample, that is, a continuous video frame input once is composed of 16 subsamples, which are recorded as {V1, V2, V3, ..., V k}, each subsample is composed of a sequence of tangent graphs (each face of the icosahedron).
[0082] B obtains the global features of the section based on the tangent image sequence.
[0083] Assume the input ERP video clip Each video frame in the image is projected into a tangent image sequence according to the tangent image calculation method. where F, 3, H, and W are the number of frames, channel dimensions (RGB), height, and width of a given video (here in ERP format), respectively. The resulting tangent images have a patch size of p × p = 224 × 224 pixels. The number of tangent images per frame, T, and the field of view (FOV), are projection hyperparameters that vary between 10 / 20 and 120° / 80°.
[0084] The tangent image sequence is trained using the ResNet-18 encoder pre-trained on ImageNet. The global features of the slice are encoded, downsampled and flattened to obtain a tangent feature vector with dimension D = 512. At this time, the ResNet-18 encoder remains frozen during the training process.
[0085] C obtains spherical geometry-aware embedding features based on the global features of the section
[0086] Since the global features of the tangent plane are obtained, the angular coordinates (φ, θ) of each pixel of the tangent viewport are mapped to the same feature dimension using a fully connected layer, and these embeddings are added to the spherical position encoding (grid sampling points of the icosahedron projection) to obtain spherical geometry-aware embedding features.
[0087] D obtains the spatiotemporal features of the cross-section by embedding features based on spherical geometry perception.
[0088] Spherical geometry-aware embedding feature x Tan ∈R F×T×D In the spherical ViT, temporal and spatial attention mechanisms are calculated separately to obtain tangential spatiotemporal features, which are then enhanced for further modal feature fusion. We use a two-stage approach to approximate spatiotemporal attention: applying temporal attention between the same tangential viewports from F consecutive frames, and then applying spatial attention between T tangential viewports in the same frame. This reduces the overall self-attention complexity from F^2 × T^2 to F^2 + T^2. In this way, the spherical ViT method effectively models the global context required for panoramic video understanding.
[0089] E further obtains visual modality features
[0090] The local spatiotemporal features of each spherical ViT-processed section and the global features of the section extracted by the frozen image encoder are spliced into the fusion stage as the visual modality features in the multimodal
[0091] The audio data processing stages are:
[0092] A performs Mel frequency conversion on the audio data to obtain a logarithmic Mel spectrum, and further obtains K groups of audio data samples of the panoramic video. The specific steps are as follows:
[0093] First, the audio data is sampled at a 16kHz sampling rate to obtain audio samples. The audio samples are then divided into several frames. Each frame is weighted using a Hann window function. Each weighted frame is transformed from the time domain to the frequency domain using a short-time Fourier transform, followed by a Mel-frequency conversion to obtain a log-Mel spectrum. After calculating the log-Mel spectrum of the audio information, it is first processed by traditional ffmpeg methods for different channels. For four-channel audio, W is treated as one channel, and X, Y, and Z are concatenated as another channel, ultimately merging them into one channel. For monophonic audio, the audio is processed normally, with the audio directly extracted and the log-Mel spectrum calculated, ultimately serving as the model input.
[0094] Short-time Fourier transform (SIFT) is used to convert audio data into Log-Mel spectrogram to extract the time-frequency features of the sound in a form that is more consistent with the perception of the human ear.
[0095] B uses the VGGish network to encode the spectrogram and obtain the global audio information features
[0096] C In order to further explore the contextual associations of audio in the temporal dimension, we proposed an audio semantic temporal attention mechanism (TSE). Specifically, this mechanism uses learnable position encodings that are updated over time to supplement the temporal dimension of audio features, thereby enhancing the model's ability to understand the temporal dynamics of audio signals. Subsequently, the audio features are encoded through two consecutive Transformer layers, which effectively capture the long-distance dependencies and time series relationships in the audio data through the self-attention mechanism, thereby improving the representational ability of audio features in the temporal dimension. This enhanced audio semantic temporal attention mechanism not only improves the temporal resolution of audio features, but also provides richer temporal context for visual information in the multimodal fusion process, thereby optimizing the overall fusion performance. This mechanism enhances the temporal information in audio features, strengthens the semantic relationships between different time segments, and generates enhanced global temporal semantic information features. h a and w a is the height and width of the global temporal semantic information feature map of the audio, C is the number of channels, t a It is the time dimension.
[0097] Modal fusion stage
[0098] A two-stage fusion strategy is used to obtain the final modal features Specifically include:
[0099] A primary fusion is to obtain primary features, namely cross-modal features Specifically:
[0100] A1 global audio information features and visual modality features Alignment is to change the number of channels through a linear layer and repeat the number of tangent images and the product of their length, width and height in space to match the dimension of the visual modality feature. Here H v , W v Represent the height and width of a single tangent image respectively, T is the number of tangent images, F is the length of the time series, where l i Represents the i-th layer of spherical surface Vit in the initial fusion.
[0101] A2 uses slicing operations to reduce the computational complexity of modal fusion.
[0102] If the visual modality features obtained from the spherical ViT are spatially and temporally fused with the global audio information features, the length of the tokens should be the result of the tangent image containing all time series. In order to predict the final frame, slicing along the time dimension is performed before decoding to extract the last frame (the conventional cross-attention mechanism is already used here). This method leads to an increase in computational complexity, and its complexity becomes a multiple of the increase in the time series. Since the spherical ViT sphere has learned the spatiotemporal features, before applying the cross-attention mechanism for modal fusion, we slice the learned spatiotemporal features within and between video frames. This method reduces the computational complexity, and the complexity then becomes the complexity of a single time series. This significant reduction in complexity is beneficial to the initial fusion process. The entire process can be described by the following formula:
[0103]
[0104] where s(·) denotes the operation of slicing the video frames, and β(·) and γ(·) denote ordinary linear layers. Represents the visual modality input features of the L-1th layer in the initial fusion process, Represents the global audio information features obtained through VGGish decoding, and then the modal interaction is realized through the conventional cross-attention mechanism calculation.
[0105] Through audio-visual interaction, each visual pixel can correspond to the entire range of auditory information, thereby achieving efficient fusion. It represents the feature map after the cross-modal interaction between the visual input from the i-th sphere Vit and the global audio information. This fusion stage uses low-level visual features to fuse and obtain the audio-visual modality interaction features.
[0106] A3 recursive thinking completes primary fusion to obtain primary features
[0107] Primary fusion is composed of multiple interaction methods as described in step A2. Each layer will interact with the previous layer. The audio and visual features of the image are added element by element, thereby effectively utilizing the visual features of different levels in the low level, realizing the interaction between different visual features and global audio, and outputting the final primary features.
[0108] B is finally fused to obtain the final modal features before decoding Achieve accurate prediction of salient areas in panoramic videos, specifically:
[0109] B1 tokens definition and update process
[0110] This method proposes an interactive token that can be iteratively updated each time the model is trained. It is directly generated by random initialization. The random initialization is generally generated by standard Gaussian distribution or uniform distribution. Represents the initialized tokens. Represents the tokens updated from the previous indirect interaction layer, so as to update f for the next layer iteration l fusion Get ready. Here and f l fusion It is a progressive relationship. These updateable fusion tokens do not participate in the fusion process of the next link, but interact between the two modalities through their own updates, thereby outputting modalities through this indirect interactive modality method, and finally outputting the updated visual modality and audio modality to be input to the next link. This process obtains the updated fusion feature f i fusion , thus preparing for the next internal iteration of this link.
[0111] B2 uses a linear layer to reduce visual Z v The total length of the
[0112] During the training process, these tokens are continuously optimized and updated as the network backpropagates. As mentioned earlier, the visual features are obtained by outputting the last layer of the spherical ViT module. The audio self-attention mechanism module provides audio features, which can better utilize the temporal semantic information of the audio modality f a , to assist in effective perception and positioning in the video space. In order to make full use of spatial information, the cross-modal attention mechanism first needs to and f a Convert them into vector form respectively. First, re represents the tensor rearrangement operation, and if we use Z directly v as the key and value (the query in this case is the tokens in B1), and (Same as above) to calculate the cross-attention mechanism with the tokens in B1, the computational complexity will be very high. Considering the visual tokens-Z v The length of Z is greater than the length of the audio tokens-Za. A more efficient approach is to first use a simple linear layer to reduce the visual Z v The total length of tokens, thus obtaining a new set of visual tokens- This reduction enables the cross-attention mechanism to update the aggregated visual The computational complexity at this stage becomes O(N′ V N a C), where N′ V It is an updated visual The length, N a It's Audio Z a The length of tokens is , where the spatial information is retained as much as possible. and Z a .
[0113] B3 performs the final interaction between modals:
[0114] It should be noted here that the initialization in B1 It is shared, for example, it is used in B2 at the beginning and and Z a Perform cross attention mechanism calculations, each update get ...and so on Until the end of training, for example, if the model inputs 16 consecutive frames of video for the first time, after this stage It will be updated three times (assuming the number of layers is 3), then after this indirect interaction is completed The model will continue to input data. The update will be performed in the same way as above, and the whole process of learning and updating parameters will be gradually refined. Each update can integrate the information of the two modalities. After this internal iterative update fusion, the updated visual and audio Use standard cross-attention sharing To produce the final updated visual and audio characteristics This further refines the interaction results of audio semantic features in vision.
[0115] B4 Strengthen the interaction between visual and audio modalities: Here we still need to use the primary features obtained by primary fusion This updateable fusion medium further provides more accurate information for the final fusion. First, the final audio semantic information is fused with the projected tangent image patch through the residual connection transformer. As a query, and As key and value, and then calculated Before decoding output, it is necessary to pass the The primary interaction characteristics are the same Combine multiple cross-attention mechanisms to obtain the final features ,
[0116] B5 obtains the saliency map through dynamic convolution decoding: Since in B4 we fuse and as well as The final modal characteristics are obtained Therefore, dynamic convolution decoding can be used after fusion. By adaptively adjusting the convolution kernel, the model can flexibly extract key salient areas in the image or video, and then output the predicted saliency map through rectangular projection. It can dynamically adjust the convolution kernel based on the local and global features of the input, thereby more accurately capturing complex contextual relationships, differences between background and foreground, and improving accuracy in multi-scale feature fusion. In addition, dynamic convolution decoding can effectively reduce computational complexity and optimize storage overhead. Especially in dynamic scenes, it enhances the model's ability to capture temporal changes and dynamic salient areas by adaptively adjusting the convolution kernel, improving the accuracy and robustness of predictions.
[0117] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. An audio-visual fusion method for panoramic video saliency prediction, characterized in that: These include: Preprocessing stage; Input a panoramic video and obtain video frames and audio data constituting the panoramic video; The video frame processing stages are: Calculate the tangent image for each set of video frames to obtain the K sets of tangent image sequences x of the panoramic video tan ∈R F×T×3×p×p ; Obtaining the global features of the section based on the tangent image sequence; Obtain spherical geometry-aware embedding features x based on the global features of the cut surface Tan ∈R F×T×D ; By embedding features based on spherical geometry perception, we can obtain the spatiotemporal features of the cross-section and further obtain the visual modality features. The audio data processing stages are: Perform Mel frequency conversion on the audio data to obtain a logarithmic Mel spectrum, and further obtain K groups of audio data samples of the panoramic video; Encode the audio data samples through the VGGish network to obtain global audio information features The global audio information features The audio features are encoded through the audio semantic time attention mechanism to generate enhanced global temporal semantic information f a ; Modal fusion stage: A two-stage fusion strategy is used to obtain the final modal features The final modal features Decode and extract salient areas in panoramic videos.
2. The audio-visual fusion method according to claim 1, characterized in that: The two-stage fusion strategy is adopted to obtain the final modal features Specifically include: Primary integration, including: Global audio information features are aligned with visual modality features; Slicing operation is used to reduce the computational complexity of modal fusion; Recursive thinking completes primary fusion to obtain primary features The final integration includes: Use updated tokens to integrate visual modality features and global temporal semantic information f a indirect interaction; The final modal features are obtained by using the multi-head cross-attention mechanism to fuse information after cross-modal indirect interaction.
3. The audio-visual fusion method according to claim 2, characterized in that: The global audio information features are aligned with the visual modality features, specifically by changing the number of channels of the global audio information features through a linear layer and repeating the number of tangent images in space and the number of times the product of their corresponding length, width and height is multiplied, so that the dimensions of the global audio information features are consistent with those of the visual modality features.
4. The audio-visual fusion method according to claim 2, characterized in that: Use updated tokens to integrate visual modality features and global temporal semantic information f a The indirect interaction is: Update tokens: Randomly initialize to generate initial values, use Represents the initialized tokens, using Indicates tokens updated from the previous indirect interaction layer, and It is a progressive relationship, and Fusiontokens are updated through the back propagation of gradients; f a and Convert them into vector form respectively, specifically: re represents the tensor rearrangement operation, Use a linear layer to reduce visual Z v The total length of the use and and Z a Perform calculations on the cross attention mechanism, and after internal iterative updates and fusion, obtain the visual and audio Use standard cross-attention sharing To produce the final updated visual and audio characteristics The final audio semantic information is fused with the projected tangent image through the residual connection transformer. As a query, and As key and value, and then calculated Before decoding output, it is necessary to pass the primary fusion feature Interaction features with primary Combine multiple cross-attention mechanisms to obtain the final modal features 5. The audio-visual fusion method according to claim 1, characterized in that: Dynamic convolutional decoding is adopted and ERP back-projection is used to obtain the predicted salient regions in the panoramic video.
6. The audio-visual fusion method according to claim 1, characterized in that: The audio data is subjected to Mel frequency conversion to obtain a logarithmic Mel spectrum, and K groups of audio data samples of the panoramic video are further obtained, specifically: The audio data is sampled at a 16KHz sampling rate to obtain audio samples; Divide the audio samples into several frames; Use Hann window function to weight each frame of data; Short-time Fourier transform is used on each frame of weighted data to convert from time domain to frequency domain, and Mel frequency conversion is performed to obtain the logarithmic Mel spectrum; Repeat the above steps to obtain K groups of audio data samples.
7. The audio-visual fusion method according to claim 1, characterized in that: The tangent image is processed to obtain the spherical geometric perception embedding feature x Tan ∈R F×T×D , specifically: The global features are obtained by passing the tangent image sequence through the ResNet-18 encoder, which is downsampled and flattened to obtain a tangent feature vector with dimension D = 512. At this time, the ResNet-18 encoder remains frozen during the training process. A fully connected layer is used to map the angular coordinates (φ, θ) of each pixel of the tangent viewport to the same feature dimension and add it to the spherical position encoding to obtain a spherical geometry-aware embedding feature.
8. The audio-visual fusion method according to claim 1, characterized in that: The spherical geometric perception embedding feature is used to obtain the spatiotemporal features of the section, and further obtain the visual modality features. Specifically: The spherical geometric perception embedding features are calculated in the spherical ViT respectively using the temporal and spatial attention mechanisms, and feature enhancement is performed to obtain the spatiotemporal features of the cross-section; The slice spatiotemporal features are concatenated with the global features obtained by the ResNet-18 encoder to obtain the visual modality features.
9. The audio-visual fusion method according to claim 1, characterized in that: The audio data also includes the following processing: Differentiate the audio channels of the audio data. If there are four channels, treat W as one channel, X, Y, and Z as another channel after splicing, and finally merge them into one channel for Mel frequency conversion. If there is only one channel, the Mel frequency conversion is performed directly.
10. The audio-visual fusion method according to claim 1, characterized in that: The audio semantic time attention mechanism includes: In the global audio information feature The time dimension increases the position encoding to learn the transformation process of time series; The audio features are encoded through two consecutive Transformer layers to obtain the enhanced global temporal semantic information f a .
Citation Information
Patent Citations
Panoramic video fixation point transfer detection and enhancement method based on global information
CN117876928A
No-reference panoramic video quality evaluation method and system based on spatial-temporal feature fusion
CN118865072A
Panoramic video processing method and device, electronic equipment and storage medium
CN119893158A
Audio-visual speech enhancement
US20210134312A1
Cited By
Method, medium and device for predicting saliency of multi-view video with spatial audio
CN122454489A
Method, medium and device for predicting saliency of multi-view video with spatial audio
CN122454489B