An audiovisual segmentation method integrating frequency-domain design

By introducing frequency-driven strategy and frequency-domain-oriented audio integration module in the audio-visual segmentation method, as well as the cross-modal fusion module based on the frequency domain, the problem of unused audio-visual features in the existing methods is solved, and the accuracy and consistency of audio-visual segmentation are significantly improved.

CN119693857BActive Publication Date: 2025-05-27DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510192687.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing audio-visual segmentation methods mainly engage in cross-modal feature interaction in the spatial domain, fail to fully utilize the full potential of audio-visual information, lack in-depth correlation analysis of the instance level of vocal object, and ignore the intrinsic connection of audio-visual features in the frequency domain, resulting in insufficient integration and alignment of multimodal features.

Method used

By introducing a frequency-driven strategy and leveraging the synergy between spatial and frequency information, an audio-visual segmentation method with integrated frequency domain design is proposed, including a frequency domain-oriented audio integration module and a cross-modal fusion module based on the frequency domain, which significantly improves the integration and alignment effect of multimodal features.

Benefits of technology

This method effectively improves the consistency of audio-visual correspondence, and achieves more accurate boundaries of vocal objects and higher accuracy of audio-visual segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693857B_ABST
    Figure CN119693857B_ABST
Patent Text Reader

Abstract

The present invention discloses an audiovisual segmentation method integrating frequency-domain design. First, a method combining the spatial domain and the frequency domain is introduced to explore comprehensive audiovisual alignment and fusion. Integrating frequency-domain information improves audiovisual consistency and enhances the fine alignment of multimodal features. A frequency-domain guided audio integration module and a frequency-domain based cross-modal fusion module are proposed. Among them, the frequency-domain guided audio integration module encodes audio information into visual representations through early fusion based on frequency-domain enhancement, thereby generating more detailed visual representations of audio perception, which helps to generate more robust multimodal representations subsequently; the frequency-domain based cross-modal fusion module aims to explore audiovisual associations by combining the spatial domain and the frequency domain, enhance cross-modal feature alignment, and thus improve the segmentation performance of the model. The frequency-domain guided audio integration module is integrated into each stage of the encoder, making full use of frequency-domain information and audio cues, reducing the differences between audiovisual modalities, and contributing to more accurate audiovisual fusion and alignment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-modal (audio-visual modality) information processing, and relates to an audio-visual segmentation method integrating frequency domain design, specifically an audio-visual segmentation method based on frequency domain driving, multi-modal fusion and attention mechanism. Background Art

[0002] With the development of multimedia information processing technology, the Audio-Visual Segmentation (AVS) technology has emerged, aiming to accurately separate the visual objects that produce sounds from video and audio data. This technology has broad application prospects in many fields such as video editing, security monitoring, and human-computer interaction. The existing audio-visual segmentation methods mainly focus on using convolutional neural networks (CNNs) or attention mechanisms in the spatial domain for cross-modal feature interaction to achieve audio-visual segmentation. For example, Zhou et al. cleverly used a cross-modal attention module in Audio–visual segmentation to achieve early fusion, and combined it with a fully convolutional network (FCN) decoder to achieve audio-visual segmentation. Although this method is simple and effective, it still has certain limitations, including the failure to fully utilize the full potential of audio-visual information, being limited to exploring pixel-level correlations, and lacking in-depth correlation analysis at the instance level of the sound-producing objects. Gao et al. in Audio-Visual Segmentation with Transformer achieved a more robust audio-visual cross-modal representation by introducing a Transformer architecture based on the attention mechanism to improve the audio-visual segmentation performance. In addition, some works introduced contrastive learning to enhance the effectiveness of cross-modal attention, thereby improving the accuracy of audio-visual segmentation. However, the above methods are all limited to spatial domain audio-visual segmentation, lacking exploration of frequency domain information and ignoring the internal connection of audio-visual features in the frequency domain, resulting in insufficient integration and alignment of multi-modal features.

[0003] Therefore, in view of the deficiencies and challenges in the above research, the present invention innovatively introduces a frequency-driven strategy, utilizes the synergistic effect of spatial and frequency information, significantly improves the integration and alignment effect of multi-modal features, and effectively improves the consistency of audio-visual features and audio-visual segmentation performance. Summary of the Invention

[0004] The present invention proposes a method for strengthening the integration and alignment strategy of multi-modal data by using frequency domain information, and then realizing efficient audio-visual segmentation. This method can effectively improve the audio-visual correspondence consistency, achieve a more accurate segmentation boundary of the sound-producing object and a higher audio-visual segmentation accuracy.

[0005] Technical Solution of the Present Invention:

[0006] An audio-visual segmentation method integrating frequency-domain design, the steps are as follows:

[0007] Step 1: Feature encoding optimized by a frequency-domain guided audio integration module

[0008] In the audio encoding stage, the VGGish model is used as the audio encoder to extract audio features; in the visual encoding stage, a frequency-domain guided audio integration module is designed, and a design combining frequency-domain enhancement and early fusion is introduced during image encoding to generate more effective visual representations;

[0009] ResNet-50 or PVT-v2 is used as the visual encoder. The visual encoder contains four stages for gradually extracting features, denoted as stage1, stage2, stage3, and tage4; ResNet-50 is based on a convolutional neural network, and gradually reduces the resolution of the feature map through residual connections and convolutional operations while increasing the number of channels; PVT-v2 adopts the Transformer architecture and extracts features through self-attention mechanisms and patch embedding methods, and effectively captures global information through spatial downsampling and optimized calculation methods in each stage; the frequency-domain guided audio integration module is inserted after the output of each stage to process the feature map generated by the corresponding stage; after being processed by the frequency-domain guided audio integration module, the obtained visually enhanced features are used as the output of the current stage and are passed to the next stage as input; the details of each frequency-domain guided audio integration module are as follows: for the audio features output by the audio encoder and the spatial-domain visual feature map from the current stage of the visual encoder , first use the fast Fourier transform to convert the spatial-domain visual features to the frequency domain:

[0010]

[0011] where represents the frequency-domain feature map, represents the fast Fourier transform, respectively represent the indices of the height and width of the visual feature map , respectively represent the indices of the height and width of the converted frequency-domain feature map ;

[0012] Then introduce a threshold to divide the frequency-domain feature map into high-frequency components and low-frequency components, specifically as follows:

[0013]

[0014]

[0015] Furthermore, by introducing learnable parameters and adapting and adjusting the decoupled frequency components to generate a frequency-domain enhanced visual representation, this process is expressed as follows:

[0016]

[0017] where represents the frequency-domain enhanced visual representation, represents the inverse fast Fourier transform, which is used to map the frequency-domain features back to the spatial domain;

[0018] After that, a cross-modal attention mechanism is introduced to incorporate the audio features generated by the audio encoder into the visual representation, and a residual connection is used to obtain the audio-aware visual representation , which serves as the final output of the current stage and the input of the next stage; this process is formulated as follows:

[0019]

[0020] where represents the cross-modal attention mechanism;

[0021] Finally, the visual encoder outputs four multi-scale encoded features { }, where , T represents the time of the input video frame, H and W represent the height and width of each frame of the input video respectively, represents the number of channels of each encoded feature;

[0022] After that, a pixel decoder based on the multi-scale deformable Transformer is used for multi-scale visual feature fusion; the specific details are as follows: the multi-scale encoded features extracted by the visual encoder are further fed into the pixel decoder. The visual features of the last three scales in the multi-scale encoded features { are unfolded, concatenated along the channel dimension, and then feature fusion is performed through the multi-scale deformable Transformer layer to obtain the enhanced multi-scale visual features { ;

[0023] Step 2: Cross-modal fusion based on the frequency domain

[0024] Design a cross-modal fusion module based on the frequency domain to fully explore the audiovisual correspondence relationship; the cross-modal fusion module based on the frequency domain consists of two branches: a spatial guidance branch and a frequency-domain perception branch;

[0025] (1) Spatial-guided Branch: It includes cross-modal attention operations in the spatial domain: using the cross-modal attention mechanism for audio-visual fusion in the spatial domain. For the audio features output by the audio encoder and the multi-scale visual features output by the pixel decoder { the visual features with the maximum resolution in , respectively, take the encoded audio features as the query, and the spatial-domain visual features as the key and value and input them into the cross-modal attention mechanism to obtain the fused multi-modal features . This process is described as follows:

[0026]

[0027] Among them, , , are learnable projection matrices that map the features to intermediate-layer features with a dimension of ; afterwards, the multi-modal features are used for subsequent optimization and adjustment of the visual features, through a residual connection weighted by a learnable parameter , combined with the encoded audio features to obtain the enhanced audio features :

[0028]

[0029] (2) Frequency-domain Perception Branch: The frequency-domain perception branch works in parallel with the spatial-guided branch; first, use the fast Fourier transform to convert the visual features with the maximum resolution in the spatial domain to the frequency domain to obtain the frequency-domain visual feature map, and then use global average pooling to reduce the dimension of the frequency-domain visual features while retaining the global information; considering that the main structural and texture information of the frequency-domain visual features is reflected in the magnitude spectrum, and the audio features are obtained by extracting the audio spectrogram and encoding it using the audio encoder, further extract the spectral features from the frequency-domain visual feature map to facilitate subsequent audio-visual interaction and alignment; this process is expressed as follows:

[0030]

[0031]

[0032]

[0033] Among them, , , , , respectively represent the frequency-domain visual feature map after Fourier transform, the frequency-domain features after global average pooling, the global average pooling operation, the real part of and the imaginary part of

[0034] Finally, the frequency-domain perception branch works in cooperation with the space-guided branch to optimize and adjust the visual representation; the spectral features and the fused multi-modal features first pass through a convolutional layer independently, and then the two are deeply fused through element-wise addition and a convolutional layer, enabling the spectral features and the fused multi-modal features to complement and enhance each other, promoting the effective interaction and fusion of multi-modal information in the frequency domain; this process is expressed as:

[0035]

[0036] where represents the convolution operation, represents the output result, which is used to weight and adjust the original visual features; is adjusted by multiplication , and also passes through a residual connection with learnable parameter weighting to obtain the enhanced visual feature ;

[0037]

[0038] where represents the weighted learnable parameter;

[0039] Through the cross-modal fusion module based on the frequency domain, the optimized audio feature and the visual feature are finally obtained;

[0040] Step 3: Generate a mask through the decoder

[0041] Adopt the Mask2Former architecture as the decoder to generate the vocal object query embedding; then the vocal object query embedding is combined with the obtained in Step 2 to generate the predicted mask;

[0042] Specifically, the optimized audio feature obtained from the cross-modal fusion module based on the frequency domain is combined with the learnable query initialized with the learnable parameter through point-wise addition as the query of the decoder; meanwhile, the enhanced multi-scale visual features { output by the pixel decoder for multi-scale feature fusion are used as the key and value in sequence, and interact with the audio-based query through the layer-by-layer cross-modal attention mechanism to generate the vocal object query embedding ; finally, Obtain classification predictions through a linear layer, Multiply with the optimized maximum-resolution visual features obtained from the cross-modal fusion module based on the frequency domain to obtain a prediction mask;

[0043] Step 4: Model training

[0044] During training, consider both the class loss and the prediction mask loss. To utilize the potential temporal coupling of audiovisual segmentation to improve the performance of the model, an adaptive inter-frame consistency loss is adopted; the three losses are combined to obtain the total loss and backpropagated to optimize the model:

[0045]

[0046]

[0047]

[0048]

[0049] Among them, represents the true class, represents the class predicted by the model, and the class loss function is calculated by the cross-entropy function; and represent the labels of a certain point on the true mask and the predicted mask respectively, and the mask loss function is composed of binary cross-entropy and Dice loss; represents the predicted mask of the t-th frame, i.e., the adaptive inter-frame consistency loss; , , represent the weighting parameters respectively, i.e., the total loss function.

[0050] Advantages of the present invention:

[0051] (1) The present invention first introduces a method of combining the spatial domain and the frequency domain to explore comprehensive audiovisual alignment and fusion, integrates frequency domain information to improve audiovisual consistency, and enhances the fine alignment of multi-modal features.

[0052] (2) Two innovative frequency-domain design modules are proposed, namely the frequency-domain oriented audio integration module and the frequency-domain based cross-modal fusion module. Among them, the frequency-domain oriented audio integration module encodes audio information into visual representations through early fusion based on frequency-domain enhancement, thereby generating more detailed visual representations of audio perception, which helps to generate more robust multi-modal representations subsequently; the frequency-domain based cross-modal fusion module aims to explore audiovisual associations by combining the spatial domain and the frequency domain, enhancing cross-modal feature alignment, and thus improving the segmentation performance of the model.

[0053] (3) The frequency-domain oriented audio integration module is integrated into each stage of the encoder, making full use of frequency-domain information and audio cues to reduce the differences between audiovisual modalities, which helps for more accurate audiovisual fusion and alignment. Different from previous works that only explore spatial-domain audiovisual interactions, the frequency-domain based cross-modal fusion module collaboratively explores spatial-domain and frequency-domain audiovisual interactions, promoting a more refined and accurate segmentation effect. Brief Description of the Drawings

[0054] Figure 1 It is the overall architecture diagram of the audiovisual segmentation method involving integrated frequency-domain design in this application.

[0055] Figure 2 It is the detailed diagram of the frequency-domain oriented audio integration module within the method involved in this application.

[0056] Figure 3 It is the detailed diagram of the frequency-domain based cross-modal fusion module within the method involved in this application. Detailed Embodiments

[0057] The following further illustrates the detailed embodiments of the present invention in conjunction with the drawings and technical solutions.

[0058] Refer to Figure 1 , an audiovisual segmentation method involving integrated frequency-domain design, includes the following steps:

[0059] Step 1: Data preprocessing, including the following steps: Dataset selection and cropping: Use the AVSBench dataset for the audio-visual segmentation task, which includes two datasets, AVSBench-object and AVSBench-sematic. Among them, AVSBench-object contains two sub-datasets: the single-source audio-visual segmentation dataset (S4) and the multi-source audio-visual segmentation dataset (MS3). S4 contains 4,932 videos. Using a semi-supervised training method, the division ratio of the training set, validation set, and test set is 70 / 15 / 15. MS3 includes 424 videos and adopts a fully supervised training method. The division of the training set, validation set, and test set is the same as that of S4. AVSBench-semantic adds semantic labels for the audio-visual semantic segmentation (AVSS) task. It contains 11,356 videos, of which 8,498 are used for training, 1,304 for validation, and 1,554 for testing. During the training and testing experiments, this design uniformly crops the video frames to a size of 224×224. In addition, two data augmentation methods are used for all input video frames, including horizontal flipping and color enhancement.

[0060] Step 2: Audio-visual encoding, including the following steps: First, send the audio mel spectrogram obtained from the audio into the pre-trained audio encoder VGGish to extract audio features . To ensure the stability of the representation ability of the audio encoder, its weights are frozen during training. After that, the pre-trained visual encoder ResNet-50 or PVT-v2 model is used to jointly encode the video frames obtained in the data processing step with the frequency-domain guided audio integration module. The specific details of this step are as follows: Keeping the overall structure of the visual encoder unchanged, the frequency-domain guided audio integration module is inserted after the four main stages of the visual encoder. During the process of extracting image features, frequency-domain enhancement is used and audio cues are gradually integrated to generate audio-aware visual representations. Among them, the specific details of the frequency-domain guided audio integration module are as follows: The spatial-domain visual feature map is converted into a frequency-domain feature map using the Fourier transform, and then the high and low frequency components of the frequency-domain feature map are divided using a threshold of 0.1. Learnable parameters 1.0 and 1.5 are set for the high and low frequency components respectively, and a multiplication operation is performed. The adjusted high and low frequency components are added together to obtain an enhanced frequency-domain feature map. Then the inverse Fourier transform is used to convert the frequency-domain feature map back to the spatial domain. After that, the frequency-domain enhanced visual representation is fused with audio cues through the multi-head cross-modal attention mechanism, and the original visual information is retained through the residual connection to obtain the audio-aware visual representation. Finally, the visual encoder of this design outputs multi-scale visual feature maps obtained from each frequency-domain guided audio integration module. Subsequently, a pixel decoder based on the multi-scale deformable Transformer is used to perform multi-scale feature fusion on the multi-scale visual feature maps. The details of the pixel decoder are as follows: Six Transformer layers are used to perform feature fusion on the three multi-scale visual feature maps with i = 2 to 4 output by the visual encoder, enhancing the multi-scale visual feature maps, and at the same time mapping all multi-scale features to the same channel dimension through convolutional layers.

[0061] Step 3: Audio-visual two-way interaction optimization, the specific steps are as follows: The highest-resolution visual features output by the pixel decoder and the encoded audio features are fed into the frequency-domain based cross-modal fusion module, which comprehensively utilizes spatial-domain and frequency-domain information to interact and two-way optimize the visual and audio features, obtaining improved audio features and visual features. The specific details are as follows: In the spatially guided branch, this design uses the multi-head cross-modal attention mechanism for spatial-domain audio-visual fusion, with 8 attention heads. Among them, the audio is used as the query for the attention operation, and the visual features are used as the key and value to obtain the fused multi-modal features . This fused multi-modal feature is further fused with the original audio feature through a residual connection with a weight of 1.0 to obtain enhanced audio features. In the frequency-domain perception branch, first, the spatial-domain visual features are converted to the frequency domain using the fast Fourier transform, the global average pooling is used to reduce the dimension of the visual features, and then the spectral features of the frequency-domain visual feature map are extracted . Further, through convolution and pointwise addition, the output of the spatially guided branch is fused and the result obtained by the frequency-domain perception branch is used, and its result is used as the weight to adjust the visual features. Finally, enhanced visual features are obtained through a weighted residual connection as well. The learnable parameters in this process are set to 1.0, and its formula is as follows:

[0062]

[0063] Step 4: Mask generation, including the following steps: Adopt a multi-level cascaded decoder architecture based on Transformer. Use a set of learnable query embeddings combined with the optimized audio features in Step 3 through pointwise addition as the initial queries, and each query corresponds to a potential vocal target object. The decoder processes the multi-scale visual features layer by layer. Each decoding layer includes a self-attention layer for the queries, a cross-attention layer for the interaction between the queries and the visual feature maps as keys and values, and a feed-forward network composed of a normalization layer and a linear layer. Finally, the decoder outputs a vocal object query embedding .

[0064] The final query embedding is respectively input into two branches: the classification head and the mask head. The classification head predicts the vocal object category. The specific details are as follows: The query embedding is mapped to an embedding space with the same number of categories through a linear layer to obtain the category prediction. Its formula is as follows:

[0065]

[0066] The mask head generates a binary segmentation mask through pixel-level interaction with the high-resolution feature map. The specific details are as follows: First, the vocal object query embedding is processed by a multi-layer perceptron, and then it is multiplied by the enhanced visual features output in Step 3 to obtain the predicted mask. Its formula is as follows:

[0067]

[0068] where () represents the processing of the multi-layer perceptron, represents matrix multiplication.

[0069] Step 5: Loss function and optimization, the specific steps are as follows: During the training process, consider the category prediction loss, the mask prediction loss, and the adaptive inter-frame consistency loss. The total loss is obtained by combining the three losses and the model is optimized through backpropagation. The formula of the category prediction loss function:

[0070]

[0071] where, represents the true category, Indicates the predicted category of the model. The formula for the masked prediction loss function:

[0072]

[0073] Among them, and respectively represent the labels of a certain point on the true mask and the predicted mask. The formula for the adaptive inter-frame consistency loss function: Among them, represents the predicted mask of the t-th frame. The formula for the total loss function can be expressed as:

[0074]

[0075] Among them, , , respectively represent the weighting parameters, which are specifically set to 2.0, 5.0, and 10.0. Finally, the total loss is calculated through the total loss function and the model is optimized by backpropagation.

[0076] Experimental verification:

[0077] I. Verification details

[0078] The method is based on the Mask2Former model architecture: Two visual encoder settings are used in the experiment: ResNet-50 pre-trained on ImageNet and PVT-v2, and the audio encoder uses the pre-trained VGGish model. The pixel decoder adopts a multi-scale deformable attention transformer architecture. In the decoder, the number of learnable queries is set to 100. This method also introduces diverse data augmentation techniques, including flipping and random cropping, etc. During the training process, the AdamW optimizer that combines momentum and adaptive learning rate is selected to improve the training stability.

[0079] During the training process, the weight coefficients of the loss function are respectively set to λcls = 2, λmas = 5, λada = 10, the learning rate is set to 0.0001, the weight decay is 0.05, the cropping size is 224×224, and the batch size is 8. The single-source subset in AVSBench-Object and AVSBench-Semantic are trained for 90000 iterations, while the single-source subset in AVSBench-Object is trained for 20000 iterations.

[0080] To ensure fairness, following previous research, comparisons are made between models using the same visual encoder, and the same evaluation metrics are used: Jaccard index and F-score .

[0081] II. Quantitative Results

[0082] Table 1 shows the comparison results between the advanced audio-visual segmentation method and the proposed design. The performance of the model designed in this application has achieved excellent results. Especially on the multi-source subset and audio-visual semantic segmentation datasets, it has surpassed the previous state-of-the-art (SOTA) methods. Whether the image encoder used is ResNet-50 or PVT-v2 network, the audio-visual segmentation method with integrated frequency-domain design in this design has demonstrated excellent performance and exceeded the latest SOTA methods in both evaluation metrics.

[0083] Table 1 shows the comparison results of using two visual encoders, ResNet-50 and PVT-v2 respectively, with the current advanced audio-visual segmentation methods on the AVSBench dataset.

[0084]

Claims

1. An audio-visual segmentation method integrating frequency domain design, characterized in that: Here are the steps: Step 1: Feature encoding via frequency-domain oriented audio integration module optimization In the audio encoding stage, the VGGish model is used as an audio encoder to extract audio features; in the visual encoding stage, a frequency domain-oriented audio integration module is designed, and a design combining frequency domain enhancement and early fusion is introduced in the image encoding process to generate more effective visual representations; the details are as follows: ResNet-50 or PVT-v2 is used as the visual encoder. The visual encoder contains four stages for gradually extracting features, which are represented as stage1, stage2, stage3 and tag4. ResNet-50 is based on a convolutional neural network. It gradually reduces the resolution of the feature map through residual connections and convolution operations, while increasing the number of channels. PVT-v2 uses the Transformer architecture to extract features through a self-attention mechanism and patch embedding. In each stage, global information is effectively captured through spatial downsampling and optimized calculation methods. The frequency-domain-oriented audio integration module is inserted after the output of each stage to process the feature map generated by the corresponding stage. After being processed by the frequency-domain-oriented audio integration module, the visual enhancement features obtained are used as the output of the current stage and are passed to the next stage as input. The details of each frequency-domain-oriented audio integration module are as follows: For the audio features output by the audio encoder and the spatial domain visual feature map from the current stage of the visual encoder First, the spatial domain visual features are transformed using fast Fourier transform Convert to the frequency domain: in, represents the frequency domain feature map, represents the fast Fourier transform, Represents the visual feature map The height and width index of Respectively represent the frequency domain feature maps after conversion The height and width index of the Introducing a threshold The frequency domain feature map Divided into high-frequency components and low-frequency components, as shown in the following formula: Furthermore, by introducing learnable parameters and Adaptively adjust the decoupled frequency components to generate a visual representation enhanced in the frequency domain. The process is expressed as follows: in, represents the visual representation after frequency domain enhancement, represents the inverse fast Fourier transform, which is used to map frequency domain features back to the spatial domain; Afterwards, a cross-modal attention mechanism is introduced to convert the audio features generated by the audio encoder to visual representation, and use residual connections to obtain audio-aware visual representation , as the final output of the current stage and the input of the next stage; the process is expressed in the following formula: in, Represents the cross-modal attention mechanism; Finally, the visual encoder outputs four multi-scale encoding features { },in , T represents the time of the input video frame, H and W represent the height and width of each frame of the input video respectively, Indicates the number of channels for each encoded feature; Afterwards, a pixel decoder based on a multi-scale deformable Transformer is used to perform multi-scale visual feature fusion; the details are as follows: Multi-scale encoded features extracted by the visual encoder is further fed into the pixel decoder, the multi-scale encoding feature { The visual features of the last three scales in are expanded, connected along the channel dimension, and then fused through the multi-scale deformable Transformer layer to obtain enhanced multi-scale visual features { ; Step 2: Cross-modal fusion based on frequency domain A frequency-domain-based cross-modal fusion module is designed to fully explore the audio-visual correspondence relationship. The frequency-domain-based cross-modal fusion module consists of two branches: a spatial guidance branch and a frequency-domain perception branch. The details are as follows: (1) Spatial-guided branch: including spatial domain cross-modal attention operation: using cross-modal attention mechanism to perform audio-visual fusion in spatial domain, for the audio features output by the audio encoder and the multi-scale visual features output by the pixel decoder { Maximum resolution visual features in , respectively, the encoded audio features As query, spatial domain visual features Input as key and value into the cross-modal attention mechanism to obtain fused multimodal features , the process is described as follows: in, , , It maps the features to a dimension of The learnable projection matrix of the intermediate layer features; then, the multimodal features For subsequent optimization and adjustment of visual features, Through a Weighted residual connections, and encoded audio features Combined to achieve enhanced audio characteristics : (2) Frequency-domain-aware branch: The frequency-domain-aware branch works in parallel with the spatial-guided branch. First, the maximum resolution visual features in the spatial domain are transformed using the fast Fourier transform. Convert to the frequency domain to obtain the frequency domain visual feature map, and then use global average pooling to reduce the dimension of the frequency domain visual feature while retaining the global information; considering that the main structure and texture information of the frequency domain visual feature are reflected in the amplitude spectrum, and the audio feature is obtained by extracting the audio spectrum map and encoding it using the audio encoder, the frequency domain visual feature map is further extracted with spectrum features to facilitate subsequent audio-visual interaction and alignment; the process is expressed as follows: in, , , , , They represent the frequency domain visual feature map after Fourier transform, the frequency domain features after global average pooling, the global average pooling operation, The real part and The imaginary part of Finally, the frequency-domain-aware branch works in tandem with the spatially guided branch to optimize and adjust the visual representation; spectral features And fused multimodal features First, each is processed independently by a convolution layer, and then the two are deeply fused through element-by-element addition and convolution layer, so that the spectral features Multimodal features with fusion Complement and enhance each other, and promote the effective interaction and fusion of multimodal information in the frequency domain; the process can be expressed as: in, represents the convolution operation, Represents the output result, which is used to weight and adjust the original visual features; Adjustment by multiplication , also through residual connections with learnable parameter weights, to obtain enhanced visual features ; in, represents the weighted learnable parameters; Through the frequency domain-based cross-modal fusion module, the optimized audio features are finally obtained and visual features ; Step 3: Implement mask generation through decoder The Mask2Former architecture is used as the decoder to generate the sound object query embedding. The sound object query embedding is then combined with the one obtained in step 2. Combine to generate a prediction mask; Specifically, the optimized audio features obtained by the frequency domain cross-modal fusion module and the learnable query obtained by initializing the learnable parameters are combined through point-by-point addition as the query of the decoder; at the same time, the pixel decoder performs multi-scale feature fusion to output the enhanced multi-scale visual features { As keys and values, they interact with audio-based queries through a layer-by-layer cross-modal attention mechanism to generate query embeddings for sounding objects. ;at last, The classification prediction is obtained through the linear layer, The optimized maximum resolution visual features obtained by the frequency domain-based cross-modal fusion module Multiply them together to get the prediction mask; Step 4: Model training In the training process, both the category loss and the predicted mask loss are considered. In order to utilize the potential temporal coupling of audio-visual segmentation to improve the performance of the model, an adaptive inter-frame consistency loss is used. The three losses are combined to get the total loss and backpropagate to optimize the model: in, represents the true category, Represents the model prediction category, category loss function Calculated by the cross entropy function; and Represents the label of a point on the real mask and the predicted mask, respectively, and the mask loss function It consists of binary cross entropy and Dice loss; represents the prediction mask of the t-th frame, That is, adaptive inter-frame consistency loss; , , They represent weighting parameters, That is the total loss function.

Citation Information

Patent Citations

  • Audio and video segmentation method and system based on multi-modal fusion attention

    CN117951335A

  • Audio data processing

    US20240212706A1