An audio-visual segmentation method based on cross-modal cognitive consensus alignment
By introducing a cross-modal cognitive consensus inference module and a cognitive consensus-guided attention module in the audio-visual segmentation method, semantic level alignment of audio and video is achieved, and the problem of dimensional differences between audio and video in the prior art is solved, and the segmentation accuracy and performance are significantly improved.
Patent Information
- Application Number
- CN202310933937.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-07-27
AI Technical Summary
The existing audio and video segmentation method cannot effectively solve the dimensional difference between audio global information and multiple visual local information through feature-level interaction, resulting in poor segmentation effect.
A cross-modal cognitive consensus inference module and a cognitive consensus-guided attention module are proposed. The modal alignment semantic labels are calculated through the classification head of the audio and video encoder, and the semantic level alignment information is injected into the segmentation framework using gradient inversion technology to achieve the combination of feature level and semantic level.
Through semantic-level cross-modal alignment, the dimensional difference between audio and video is effectively compensated, and the accuracy and segmentation performance of audio and video segmentation are significantly improved, achieving the most advanced segmentation effect at present.
Smart Images

Figure CN117079181B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-modal image segmentation. Given a video and corresponding audio, with the audio signal as a reference, the target in the video that emits the sound is extracted and a pixel-level mask is generated. Through the proposed cross-modal cognitive consensus inference module and cognitive consensus-guided attention module, the present invention performs explicit semantic-level cross-modal alignment on the audio and video, and obtains good target segmentation results. Background Art
[0002] With the continuous development of the field of computer vision, fine-grained segmentation technologies for visual image targets such as semantic segmentation, instance segmentation, and panoptic segmentation have achieved remarkable achievements. The above methods equally treat every target and background in the image and segment them. However, in real multimedia application scenarios, it is often only necessary to highlight the truly interesting targets, which cannot be achieved by the above image segmentation methods. The purpose of audio-visual segmentation is to refine and extract the interesting targets (sound-emitting targets) in the image under the guidance of audio information. This segmentation method has extensive potential uses and important significance in real application scenarios.
[0003] The main challenges of audio-visual segmentation lie in the following two aspects: on the one hand, the model needs to fully understand the semantic content and long-distance context information of both visual and audio modalities; on the other hand, the model needs to perform explicit and accurate alignment of the visual and audio modalities. Specifically, a piece of audio information usually only contains global audio label information, but each frame of the video often contains different local targets. Achieving alignment from global to local and highlighting the interesting targets is the key difficulty of this task.
[0004] In terms of video and audio data encoders, many excellent models have been proposed. In the visual modality, researchers usually use visual encoders based on convolutional neural networks (CNNs), such as ResNet, VGGNet, etc., or use higher-performance visual encoders based on Transformers, such as ViT, Swin Transformer, PVT, etc.; in the audio modality, the current mainstream method is to convert the audio into a spectrogram and use an encoder with a convolutional network structure to extract features. Widely used audio encoders include VGGish, PANNs, etc. The above high-performance visual and audio encoders provide a solid foundation and stable guarantee for this method.
[0005] This method is further improved based on the paper "Audio-Visual Segmentation" in ECCV 2022. In the method proposed in this paper, the author uses a cross-modal attention module to perform dense cross-modal interactions on the extracted audio and video features, and inputs the multi-modal features after the interaction into the segmentation head to achieve audio-visual segmentation. However, the above method only performs feature-level interactions and alignments on the audio-visual modalities, and single feature-level alignment cannot effectively solve the above-mentioned dimensional gap problem from global to local. Therefore, to solve the above problem, a method based on semantic-level cross-modal cognitive consensus is proposed, which further performs semantic-level interactions on the basis of audio-visual feature-level interactions, effectively compensates for the dimensional gap and achieves more accurate segmentation.
[0006] This solution has not been publicly published in domestic or foreign publications, has not been publicly used at home and abroad, or made known to the public in other ways. Summary of the Invention
[0007] The object of the present invention is to solve the following technical problems:
[0008] First: The existing audio-visual segmentation methods only use single feature-level interaction to achieve cross-modal alignment, and cannot solve the dimensional difference problem between the global audio information and multiple local visual information; to solve this problem, the present invention proposes a cross-modal cognitive consensus inference module to achieve semantic-level alignment between modalities; specifically, the present invention classifies the audio-visual modalities respectively through the classification heads of the audio-visual encoders, and obtains the classification confidence of each audio-visual modality, and weights and scores the semantic similarity between the above confidences and the audio-visual classification labels to obtain the semantic label for modality alignment.
[0009] Second: After obtaining the semantic label for modality alignment, the present invention uses the gradient backpropagation technique to send the modality-aligned label back to the visual encoder and obtain the weight vector corresponding to this semantic category; the present invention proposes a cognitive consensus-guided attention module to inject the semantic-level alignment information into the audio-visual segmentation framework, so as to achieve the combination of feature-level alignment and semantic-level alignment of the audio-visual modalities; subsequently, the present invention inputs the audio-visual aligned features into a general fully convolutional segmentation network to achieve the segmentation of the sound-emitting target; the present invention combines the inference of cross-modal cognitive consensus with feature-level alignment, achieving the most advanced segmentation performance. The existing audio-visual segmentation methods use cross-modal attention modules to achieve dense audio-visual feature-level interactions, but due to the lack of higher-level semantic-level cross-modal alignment, it is difficult for the existing methods to solve the dimensional difference problem between the global audio label and multiple local regions of the video.
[0010] The technical solution of the present invention is: An audio-visual segmentation method based on cross-modal cognitive consensus alignment, the method comprising:
[0011] Step 1: Obtain a video frame and its corresponding audio segment; the visual encoder has four feature extraction stages. Input the video frame into the visual encoder, and take the visual features output by the four stages of the visual encoder as hierarchical visual features, denoted as V i , where i = 1, 2, 3, 4; in addition, input the audio segment into the audio encoder to extract the audio feature F a ; the hierarchical visual features V i and the audio feature F a will be used for subsequent calculations;
[0012] Step 2: Utilize the classification heads and their classification weights preset in the audio encoder and the visual encoder; among the hierarchical visual features V i , where i = 1, 2, 3, 4, V 4 is the highest-level visual feature and contains the global semantic information of the image; respectively perform class confidence scoring on the visual feature V 4 and the audio feature F a to obtain the visual classification confidence and the audio classification confidence Then, calculate the semantic-level similarity m between the visual label text and the audio label text, and the specific formula is as follows: jk :
[0013]
[0014] where ||·|| F represents the Frobenius norm, and j and k respectively represent the row and column indices of the finally calculated semantic similarity matrix M sim ; then, calculate the confidence reweighting matrix M cof (j, k), and the specific formula is as follows:
[0015]
[0016] where α and β are balance coefficients, and the values in the confidence reweighting matrix M cof (j, k) can be regarded as the cognitive consensus scores of the corresponding visual semantics and text semantics; after obtaining the confidence reweighting matrix M cof (j, k), find the maximum scoring value in the matrix and obtain the visual label corresponding to the maximum value as the semantic label for modality alignment; transmit the semantic label for modality alignment back to the four hierarchical stages of the visual encoder in the form of gradient backpropagation and obtain the class activation weights
[0017] Step 3: Obtain the class activation weights After that, the weights containing semantic-level alignment information are integrated into the features extracted by the encoder, and the specific formula is as follows:
[0018]
[0019]
[0020]
[0021]
[0022] Among them, σ represents the sigmoid function operation, Avg refers to the average operation, represents the element-wise multiplication with the broadcasting mechanism; the cognitive consensus-guided attention module integrates the semantic-level cognitive consensus weights and the video feature-level representations in a channel-space form to obtain the integrated information to guide the network for subsequent segmentation;
[0023] Step 4: First, perform mapping and repetition operations on the audio feature F a to obtain Then, pass the visual feature V at this level i through the dilated convolution module to obtain V i a and input it together with the audio feature into the non-local module for audio-visual feature-level cross-modal interaction. The specific formula is as follows:
[0024]
[0025] M i = V i a + θ 4 (Φ · θ 3 (V i a )) (8)
[0026] In the formula, θ 1 , θ 2 , θ 3 and θ 4 respectively represent different three-dimensional convolutional layers, N is the number of pixels in the feature spectrum, Φ is the cross-modal attention matrix, and M i is the multi-modal feature at the i-th level;
[0027] Step 5: Fuse the hierarchical multi-modal features M i The fusion formula is as follows:
[0028]
[0029] In the formula, Conv represents the convolutional layer, and Upsample represents the upsampling operation. Y 1 is fed into the fully convolutional network to obtain the predicted value of the network Finally, the network is trained using binary cross-entropy loss:
[0030]
[0031] where Y represents the ground truth of the segmentation mask, and L seg represents the loss value.
[0032] In the present invention, a method for audio-visual cross-modal semantic-level cognitive consensus is proposed, and a new cross-modal cognitive consensus module and a cognitive consensus-guided attention module are proposed. The cross-modal cognitive consensus module calculates the classification confidence of audio and vision respectively, and measures the mutual similarity of audio-visual semantic labels. Then, the classification confidence is used to weight the mutual similarity to obtain the semantic-level cross-modal cognitive consensus score and select the semantically aligned labels. Subsequently, the gradient of the semantically aligned labels is backpropagated to the visual encoder to obtain the class activation information. Through the cognitive consensus-guided attention module, the visual targets with high semantic consistency are highlighted to guide the subsequent segmentation process. On the one hand, the method of the present invention achieves the current state-of-the-art performance on the audio-visual segmentation dataset; on the other hand, the method of the present invention can accurately and effectively segment the vocal targets in the video and output the pixel-level mask. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a cross-modal cognitive consensus alignment idea diagram;
[0034] Figure 2 is an audio-visual segmentation method based on cross-modal cognitive consensus alignment;
[0035] Figure 3 is a schematic diagram of the subjective comparison of the segmentation effects between the method of this study and the baseline method. EMBODIMENTS
[0036] Step 1: The video frames and their corresponding audio segments are respectively input into the visual encoder and the audio encoder to extract the corresponding hierarchical visual features V i , i = 1, 2, 3, 4 and the audio feature F a ; The hierarchical visual features V i and the audio feature F a are used for subsequent calculations;
[0037] Step 2: In the cross-modal cognitive consensus inference module shown in the lower left corner as Figure 2 , using the classification heads and their classification weights preset in the audio encoder and the visual encoder, respectively, for the visual features V4 Perform class confidence scoring on the audio feature F a and calculate the obtained visual classification confidence and the audio classification confidence Next, calculate the semantic-level similarity between the visual label text and the audio label text The specific formula is as follows:
[0038]
[0039] In the formula, ||·|| F represents the Frobenius norm, and j and k respectively represent the row and column indices of the finally calculated semantic similarity matrix M sim ; Next, calculate the confidence reweighting matrix M cof , and the specific formula is as follows:
[0040]
[0041] In the formula, α and β are balance coefficients, which are respectively set to 0.1; the values in the confidence reweighting matrix M cof can be regarded as the cognitive consensus scores corresponding to the visual semantics and the text semantics. Note: 1000 and 527 are respectively the preset number of visual categories and the number of audio categories of the encoder; after obtaining the confidence reweighting matrix M cof , find the maximum score value in the matrix, and obtain the corresponding visual label at the maximum value as the semantic label for modality alignment; transmit the semantic label for modality alignment back to the four hierarchical stages of the visual encoder in the form of gradient backpropagation, and obtain the class activation weights (i = 1, 2, 3, 4);
[0042] Step 3: Obtain the class activation weights (i = 1, 2, 3, 4), then use the cognitive consensus-guided attention module CCAM shown in the lower right corner Figure 2 to integrate the weights containing semantic-level alignment information into the features extracted by the encoder. The specific formula is as follows:
[0043]
[0044]
[0045]
[0046]
[0047] In the above formula, σ represents the sigmoid function operation, and Avg refers to the averaging operation Represents pointwise multiplication with a broadcast mechanism. The attention module guided by cognitive consensus integrates semantic-level cognitive consensus weights and video feature-level representations in a channel-space form to guide the network for subsequent segmentation.
[0048] Step 4: After integrating the semantic-level alignment information in Step 4, the present invention then performs feature-level fusion. The present invention first maps and repeats the audio feature F a to obtain Next, the visual feature V at this level i is passed through the dilated convolution module to obtain V i a and is jointly input into the non-local module for audio-visual feature-level cross-modal interaction. The specific formula is as follows:
[0049]
[0050] M i = V i a + θ 4 (Φ · θ 3 (V i a )) (8)
[0051] In the formula, θ 1 , θ 2 , θ 3 and θ 4 represent different three-dimensional convolutional layers respectively, N is the number of pixels in the feature spectrum, Φ is the cross-modal attention matrix, and M i is the multi-modal feature at the i-th level.
[0052] Step 5: The hierarchical multi-modal features M i (i = 1, 2, 3, 4) are fused. The fusion formula is as follows:
[0053]
[0054] In the formula, Conv represents the convolutional layer and Upsample represents the upsampling operation. Y 1 is fed into the fully convolutional network to obtain the predicted value of the network Finally, the network is trained using binary cross-entropy loss (BCELoss):
[0055]
[0056] In the above formula, Y represents the true value of the segmentation mask, and L seg represents the loss value.
[0057] The idea of this method is asFigure 1 As shown: Cross-modal cognitive consensus is the semantic-level consensus across audio and video modalities. In Figure 1 , for each input video frame image and its corresponding audio segment, the invention extracts the corresponding visual semantics and audio semantics respectively; in an audio segment, in most cases, there is only one global audio semantic label, but in a video image, there are often multiple local regions, and these local regions correspond to different visual semantic labels. Then, according to the similarity of semantic labels, the most relevant audio and visual semantic labels are selected as the semantic alignment labels for modal alignment; finally, the inferred semantic alignment labels for modal alignment are used as guiding information to obtain an accurate vocal target segmentation mask.
[0058] In Figure 2 , the invention shows the specific network architecture of the proposed method. In the figure, V i (i = 1, 2, 3, 4) represents the hierarchical visual features extracted by the visual encoder; sigmoid refers to the sigmoid function operation; (i = 1, 2, 3, 4) represents the four stages of modal alignment backpropagating to the visual encoder, and the class activation weights obtained for different hierarchical visual features respectively; CCAM refers to the cognitive consensus-guided attention module; M i (i = 1, 2, 3, 4) refers to the hierarchical multimodal features obtained through feature-level and semantic-level cross-modal alignment.
[0059] In terms of the visual encoder, the invention selects two high-performance visual encoders: Swin Transformer (abbreviated as Swin) and Pyramid Vision Transformer v2 (abbreviated as PVTv2); in terms of the audio encoder, the invention selects VGGish and PANNs. The invention conducts experimental evaluations on the current mainstream audio-visual segmentation dataset AVSBench and compares with other mainstream methods. AVSBench is divided into a single-source subset (abbreviated as S4) and a multi-source subset (abbreviated as MS3) to comprehensively evaluate audio-visual segmentation methods. In the quantitative experiment, the invention uses the mean intersection over union (mIoU) and F-score to evaluate the method of the invention. As shown in Table 1, in the S4 setting, the method of the invention increases by 1.7% mIoU and 2.0% F-score compared with the current state-of-the-art method; in the MS3 setting, the method of the invention increases by 3.0% mIoU and 3.5% F-score compared with the current state-of-the-art method.
[0060] Table 1 Objective performance evaluation table of the method in this study on the AVSBench dataset
[0061]
[0062] In Figure 3 , the present invention shows the comparison of the subjective segmentation effects of the method of the present invention and the baseline method under different combinations of audio - video encoders. It can be seen from the comparison results of Figure 3 that the method of the present invention based on cross - modal cognitive consensus alignment can more accurately locate and segment the target making sounds compared with the compared baseline method.
Claims
1. An audio - video segmentation method based on cross - modal cognitive consensus alignment, the method comprises: Step 1: Obtain video frames and their corresponding audio segments; The visual encoder has four feature extraction stages. The video frames are input into the visual encoder, and the visual features output by the four stages of the visual encoder are taken as hierarchical visual features, denoted as V i , where i = 1, 2, 3, 4; in addition, the audio clip is input into the audio encoder to extract the audio feature F a ; the hierarchical visual feature V i and the audio feature F a will be used for subsequent calculations; Step 2: Use the classification heads and classification weights preset by the audio encoder and visual encoder; the hierarchical visual features V output by the visual encoder i , i=1,2,3,4, V 4 is the highest level visual feature and contains the global semantic information of the image; 4 With audio feature F a Perform category confidence scoring and calculate the visual classification confidence and audio classification confidence Next, calculate the visual label text Text with audio tag The semantic similarity m between jk , the specific formula is as follows: Among them, ||·|| F represents the Frobenius norm, where j and k respectively represent the row and column indices of the finally calculated semantic similarity matrix M sim ; then, calculate the confidence reweighting matrix M cof (j, k), and the specific formula is as follows: where α and β are balance coefficients, and the confidence reweighting matrix M cof (the value within (j, k)) can be regarded as the cognitive consensus score corresponding to the visual semantics and the text semantics; after obtaining the confidence reweighting matrix M cof (j, k), find the maximum score value in the matrix, and obtain the corresponding visual label at the maximum value as the semantic label for modality alignment; transmit the semantic label for modality alignment back to the four hierarchical stages of the visual encoder in the form of gradient backpropagation, and obtain the class activation weights Step 3: Obtain class activation weights After that, integrate the weights containing semantic-level alignment information into the features extracted by the encoder. The specific formula is as follows: Among them, σ represents the sigmoid function operation, and Avg refers to the averaging operation. represents the element-wise multiplication with the broadcasting mechanism; the attention module guided by cognitive consensus integrates the semantic-level cognitive consensus weights and the video feature-level representations in a channel-spatial form to obtain the integrated information V i r , which is used to guide the network for subsequent segmentation. Step 4: First, map and repeat the audio feature F a to obtain Next, pass the visual feature V at this level i through the dilated convolution module to obtain V i a and input it together with the audio feature into the non-local module for audio-visual feature-level cross-modal interaction. The specific formula is as follows: M i = V i a + θ 4 (Φ · θ 3 (V i a )) (8) θ in the formula 1 , θ 2 , θ 3 and θ 4 respectively represent different three-dimensional convolutional layers, N is the number of pixels in the feature spectrum, Φ is the cross-modal attention matrix, and M i is the multi-modal feature at the i-th level; Step 5: Fuse the hierarchical multi-modal feature M i The fusion is performed according to the following formula: In the formula, Conv represents the convolutional layer, and Upsample represents the upsampling operation. Send Y 1 into the fully convolutional network to obtain the predicted value of the network Finally, train the network using binary cross-entropy loss: Among them, Y represents the true value of the segmentation mask, and L seg represents the loss value.