Semantic consistency-based open vocabulary audiovisual segmentation method
By introducing the CLIP and CLAP models and combining cross-modal attention and self-attention mechanisms, the generalization and semantic drift problems of audio-visual segmentation methods in open-world scenarios are solved, more efficient audio-visual semantic segmentation is achieved, and the robustness and accuracy of the model are improved.
Patent Information
- Application Number
- CN202511311730.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing audio-visual segmentation methods lack generalization capabilities in open-world scenarios, have difficulty handling multi-source sound overlap and semantic drift, and lack cross-modal audio-visual consistency, which affects the robustness and flexibility of audio-visual semantic segmentation.
The Bridging Vision-Language Model (CLIP) and Audio-Language Model (CLAP) are introduced. Through a symmetric cross-modal attention guidance module and a hierarchical modal fusion decoder, combined with the cross-modal attention mechanism and self-attention mechanism, they explicitly model audio-visual semantic associations, enhance the audio semantic discrimination ability, and improve the model generalization ability through auxiliary audio supervision loss.
It significantly improves the semantic consistency between audiovisual modalities and the generalization performance of the model, improves the recognition and segmentation accuracy of sound-making objects, and enhances the segmentation and classification capabilities of unknown categories.
Smart Images

Figure CN120822079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and multimodal (audio-visual) information processing, and relates to an open vocabulary audio-visual segmentation method based on semantic consistency. Background Art
[0002] With the advancement of multimedia information processing technology, audio-visual segmentation (AVS) technology has emerged, aiming to accurately separate the sound-producing visual objects from video and audio data. This technology has broad application prospects in a variety of fields, including virtual reality, security surveillance, and human-computer interaction. Existing AVS methods primarily focus on identifying a limited number of predefined categories in the training data and lack generalization capabilities to new, unseen categories. This limits their application in open-world scenarios and for identifying unknown sound sources, and also affects the robustness and flexibility of AVS systems. To address these challenges, researchers recently proposed the new task of Open-Vocabulary Audio-Visual Semantic Segmentation (OVAVSS), formally defining the problem and constructing a related evaluation benchmark. OVAVSS aims to segment and identify the corresponding sound-producing objects given video and audio information, including novel sound-producing objects not seen during training. However, OV-AVSS employs video-level decoding and a category-independent mask generation strategy, which presents challenges in temporal alignment and fails to explicitly align audiovisual semantics. This makes it difficult to effectively handle issues such as multi-source sound overlap in complex scenarios. Furthermore, due to its over-reliance on pre-trained visual-language models, semantic drift and misclassification can easily occur when the audio signal is ambiguous or semantically unclear. Furthermore, the two-stage audiovisual semantic segmentation process of first segmentation and then classification weakens cross-modal audiovisual consistency, affecting the inherent correlation and alignment of the overall audiovisual semantics.
[0003] Therefore, this paper addresses the shortcomings and challenges of existing research by innovatively introducing a strategy of bridging the vision-language model (CLIP) and the audio-language model (CLAP). This strategy leverages the pre-trained knowledge and synergy of these two models to effectively address semantic drift and misclassification, while improving audiovisual consistency. By deeply mining audio cues, it achieves accurate recognition and classification of sounding objects, significantly improving the deep fusion and semantic alignment of audiovisual modalities, and thus achieving open-vocabulary audiovisual segmentation with strong generalization capabilities. Summary of the Invention
[0004] This paper proposes a semantically consistent open-vocabulary audiovisual segmentation method, achieving robust open-vocabulary audiovisual semantic segmentation. This method effectively improves semantic consistency between audiovisual modalities, enhances the ability to identify and segment sound-producing objects, and significantly improves the generalization performance of the model.
[0005] The technical solution of the present invention:
[0006] An open vocabulary audiovisual segmentation method based on semantic consistency is as follows:
[0007] (1) Symmetric cross-modal attention guidance module, the specific implementation steps are as follows:
[0008] Audio features encoded by the CLAP audio codec Multi-scale visual features encoded by CLIP image encoder First, the multi-scale visual features are projected into a unified feature dimension through convolution , then expanded along the spatial dimension and merged into multi-scale features, we get ; At the same time, the audio features are mapped to the same feature dimension through the multi-layer perceptron, denoted as ; Then use the cross-modal attention mechanism to fuse and , we get the fused multimodal features, which are expressed as :
[0009]
[0010] in, Represents row-by-row normalization activation function, represents a learnable projection matrix;
[0011] In order to fully capture the complex audio-visual association, two cross-modal interaction layers are designed, which use visual features to and audio characteristics For query, the fused multimodal features are used As a key or value, through the cross-modal attention mechanism The fusion information in the selective injection of visual features and audio characteristics , respectively, to obtain enhanced visual features and audio features, guiding the enhanced visual features and audio features to focus more accurately on the sound object; the corresponding calculation formula is as follows:
[0012]
[0013]
[0014] in, and Represent the enhanced visual features and audio features respectively; 、 Both are learnable projection matrices;
[0015] Then, It is input into the pixel decoder based on multi-scale deformable attention for multi-scale feature interaction and outputs more accurate visual features, denoted as ;
[0016] (2) Hierarchical modal fusion decoder, its specific implementation steps are:
[0017] First, the enhanced audio features of the input are temporally expanded, and the audio features are optimized through the self-attention mechanism and temporal context association to obtain the temporally enhanced audio features, which are denoted as Then, a set of target queries initialized with learnable parameters are fed with temporally enhanced audio features via cross-modal attention. Interact to get the initial target query embedded in the audio clue ;
[0018] The hierarchical modality fusion decoder consists of three decoding layers, in each of which the object query is combined with the visual features output by the pixel decoder. Perform cross-modal attention interaction to extract the visual information of the object corresponding to the vocalization clue and obtain the updated target query :
[0019]
[0020] in, ; Represents the target query output from the previous decoding layer. When the decoding layer is the first layer, represents the cross-modal attention mechanism;
[0021] Subsequently, the updated target query is further subjected to fine-grained association modeling through the self-attention layer to improve its discriminative ability. In order to maintain audio clues and enhance the recognition ability of complex multi-source scenes, the updated target query is further interacted with the audio features to obtain a target query that integrates audio information. :
[0022]
[0023] Next, the target query is explicitly modeled by performing self-attention calculation on the target query that integrates audio information in the spatiotemporal dimension. Cross-frame association further enhances the temporal consistency of the target query:
[0024]
[0025] in, Represents the self-attention mechanism; Indicates that the shape is The tensor is expanded to the shape The target query, The number of target queries, represents the channel dimension; Represents the reverse operation of restoring the result of the above expansion operation to the original shape; finally, the target query with enhanced temporal consistency is performed After being processed by the feedforward network, the target query that integrates the audio-visual information is output, which is denoted as ;
[0026] (3) Audio semantic enhancement module, its specific implementation steps are as follows:
[0027] First target query Audio features with timing enhancement By interacting with cross-modal attention, each target query selectively aggregates relevant audio information and explicitly integrates audio semantics; residual connections are then applied to the attention outputs, and the fused features are further integrated through a multi-layer perceptron. ; For the category, the CLAP text prompt template "This sound contains the {}" is used, and the text embedding is obtained by the CLAP text encoder ; Finally, the fusion features and text embedding Calculate the similarity through dot product to obtain the final audio semantic classification result ; The above process is expressed as:
[0028]
[0029]
[0030] In order to make full use of CLAP's pre-training knowledge, enhance the ability to distinguish audio semantics, and improve the ability of the open vocabulary audio-visual segmentation model to accurately use audio to identify the sounding object, the audio semantic enhancement module introduces a semantic classification loss for auxiliary audio supervision. ; The loss formula is expressed as:
[0031]
[0032] in, represents the cross entropy loss calculation, Indicates the true category label corresponding to the audio.
[0033] (4) Model mask generation and overall loss calculation, the specific steps are:
[0034] In order to achieve open vocabulary audio-visual semantic segmentation, a CLIP-based classification head and a mask head are used for audio-visual semantic segmentation; in the mask head, the target query After being processed by a multi-layer perceptron, the highest resolution visual features output by the pixel-level decoder are Recorded as mask features, the corresponding segmentation mask prediction is generated by matrix product operation In the CLIP-based classification head, the segmentation mask prediction and mask features are pooled through mask pooling to obtain mask pooling features, which are then combined with the target query Perform residual connection and obtain category features through multi-layer perceptron processing; at the same time, text embedding The CLIP text prompt template "A photo of {}" is used for the category and generated by the CLIP text encoder; finally, the category features are embedded with the text Calculate similarity through dot product and generate corresponding category prediction ;
[0035] During the training process, this design considers both semantic classification loss and prediction mask loss; the semantic classification loss consists of two parts: one is the main classification loss based on CLIP from the classification head , and the second is the CLAP-based auxiliary audio supervision classification loss introduced by the audio semantic enhancement module ; the total loss is as follows:
[0036]
[0037]
[0038] Among them, the mask loss function By predicting the mask and the true mask Calculate focus and Diceloss; 、 、 represent weighting parameters, That is, the total loss function;
[0039] In the inference phase, the segmentation prediction mask generated by the mask header is used , another set of semantic classification predictions is obtained by performing similarity calculations with the text embedding containing unknown categories encoded using the frozen CLIP text encoder through mask pooling ; Among them, the classification head output Responsible for classification and identification of known categories, Focuses on the recognition of unknown categories; subsequently, and Through geometric integration, the final semantic category prediction is generated, which has both the ability to discriminate known categories and the generalization ability of open vocabulary. Finally, the semantic category prediction is combined with the prediction mask to generate the audiovisual semantic segmentation mask.
[0040] Beneficial effects of the present invention:
[0041] (1) This paper explicitly enhances the audio semantic recognition capability by designing a novel audio semantic enhancement module, improves the model's cross-modal alignment and semantic recognition accuracy, and thus enhances the robustness and accuracy of audio-visual semantic segmentation.
[0042] (2) Two innovative cross-modal fusion interaction modules are proposed: a symmetrical cross-modal attention guidance module and a hierarchical modal fusion decoder. Through refined cross-modal interaction and multimodal decoding, the spatiotemporal semantics of audiovisual information are fully explored, audiovisual features are aggregated in spatial and temporal dimensions, and the precise positioning and classification of sound-generating objects are effectively ensured.
[0043] (3) By jointly using CLIP and CLAP and aligning audiovisual features based on shared true labels, the present invention not only enhances the segmentation performance of known categories of sound-making objects, but also significantly improves the segmentation and classification capabilities of unknown categories and the generalization ability of the model in open vocabulary scenarios through the knowledge of pre-trained basic models. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is the overall architecture diagram of the open vocabulary audio-visual segmentation method based on semantic consistency involved in this application (the figure includes the structure diagram of the audio semantic enhancement module).
[0045] Figure 2 This is a detailed diagram of the symmetrical cross-modal attention guidance module within the method involved in this application.
[0046] Figure 3 This is a detailed diagram of the hierarchical modal fusion decoder involved in the method of this application. DETAILED DESCRIPTION
[0047] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.
[0048] Reference Figure 1 , an open vocabulary audio-visual segmentation method based on semantic consistency, comprising the following steps:
[0049] Step 1: Audiovisual encoding, including the following steps:
[0050] For image encoding, this design uses the CLIP image encoder to extract visual features from each video frame, generating multi-scale visual features. Simultaneously, the CLAP audio encoder processes the audio and extracts corresponding audio features. Furthermore, during retraining, all training categories (known categories) are combined with the CLIP and CLAP text hint templates, respectively. The respective text encoders (CLIP text encoder and CLAP text encoder) then output corresponding CLIP-encoded and CLAP-encoded text embeddings. During inference, all categories (both known and unknown) are used for text encoding.
[0051] Step 2: Optimize the audiovisual interaction through a symmetric cross-modal attention guidance module. The specific steps are as follows:
[0052] The multi-scale visual features encoded by CLIP are mapped to a unified feature dimension through convolutional layers, then spatially expanded and merged along the sequence length dimension. Simultaneously, the audio features encoded by CLAP are mapped to the same feature dimension as the multi-scale visual features through a multi-layer perceptron. The pre-processed audiovisual features are input to a symmetric cross-modal attention guidance module, which utilizes complementary audio-visual information to bidirectionally optimize the visual and audio features, resulting in improved audio and visual features.
[0053] Step 3: The decoding phase based on the hierarchical modality fusion decoder includes the following steps:
[0054] First, the improved audio features output from step 2 are temporally expanded, and the audio features are optimized through the self-attention mechanism and temporal context association to obtain temporally enhanced audio features. Secondly, a set of target queries initialized with learnable parameters interact with the temporally enhanced audio features through cross-modal attention to obtain the initial target query embedded with audio clues. The target query interacts with the audio-visual features layer by layer through the decoding layer of the hierarchical modal fusion decoder, aggregates the audio-visual information of the sounding object, and outputs the final target query. .
[0055] Step 4: Predict the training loss function assisted by the generation and audio semantic enhancement modules. The specific steps are as follows:
[0056] This design proposes a one-stage open vocabulary audiovisual semantic segmentation framework to uniformly generate prediction masks and corresponding semantic categories. Specifically, in the mask header, the target query After being processed by a multi-layer perceptron, the highest resolution visual features output by the pixel-level decoder are Recorded as mask features, the corresponding segmentation mask prediction is generated by matrix product operation In the CLIP-based classification head, the segmentation mask prediction and mask features are pooled through mask pooling to obtain mask pooling features, which are then combined with the target query Perform residual connection and obtain category features through multi-layer perceptron processing; at the same time, text embedding The CLIP text prompt template "A photo of {}" is used for the category and generated by the CLIP text encoder; finally, the category features are embedded with the text Calculate similarity through dot product and generate corresponding category prediction The audio semantic enhancement module explicitly models the audio semantics and outputs an auxiliary audio semantic classification based on the frozen CLAP text encoder. .
[0057] During the training process, this design considers both semantic classification loss and prediction mask loss; semantic classification loss consists of two parts: one is the main classification loss based on CLIP from the classification head , and the second is the CLAP-based auxiliary audio supervision classification loss introduced by the audio semantic enhancement module The total losses are as follows:
[0058]
[0059] Among them, the mask loss function By predicting the mask and the true mask Calculate focus and Diceloss; 、 、 represent weighting parameters, That is the total loss function.
[0060] In the inference phase, the present invention draws on the FC-CLIP strategy to achieve accurate and robust open vocabulary audio-visual semantic segmentation. Specifically, the segmentation prediction mask generated by the mask header is used , another set of semantic classification predictions is obtained by performing similarity calculations with the text embedding containing unknown categories encoded using the frozen CLIP text encoder through mask pooling ; Among them, the classification head output Responsible for classification and identification of known categories, Focuses on the recognition of unknown categories; subsequently, and Through geometric integration, the final semantic category prediction is generated, which has both the ability to discriminate known categories and the generalization ability of open vocabulary. Finally, the semantic category prediction is combined with the prediction mask to generate the audiovisual semantic segmentation mask.
[0061] Experimental verification:
[0062] 1. Verification details
[0063] This method uses the ConvNeXt-Large CLIP model and the HTS-AT CLAP model. The pixel decoder employs a multi-scale deformable attention deformer architecture. In the decoder, the number of learnable queries is set to 100. Following previous work on audio-visual segmentation, various data augmentation techniques are introduced, including flipping and random cropping. Training uses the AdamW optimizer with a batch size of 1 and a learning rate of 0.0001 to improve training stability.
[0064] During the training process, the weight coefficients of the loss function are set to =2.0, , = 5.0, trained on the AVSBench_OV dataset for 96228 iterations.
[0065] To ensure fairness, we follow previous studies and use the same evaluation metrics: mean Intersection-over-Union (mIoU), mean Intersection-over-Union (IoU) of known classes (Base mIoU), mean Intersection-over-Union (IoU) of unknown classes (Novel mIoU), and the harmonic mean of the two (Harmonic mIoU). The calculation formula for Harmonic mIoU is as follows:
[0066]
[0067] Quantitative Results
[0068] Table 1 compares the open vocabulary audiovisual semantic segmentation model proposed in this paper with existing state-of-the-art methods. As can be seen from the data in this table, the model designed in this application achieves excellent performance across all evaluation metrics, significantly outperforming the OV-AVSS method. These results fully demonstrate the balanced and robust generalization capabilities of this model.
[0069] Table 1 Comparison results with current advanced methods on the AVSBench-OV dataset
[0070] .
Claims
1. An open vocabulary audio-visual segmentation method based on semantic consistency, characterized by: The details are as follows: (1) Symmetric cross-modal attention guidance module, the specific implementation steps are as follows: Audio features encoded by the CLAP audio codec Multi-scale visual features encoded by CLIP image encoder First, the multi-scale visual features are projected into a unified feature dimension through convolution , then expanded along the spatial dimension and merged into multi-scale features, we get ; At the same time, the audio features are mapped to the same feature dimension through the multi-layer perceptron, denoted as ; Then use the cross-modal attention mechanism fusion and , we get the fused multimodal features, which are expressed as : in, Represents row-by-row normalization activation function, represents a learnable projection matrix; In order to fully capture the complex audio-visual association, two cross-modal interaction layers are designed, which use visual features to and audio characteristics For query, the fused multimodal features are used As a key or value, through the cross-modal attention mechanism The fusion information in the selective injection of visual features and audio characteristics , respectively, to obtain enhanced visual features and audio features, guiding the enhanced visual features and audio features to focus more accurately on the sound object; the corresponding calculation formula is as follows: in, and Represent the enhanced visual features and audio features respectively; 、 Both are learnable projection matrices; Then, It is input into the pixel decoder based on multi-scale deformable attention for multi-scale feature interaction and outputs more accurate visual features, denoted as ; (2) Hierarchical modal fusion decoder, its specific implementation steps are: First, the enhanced audio features of the input are temporally expanded, and the audio features are optimized through the self-attention mechanism and temporal context association to obtain the temporally enhanced audio features, which are denoted as Then, a set of target queries initialized with learnable parameters are fed with temporally enhanced audio features via cross-modal attention. Interact to get the initial target query embedded in the audio clue ; The hierarchical modality fusion decoder consists of three decoding layers, in each of which the object query is combined with the visual features output by the pixel decoder. Perform cross-modal attention interaction to extract the visual information of the object corresponding to the vocalization clue and obtain the updated target query : in, ; Represents the target query output from the previous decoding layer. When the decoding layer is the first layer, represents the cross-modal attention mechanism; Subsequently, the updated target query is further interacted with the audio features to obtain the target query that integrates the audio information : Next, the target query is explicitly modeled by performing self-attention calculation on the target query that integrates audio information in the spatiotemporal dimension. Cross-frame association further enhances the temporal consistency of the target query: in, Represents the self-attention mechanism; Indicates that the shape is The tensor is expanded to the shape The target query, The number of target queries, represents the channel dimension; Represents the reverse operation of restoring the result of the above expansion operation to the original shape; finally, the target query with enhanced temporal consistency is performed After being processed by the feedforward network, the target query that integrates the audio-visual information is output, which is denoted as ; (3) Audio semantic enhancement module, its specific implementation steps are as follows: First target query Audio features with timing enhancement By interacting with cross-modal attention, each target query selectively aggregates relevant audio information and explicitly integrates audio semantics; residual connections are then applied to the attention outputs, and the fused features are further integrated through a multi-layer perceptron. ; For the category, the CLAP text prompt template "This sound contains the {}" is used, and the text embedding is obtained by the CLAP text encoder ; Finally, the fusion features and text embedding Calculate the similarity through dot product to obtain the final audio semantic classification result ; The above process is expressed as: The audio semantic enhancement module introduces a semantic classification loss for auxiliary audio supervision ; The loss formula is expressed as: in, represents the cross entropy loss calculation, Indicates the true category label corresponding to the audio; (4) Model mask generation and overall loss calculation, the specific steps are: In order to achieve open vocabulary audio-visual semantic segmentation, a CLIP-based classification head and a mask head are used for audio-visual semantic segmentation; in the mask head, the target query After being processed by a multi-layer perceptron, the highest resolution visual features output by the pixel-level decoder are Recorded as mask features, the corresponding segmentation mask prediction is generated by matrix product operation In the CLIP-based classification head, the segmentation mask prediction and mask features are pooled through mask pooling to obtain mask pooling features, which are then combined with the target query Perform residual connection and obtain category features through multi-layer perceptron processing; at the same time, text embedding By using the CLIP text prompt template "A photo of {}" for the category, it is generated by the CLIP text encoder; finally, the category features are embedded with the text Calculate similarity through dot product and generate corresponding category prediction ; During the training process, this design considers both semantic classification loss and prediction mask loss; the semantic classification loss consists of two parts: one is the main classification loss based on CLIP from the classification head , and the second is the CLAP-based auxiliary audio supervision classification loss introduced by the audio semantic enhancement module ; the total loss is as follows: Among them, the mask loss function By predicting the mask and the true mask Calculate focus and Diceloss; 、 、 represent weighting parameters, That is, the total loss function; In the inference phase, the segmentation prediction mask generated by the mask header is used , another set of semantic classification predictions is obtained by performing similarity calculations with the text embedding containing unknown categories encoded using the frozen CLIP text encoder through mask pooling ; Among them, the classification head output Responsible for classification and identification of known categories, Focuses on the recognition of unknown categories; subsequently, and Through geometric integration, the final semantic category prediction is generated, which has both the ability to discriminate known categories and the generalization ability of open vocabulary. Finally, the semantic category prediction is combined with the prediction mask to generate the audiovisual semantic segmentation mask.
Citation Information
Patent Citations
Speech recognition method and device and storage medium
CN118116384A
Audio-visual video question and answer method and system based on multi-modal heterogeneous graph
CN119311842A
Audio-visual segmentation method integrated with frequency domain design
CN119693857A
Weakly supervised semantic segmentation method and apparatus based on attention mask
WO2025060272A1
Cited By
Power equipment image comparison retrieval method and system based on semantic object relationship
CN121030032A
Weak supervision scene understanding method, system and equipment for multi-modal information interaction
CN121305053A
Construction method, device and equipment of model for remote sensing image target detection
CN121789067A
Training method, audiovisual segmentation method, electronic device and storage medium
CN122090357A
Training method, audiovisual segmentation method, electronic device, and storage medium
CN122090357B