Audio-visual segmentation method based on multimode understanding
By introducing a multi-mode understanding architecture into the audio-visual segmentation method, and utilizing the fusion of visual and audio features, the problem of insufficient audio-visual segmentation performance in the prior art is solved, and higher robustness and accuracy are achieved.
Patent Information
- Application Number
- CN202510357024.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing audio-visual segmentation methods have shortcomings in performance, and it is difficult to effectively improve the robustness and accuracy of audio-visual segmentation tasks.
The audio-visual segmentation method based on multi-mode understanding is adopted, and audio-visual features are extracted through the visual backbone and the audio-visual backbone, and combined with the visual encoder and the audio-visual mixing module, the features are further processed and fused to generate the final segmentation mask.
It significantly improves the robustness and accuracy of audio-visual segmentation tasks, especially in mono- and multi-sound source scenarios, and improves the quality and efficiency of segmentation.
Smart Images

Figure CN120219746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an audio-visual segmentation method and model, in particular to an audio-visual segmentation method based on multimodal understanding. Background Art
[0002] In recent years, multimodal tasks have attracted great attention in the research community. Among them, text-visual tasks have drawn great interest from researchers. Many works have focused on related tasks such as visual question answering and visual grounding. In addition to text-visual tasks, audio-visual tasks are also becoming a hot topic. Related tasks include audio-visual correspondence, audio-visual event localization, and sound source localization. At the same time, many works have proposed unified architectures to process multimodal inputs. Most of these works are based on the Transformer architecture and demonstrate strong cross-modal capabilities. Their success highlights the reliability of transformers in the multimodal field.
[0003] Benefiting from the research and development in the multimodal field, deep learning methods have also made remarkable progress in audio-visual segmentation tasks. Among them, structures such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are used for feature extraction and fusion of audio and video. Many studies are dedicated to leveraging the correlation between audio and video signals for cross-modal information fusion to improve the accuracy of segmentation. Some studies use large-scale multimodal datasets for training and evaluation to promote the generalization ability and performance of the model. Transfer learning is used to transfer knowledge learned from one domain to another to improve the performance in audio-visual segmentation tasks. Some studies explore unsupervised or weakly supervised methods that do not require manually labeled data to reduce the cost of manual labeling. Domestic researchers have also borrowed deep learning methods, especially deep learning techniques in audio and video signal processing. Some domestic researchers are committed to constructing domestic datasets suitable for audio-visual segmentation tasks to better adapt to domestic speech and video scenarios. In China, audio-visual segmentation technology is also applied to fields such as speech recognition, intelligent monitoring, and video retrieval, similar to abroad. Some studies focus on comparing with traditional feature engineering-based methods to verify the superiority of deep learning methods. In China, audio-visual segmentation technology is also applied to fields such as audio and video websites, intelligent voice assistants, and smart homes to provide better experiences and services for users.
[0004] Generally speaking, the research on audio-visual segmentation tasks at home and abroad is constantly progressing, using technologies such as deep learning and cross-modal information fusion to continuously improve the accuracy and efficiency of segmentation to meet the increasingly complex and diverse requirements of audio and video scenarios. This also poses higher requirements for the research on audio-visual segmentation tasks. How to achieve a more ideal segmentation effect in a relatively simple way has important practical significance. Summary of the Invention
[0005] Objective of the Invention: The technical problem to be solved by the present invention is to provide an audiovisual segmentation method based on multimodal understanding in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, the present invention discloses an audiovisual segmentation method based on multimodal understanding, wherein the method includes:
[0007] Step 1, pre-train the model on an audiovisual segmentation public dataset (such as AVSBench) to obtain pre-trained model weights for subsequent inference;
[0008] Step 2, extract audiovisual features using a visual backbone and an audio backbone respectively;
[0009] Step 3, further process the visual features using a visual encoder to obtain deep-level information that may be missed by the visual backbone in the visual features, and improve the model's recognition ability of target objects;
[0010] Step 4, use an audiovisual mixing module to mix the audiovisual features, utilize the audio features to strengthen the target object features in the visual features, and generate a segmentation mask feature;
[0011] Step 5, use a multimodal decoder to match the query and the best visual features, and combine them with the mask feature generated in the previous step to generate a final segmentation mask, completing the segmentation.
[0012] Further, the visual backbone described in Step 2 receives the input video frame image and outputs multi-scale visual features. The input image is denoted as x v :
[0013]
[0014] wherein, represents the range space, H represents the height of the input image x, and W represents the width of the input image x; the multi-scale visual features extracted by the visual backbone are represented as:
[0015]
[0016] wherein, represents the i-th multi-scale visual feature, and D represents the number of channels of the feature ; the visual backbone uses a classic ResNet structure, and its specific calculation process is as follows:
[0017]
[0018] wherein, wherein represents the i-th multi-scale visual feature, and f i,jDenotes the j-th residual block in stage i of the ResNet structure, c i Is a hyperparameter representing the number of residual blocks in stage i. The residual block consists of convolutional layers, which add the input features after convolution to the original features to alleviate the degradation phenomenon during training. The calculation method is as follows:
[0019]
[0020] Among them, And Represent the input and output features respectively; Conv 1×1 And Conv 3×3 Represent the 1×1 and 3×3 convolutional layers respectively.
[0021] Furthermore, the audio backbone described in step 2 receives the input audio data and outputs a one-dimensional vector audio feature. The input audio data has been pre-resampled to a unified 16kHz mono audio, and the audio data is denoted as x a :
[0022]
[0023] Among them, N samples Is related to the audio duration and represents the number of audio samples; Subsequently, a short-time Fourier transform is performed on it to obtain the Mel spectrogram, and the Mel spectrogram is calculated by mapping the spectrum to a 64-order Mel filter bank; Finally, this spectrogram is fed into the VGGish network to extract the final audio feature. The specific calculation process is as follows:
[0024]
[0025] Among them, Fourier represents the short-time Fourier transform, Mel represents the spectrum mapping, Represents the audio backbone network, and Represents the final audio feature; Here, T represents the number of video frames corresponding to the audio, and D represents the embedding dimension of the model.
[0026] Furthermore, the use of the visual encoder to further process the visual features described in step 3 specifically includes:
[0027] Step 3-1, extract the last three layers of features from the multi-scale visual features, and splice them after unfolding to obtain the visual feature encoding vector, denoted as
[0028]
[0029] Among them, N vrepresents the total length after expanding three visual features, and D represents the embedding dimension of the model.
[0030] Step 3-2: Use a visual encoder to perform self-attention calculation on the concatenated features above to obtain a new visual feature encoding vector Its shape remains unchanged and is consistent with Keep consistent.
[0031] Step 3-3: Split the new visual feature encoding, restore it to the original visual feature size, upsample it, and add it to the original 1 / 4 resolution feature output in the visual backbone to obtain the final mask feature.
[0032] Furthermore, the feature flattening and concatenation described in Step 3-1, the specific process is as follows:
[0033]
[0034] Among them, Flatten represents the flattening operation, Concat represents the concatenation operation, and is the visual feature encoding vector, where N v represents the total length after expanding three visual features, and D represents the embedding dimension of the model.
[0035] Furthermore, the visual encoder described in Step 3-2 is composed of several self-attention layers. Each self-attention layer takes the visual feature encoding vector output by the previous layer as input to perform a new self-attention operation and outputs to the next layer; the result output by the last layer is the final visual feature encoding, and the specific process is as follows:
[0036]
[0037] Among them, s i represents the i-th self-attention layer, n is a hyperparameter representing the total number of self-attention layers, represents the composite operation; the self-attention layer performs attention operation on the visual feature encoding as both the query Q, key K, and key value V at the same time, and the specific calculation method is as follows:
[0038]
[0039] Among them and represent the input and output feature encodings respectively; D represents the embedding dimension, that is, the number of channels in the visual feature; represents matrix multiplication.
[0040] Further, in the splitting and restoring process described in step 3-3, the visual feature encoding output by the visual encoder is disassembled according to the original feature size and restored to the original size, and then these features are sampled to the same size as the largest feature through upsampling, and added together to obtain the mask feature. The specific process is as follows: Consistent, and add them together to get the mask feature. The specific process is as follows:
[0041]
[0042] Among them, slice i Indicates cutting out the corresponding part from the visual feature encoding according to the size of the original visual feature In the visual feature encoding The corresponding part is cut out, and reshape represents the operation of restoring the visual feature from the visual feature encoding. That is, the restored visual feature, whose size is the same as the original visual feature Consistent; upsample represents the upsampling operation; That is the mask feature.
[0043] Further, the audio-visual mixing module described in step 4 adopts a channel attention mechanism, performs cross-attention calculation on the audio feature and the mask feature, and then multiplies them channel by channel to enhance the features of the target object. The channel attention calculation takes the audio feature as the query Q, the mask feature as the key K and the value V for attention operation. The specific calculation method is as follows:
[0044]
[0045] Among them Represents the intermediate quantity after cross-attention operation, Is the audio feature extracted by the audio backbone, Is the initial mask feature generated by the visual encoder, and D represents the embedding dimension of the model. Subsequently, multiply the intermediate quantity of the audio feature by the initial mask feature, that is, the final mask feature is obtained. The calculation method is as follows:
[0046]
[0047] Among them Represents the mask feature after audio-visual mixing enhancement, Represents matrix multiplication.
[0048] Further, the multi-modal decoder described in step 5 completes the segmentation, specifically including:
[0049] Step 5-1, use the multi-modal decoder to copy the audio feature as the initial query, perform cross-attention calculation with the visual feature, and obtain the audio-visual mixed query.
[0050] Step 5-2: Multiply the mask feature by the hybrid query to obtain a feature matching vector.
[0051] Step 5-3: Use a multi-layer perceptron to filter the feature matching vector and generate a final segmentation mask.
[0052] Furthermore, the multi-modal decoder described in Step 5-1 is similar in structure to the visual encoder and both adopt the attention mechanism. First, the decoder copies the audio feature N q times and concatenates them as the initial query, expressed as:
[0053]
[0054] where represents the initial query, N q is a hyperparameter representing the number of queries, and Concat is the concatenation operation. The decoder consists of multiple cross-attention layers and performs layer-by-layer operations similar to the encoder. The process is as follows:
[0055]
[0056] where represents the audio-visual hybrid query, d i represents the i-th cross-attention layer in the decoder, and n is a hyperparameter representing the number of cross-attention layers. Each cross-attention layer uses the query from the previous layer as the query Q, and encodes the visual feature as the key K and value V for operation. The calculation method is:
[0057]
[0058] where represents the hybrid query output by the i-th cross-attention layer.
[0059] Furthermore, for the feature matching vector described in Step 5-2, its calculation method is:
[0060]
[0061] where is the feature matching vector.
[0062] Furthermore, the multi-layer perceptron described in Step 5-3 consists of several linear layers. Through operation, it changes the number of channels of the feature matching vector to the number of semantic categories, and finally obtains the segmentation mask after upsampling the features. The specific process is as follows:
[0063]
[0064] where L represents the linear layer and upsample represents the upsampling operation. Represents the final segmentation mask, where N cls Represents the number of semantic categories for segmentation.
[0065] Beneficial effects:
[0066] (1) In view of the performance deficiency problems existing in the existing audio-visual segmentation methods, the present invention proposes a method based on multi-modal understanding. This method adopts an encoder-decoder architecture, extracts deeper visual features, and fully analyzes audio features to locate sound sources, greatly improving the robustness of the audio-visual segmentation task.
[0067] (2) The present invention innovatively proposes an audio-visual hybrid module, which uses audio information to correct visual features, strengthens the target objects therein, and at the same time suppresses the remaining non-target objects, effectively improving the audio-visual segmentation ability of the model. Description of the drawings
[0068] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0069] Figure 1 Is the overall architecture diagram of the present invention.
[0070] Figure 2 Is the display of the effect comparison of the present invention.
[0071] Figure 3 Is the effect diagram of the influence of the key components of the present invention on visual features. Specific embodiments
[0072] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the protection scope of the present invention.
[0073] The following describes in detail the application principle of the present invention in conjunction with the drawings.
[0074] Embodiment 1
[0075] This embodiment proposes an audio-visual segmentation method based on multi-modal understanding. Refer to Figure 1 , specifically including:
[0076] Step 1, pre-train the model on the audio-visual segmentation dataset to obtain model weights;
[0077] Step 2, extract audio-visual features using the visual backbone and audio backbone respectively;
[0078] Step 3, further process the visual features using the visual encoder to obtain deeper visual information and richer semantic information;
[0079] Step 4: Use the audio-visual hybrid module to mix the audio-visual features, enhance the target object features in the visual features with the audio features, and generate a segmentation mask feature.
[0080] Step 5: Use the multi-modal decoder to match the query and the best visual features, and combine them with the mask feature generated in the previous step to generate the final segmentation mask, completing the segmentation.
[0081] Preferably, for the visual backbone described in step 2, it receives the input video frame image and outputs multi-scale visual features. The input image is represented as:
[0082]
[0083] where represents the value range space, H represents the height of the input image x, and W represents the width of the input image x; the multi-scale visual features extracted by the visual backbone are represented as:
[0084]
[0085] where represents the i-th multi-scale visual feature, and D represents the number of channels of the feature ; the visual backbone uses the classic ResNet structure, and its specific calculation process is as follows:
[0086]
[0087] where represents the i-th multi-scale visual feature, f i,j represents the j-th residual block in stage i of the ResNet structure, and c i is a hyperparameter representing the number of residual blocks included in stage i. The residual block is composed of convolutional layers, and the input feature is added to the original feature after convolution operation to alleviate the degradation phenomenon during training. The calculation method is:
[0088]
[0089] where and represent the input and output features respectively; Conv 1×1 and Conv 3×3 represent the 1×1 and 3×3 convolutional layers respectively.
[0090] Preferably, for the audio backbone described in step 2, it receives the input audio data and outputs one-dimensional vector audio features. The input audio data has been pre-resampled to a unified 16kHz mono audio, represented as:
[0091]
[0092] where N samples is related to the audio duration and represents the number of audio samples; subsequently, a short-time Fourier transform is performed on it to obtain the Mel spectrum, and the Mel spectrogram is calculated by mapping the spectrum to a 64-order Mel filter bank; finally, this spectrogram is fed into the VGGish network to extract the final audio features, and the specific calculation process is as follows:
[0093]
[0094] where Fourier represents the short-time Fourier transform, Mel represents the spectrum mapping, represents the audio backbone network, and represents the final audio features; here, T represents the number of video frames corresponding to the audio, and D represents the embedding dimension of the model.
[0095] Preferably, the further processing of the visual features using the visual encoder in step 3 specifically includes:
[0096] Step 3-1, extract the last three layers of features from the multi-scale visual features, and splice them after unfolding to obtain the visual feature encoding vector, denoted as
[0097]
[0098] where N v represents the total length after unfolding the three visual features, and D represents the embedding dimension of the model.
[0099] Step 3-2, perform self-attention calculation on the above-mentioned spliced features using the visual encoder to obtain a new visual feature encoding vector whose shape remains unchanged and is consistent with
[0100] Step 3-3, split the new visual feature encoding, restore it to the original visual feature size, perform upsampling on it, and add it to the original 1 / 4 resolution feature output from the visual backbone to obtain the final mask feature.
[0101] More preferably, the flattening and concatenation of the features described in step 3-1, the specific process is as follows:
[0102]
[0103] where Flatten represents the flattening operation, Concat represents the concatenation operation, and is the visual feature encoding vector, and here Nv represents the total length after the expansion of three visual features, and D represents the embedding dimension of the model.
[0104] Further preferably, the visual encoder described in step 3-2 is composed of a plurality of self-attention layers. Each self-attention layer takes the visual feature encoding vector output by the previous layer as input for a new self-attention operation and outputs it to the next layer; the result output by the last layer is the final visual feature encoding. The specific process is as follows:
[0105]
[0106] where s i represents the i-th self-attention layer, n is a hyperparameter representing the total number of self-attention layers, represents a composite operation; the self-attention layer performs an attention operation on the visual feature encoding as both the query Q, the key K, and the key value V at the same time. The specific calculation method is as follows:
[0107]
[0108] where and represent the input and output feature encodings respectively; D represents the embedding dimension, that is, the number of channels in the visual feature; represents matrix multiplication.
[0109] Further preferably, in the splitting and restoring process described in step 3-3, the visual feature encoding output by the visual encoder is disassembled according to the original feature size and restored to the original size, and then these features are sampled to be consistent with the largest feature through upsampling, and they are added together to obtain the mask feature. The specific process is as follows:
[0110]
[0111] where slice i represents cutting out the corresponding part from the visual feature encoding according to the size of the original visual feature , reshape represents the operation of restoring from the visual feature encoding to the visual feature, is the restored visual feature, and its size is the same as the original visual feature ; upsample represents the upsampling operation; is the mask feature.
[0112] Preferably, the audiovisual mixing module described in step 4 adopts a channel attention mechanism to perform cross-attention calculation on the audio feature and the mask feature and then multiply them channel by channel to enhance the feature of the target object. The channel attention calculation takes the audio feature as the query Q, the mask feature as the key K and the value V for attention operation. The specific calculation method is as follows:
[0113]
[0114] where represents the intermediate quantity after cross-attention operation, is the audio feature extracted by the audio backbone, is the initial mask feature generated by the visual encoder, and D represents the embedding dimension of the model. Subsequently, multiply the intermediate quantity of the audio feature by the initial mask feature, and the final mask feature can be obtained. The calculation method is:
[0115]
[0116] where represents the mask feature after audiovisual mixing enhancement, represents matrix multiplication.
[0117] Preferably, the multi-modal decoder described in step 5 completes segmentation, specifically including:
[0118] Step 5-1, use the multi-modal decoder to copy the audio feature as the initial query and perform cross-attention calculation with the visual feature to obtain an audio-visual mixed query.
[0119] Step 5-2, multiply the mask feature by the mixed query to obtain a feature matching vector.
[0120] Step 5-3, use a multi-layer perceptron to filter the feature matching vector to generate a final segmentation mask.
[0121] Further preferably, the multi-modal decoder described in step 5-1 has a similar structure to the visual encoder and both adopt the attention mechanism. The decoder first copies the audio feature N q times and concatenates them as the initial query, which is expressed as:
[0122]
[0123] where represents the initial query, N q is a hyperparameter representing the number of queries, and Concat is the concatenation operation. The decoder consists of multiple cross-attention layers and performs layer-by-layer operation similar to the encoder. The process is:
[0124]
[0125] Among them represents audio-visual hybrid query, d i represents the i-th cross-attention layer in the decoder, and n is a hyperparameter representing the number of cross-attention layers. Each cross-attention layer uses the query of the previous layer as query Q, and encodes the visual features as key K and value V for operation, and the calculation method is:
[0126]
[0127] Among them represents the hybrid query output by the i-th cross-attention layer.
[0128] Further preferably, for the feature matching vector described in step 5-2, its calculation method is:
[0129]
[0130] Among them is the feature matching vector.
[0131] Further preferably, the multi-layer perceptron described in step 5-3 is composed of several linear layers. Through operation, it changes the number of channels of the feature matching vector to the number of semantic categories, and finally obtains the segmentation mask after upsampling the features. The specific process is:
[0132]
[0133] where L represents the linear layer, and upsample represents the upsampling operation, represents the final segmentation mask, and here N cls represents the number of semantic categories of the segmentation.
[0134] In this embodiment, the test was finally carried out on the audio-visual segmentation dataset AVSBench, and the performance was compared with the existing baseline method (AVSBench). As shown in Table 1 below, the performance comparison results are shown, Figure 2 showing the specific effect of audio-visual segmentation.
[0135] Table 1 is the performance index comparison result of the present invention
[0136]
[0137] In this embodiment, the model was tested in three different scenarios: a single sound source scenario, a multi - sound source scenario, and a semantic segmentation scenario. The F - score and the mean intersection over union (mIoU) were used as evaluation metrics. The F - score is a comprehensive metric that measures the accuracy and recall of the model, balancing the influence of accuracy and recall, and comprehensively evaluating a classifier. The intersection over union (IoU), on the other hand, can measure the correlation between the predicted mask and the annotated mask, reflecting the accuracy of the overall segmentation. The specific calculation method of the F - score is as follows:
[0138]
[0139] Where TP represents the number of pixels accurately predicted, FP represents the number of background pixels misjudged as foreground, and FN represents the number of foreground pixels misjudged as background. The calculation method of IoU is:
[0140]
[0141] Where is the predicted mask, is the annotated mask, and the specific meaning represents the proportion of the number of accurately predicted pixels to the total number of pixels.
[0142] As can be seen from Table 1, on the single sound source subset, this method improved the mIoU by 3.32 and the F - score by 2.0 compared to the baseline method. On the multi - sound source subset, this method achieved a qualitative leap with an mIoU of 1.65 and an F - score of 5.0. On the semantic segmentation subset, this method also obtained a significant performance improvement, with the mIoU increasing by 6.89 percentage points and the F - score increasing by 6.8 percentage points.
[0143] The success of this method is mainly due to the visual encoder and the audio - visual hybrid module in the model design. The visual encoder well analyzes the deep - level information in visual features, such as fine edges, object textures, etc., which are difficult to fully analyze only by the visual backbone network. At the same time, the audio - visual hybrid module effectively utilizes audio information, strengthens the target objects in visual features, and suppresses non - target objects such as the background. Figure 3 respectively show the effects on the features of different target objects in visual features when not using / using these two components. It can be clearly seen from the figure that for the vocal objects (such as the right girl and the puppy), the visual encoder and the audio - visual hybrid module effectively enhance their features. At the same time, for non - sound source objects (such as the left girl, the table / bed, or the background), the audio - visual hybrid module and the visual encoder suppress them to a certain extent. These results indicate that these two components greatly improve the performance of this method, making it the current most advanced audio - visual segmentation method.
[0144] Embodiment 2
[0145] This embodiment provides an electronic device, including one or more processors;
[0146] a storage device for storing one or more computer programs;
[0147] When the computer program is executed by the processor, the steps of the method described in Embodiment 1 are implemented.
[0148] This embodiment provides the program code for the entire set of audio-visual segmentation research based on multi-modal understanding, including the model structures and implementation logics of its respective core modules.
[0149] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the inventive content of a method and system for audio-visual segmentation based on multi-modal understanding provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.
[0150] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium, including several instructions for causing a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.
[0151] The present invention provides a model idea for a method and system for audio-visual segmentation based on multi-modal understanding. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. An audio-visual segmentation method based on multimodal understanding, characterized in that: The specific steps include: S1, pre-train the model on the public audio-visual segmentation dataset to obtain the pre-trained model weights; S2, extracts audio and video features using visual backbone and audio backbone respectively; S3, using the visual encoder to further process the visual features and obtain the deep information in the visual features that is missed by the visual backbone; S4, using the audio-visual mixing module to mix the audio and video features, using the audio features to enhance the target object features in the visual features, and generate segmentation mask features; S5, uses the multimodal decoder to match the query and the best visual features, and combines them with the mask features generated in the previous step to generate the final segmentation mask and complete the segmentation.
2. The method for audio-visual segmentation based on multimodal understanding according to claim 1, characterized in that: The visual backbone described in step 2 receives the input video frame image and outputs multi-scale visual features; the input image is denoted by x v : in, represents the range space, H represents the height of the input image x, and W represents the width of the input image x; the multi-scale visual features extracted by the visual backbone are expressed as: in, represents the i-th multi-scale visual feature, and D represents the feature The number of channels; the visual backbone uses the classic ResNet structure, and its specific calculation process is as follows: Among them, represents the i-th multi-scale visual feature, f i,j represents the jth residual block of stage i in the ResNet structure, c i is a hyperparameter, which indicates the number of residual blocks contained in stage i. The residual block consists of a convolutional layer, which adds the input features to the original features after convolution operation, thereby reducing the degradation phenomenon during training. The calculation method is: in, and Represents input and output features respectively; Conv 1×1 With Conv 3×3 They represent 1×1 and 3×3 convolutional layers respectively.
3. The method for audio-visual segmentation based on multimodal understanding according to claim 1, characterized in that: The audio backbone described in step S2 receives the input audio data and outputs a one-dimensional vector audio feature; the input audio data has been resampled in advance and unified into 16kHz mono audio, and the audio data is recorded as x a : Where N samples It is related to the audio duration and represents the number of audio samples. A short-time Fourier transform is then performed to obtain the Mel spectrum, and the Mel spectrum graph is calculated by mapping the spectrum to a 64-order Mel filter bank. Finally, this spectrum graph is sent to the VGGish network to extract the final audio features. The specific calculation process is as follows: Among them, Fourier represents short-time Fourier transform, Mel represents spectrum mapping, represents the audio backbone network, and It represents the final audio feature; here T represents the number of video frames corresponding to the audio, and D represents the embedding dimension of the model.
4. The method for audio-visual segmentation based on multimodal understanding according to claim 1, characterized in that: The visual features are further processed using the visual encoder described in step S3, specifically including: S3-1, extract the last three layers of feature features from the multi-scale visual features, expand them and concatenate them to obtain the visual feature encoding vector, which is expressed as: Where N v represents the total length of the three visual features after expansion, and D represents the embedding dimension of the model; S3-2, use the visual encoder to perform self-attention calculation on the above spliced features to obtain a new visual feature encoding vector Its shape remains unchanged, Stay consistent; S3-3, splits the new visual feature encoding, restores it to the original visual feature size, upsamples it, and adds it to the original 1 / 4 resolution feature output in the visual backbone to obtain the final mask feature.
5. The method for audio-visual segmentation based on multimodal understanding according to claim 4, characterized in that: The specific process of feature flattening and stitching described in step S3-1 is as follows: Flatten represents the flattening operation, Concat represents the concatenation operation, and is the visual feature encoding vector, N v represents the total length of the three visual features after expansion, and D represents the embedding dimension of the model.
6. The method for audio-visual segmentation based on multimodal understanding according to claim 4, characterized in that: The visual encoder described in step S3-2 is composed of several self-attention layers; each self-attention layer takes the visual feature encoding vector output by the previous layer as input to perform a new self-attention operation and outputs it to the next layer; the result output by the last layer is the final visual feature encoding, and the specific process is as follows: where s i represents the i-th self-attention layer, n is a hyperparameter representing the total number of self-attention layers, Represents a composite operation; the self-attention layer encodes the visual feature as the query Q, key K, and key value V for attention operation. The specific calculation method is: in and Represent the input and output feature encoding respectively; D represents the embedding dimension, i.e., the number of channels in the visual features; Represents matrix multiplication.
7. The method for audio-visual segmentation based on multimodal understanding according to claim 4, characterized in that: The splitting and restoring process described in step S3-3 is to split the visual feature encoding output by the visual encoder according to the original feature size and restore it to the original size, and then upsample these features to the same size as the maximum feature. Consistent, and add them to get the mask feature. The specific process is: Among them, slice i According to the original visual features The size is encoded from visual features Cut out the corresponding part, reshape means to restore the visual feature encoding to the visual feature operation, This is the restored visual feature, whose size is the same as the original visual feature. Consistent; upsample indicates upsampling operation; This is the mask feature.
8. The method for audio-visual segmentation based on multimodal understanding according to claim 1, characterized in that: The audio-visual mixing module described in step S4 adopts a channel attention mechanism to perform cross-attention calculation on the audio features and the mask features and then multiply them channel by channel to enhance the features of the target object; the channel attention calculation uses the audio features as the query Q and the mask features as the key K and value V for attention calculation. The specific calculation method is: in represents the intermediate quantity after the cross attention operation, The audio features extracted for the audio backbone, is the initial mask feature generated by the visual encoder, and D represents the embedding dimension of the model; the final mask feature is obtained by multiplying the intermediate audio feature with the initial mask feature, and the calculation method is: in represents the mask feature after audio-visual mixing enhancement, Represents matrix multiplication.
9. The method for audio-visual segmentation based on multimodal understanding according to claim 1, characterized in that: The multimodal decoder in step S5 completes the segmentation, specifically including: Step S5-1, using a multimodal decoder, copying the audio features as the initial query, and performing cross-attention calculation with the visual features to obtain an audio-video mixed query; Step S5-2, multiplying the mask feature with the mixed query to obtain a feature matching vector; Step S5-3, using a multi-layer perceptron to screen the feature matching vector and generate the final segmentation mask.
10. The method for audio-visual segmentation based on multimodal understanding according to claim 9, characterized in that: The multimodal decoder described in step 5-1 is similar to the visual encoder in structure and both use the attention mechanism. The decoder first copies the audio features to N q The initial query is expressed as: in represents the initial query, N q is a hyperparameter indicating the number of queries, and Concat is a concatenation operation. The decoder consists of multiple cross-attention layers, which perform layer-by-layer operations similar to the encoder. The process is as follows: in Indicates audio and video mixed query, d i represents the i-th cross-attention layer in the decoder, and n is a hyperparameter representing the number of cross-attention layers; Each cross-attention layer takes the query of the previous layer as query Q and encodes the visual features As key K and value V, the calculation method is: in represents the mixed query output by the i-th cross-attention layer; The feature matching vector described in step S5-2 is calculated as follows: in That is the feature matching vector; The multilayer perceptron described in step S5-3 is composed of several linear layers, which changes the number of channels of the feature matching vector into the number of semantic categories through calculation, and finally obtains the segmentation mask after upsampling the features. The specific process is as follows: Where L represents the linear layer, upsample represents the upsampling operation, Represents the final segmentation mask, where N cls Indicates the number of semantic categories for segmentation.
Citation Information
Cited By
Training method, audiovisual segmentation method, electronic device and storage medium
CN122090357A