Multi-granularity fusion video clip retrieval method based on audio importance perception
By building branches of three modalities: vision, audio and text, using the audio importance perception module and multi-grained fusion strategy, the problem of insufficient utilization of audio modes in the existing methods is solved, and a more flexible semantic reasoning mechanism is achieved, which improves the accuracy and robustness of video clip retrieval.
Patent Information
- Application Number
- CN202510752246.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing video clip search methods fail to effectively utilize the importance of audio modes, resulting in performance bottlenecks in complex and diverse cross-modal semantic inference tasks and lack of dynamic regulation of the degree of contribution to audio semantics.
A learning framework is built, including branches of three modalities: vision, audio and text. Through the audio importance perception module, the fusion of audio and visual modes is dynamically regulated, and a multi-grained fusion strategy and loss function optimization model is adopted to achieve automatic perception and dynamic weighting of audio importance.
It improves the accuracy and robustness of video clip retrieval, adapts to the uncertainty of audio modes, enhances the practicality and compatibility of the system, and significantly improves the search accuracy in multimodal complex scenarios.
Smart Images

Figure CN120256674A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-modal video clip retrieval, and particularly relates to a multi-granularity fusion video clip retrieval method based on audio importance perception. Background Art
[0002] Video Moment Retrieval (VMR) refers to accurately locating the start and end time segments related to semantics from a long video according to a natural language query. This task is widely applied in fields such as video understanding, intelligent monitoring, human-computer interaction, and video search, and is one of the core issues in multi-modal content understanding. Due to its involvement of complex cross-modal semantic alignment and reasoning mechanisms, it has long been a research direction that has received key attention from academia and industry at home and abroad.
[0003] Currently, mainstream VMR methods mostly focus on the joint modeling of the visual and text modalities. Such methods usually extract the semantic features of video frames and language queries by constructing two-stream networks, cross-modal attention mechanisms, or contrast learning frameworks, and establish the corresponding relationship between the two, so as to achieve accurate positioning of semantic segments. Although such methods have improved the visual semantic alignment effect to a certain extent, they generally ignore the equally important audio modality in the video.
[0004] In fact, the audio modality carries a large amount of semantic clues in the video. For example, conversations, human voices, background sounds, unstructured noises, etc. can all provide key supplementary information for the understanding of query statements. In some scenes with visual blurring or similar actions, audio is even the only basis for distinguishing semantic behaviors. For example, a person laughing and a person talking may be almost the same visually, but they have very different performances in audio. Therefore, in recent years, some studies have begun to try to introduce audio information into the VMR model to improve the semantic understanding ability of the model.
[0005] Representative works include methods such as PMI-LOC, UMT, and ADPN. PMI-LOC introduces RGB, optical flow, and audio three-modal features and designs a cross-modal attention mechanism for fusion; UMT constructs a unified multi-modal Transformer network to integrate visual and audio information; ADPN, from the perspective of semantic consistency and complementarity, proposes a variant of the Transformer architecture to optimize the audio fusion path. These methods have proven the potential of the audio modality in the clip retrieval task to a certain extent and have achieved performance improvements superior to traditional bimodal methods.
[0006] However, there is a key problem commonly existing in the above methods, that is: the importance of audio information is not modeled and perceived, and there is a lack of a selection mechanism based on semantic relevance. Specifically, these methods usually treat audio, visual, and text features equally, and perform joint reasoning through average fusion or simple attention mechanisms, without fully considering the contribution differences of audio under different video and query semantics. In practical applications, audio signals exhibit extremely high diversity and uncertainty: on the one hand, in semantics such as laughter, shouting, and clapping, audio is an indispensable modality; but on the other hand, the audio signals in a large number of videos may be only background music, environmental noise, or even silent segments without semantics. Such information not only fails to help complete the task, but may even introduce interference and mislead the model's judgment.
[0007] A further difficulty lies in that: there is currently no publicly available large-scale audio importance label dataset, and the model cannot obtain supervision signals for the degree of audio semantic contribution during the training phase. This results in the existing methods being unable to learn from the data when to pay attention to audio and when to suppress its interference, thus limiting the effective utilization of the audio modality. This problem has not been fundamentally solved under the existing technical system.
[0008] In summary, although the existing video clip retrieval methods have initially introduced the audio modality, they generally lack the ability to perceive the importance of audio semantics and cannot dynamically adjust the weights and roles of audio features according to different query contents and video scenarios, resulting in obvious performance bottlenecks in dealing with complex and diverse cross-modal semantic reasoning tasks. Therefore, there is an urgent need to propose a retrieval method that can dynamically perceive the importance of audio semantics and adaptively regulate the fusion method of audio and visual modalities according to its contribution, so as to further improve the accuracy and robustness of video clip retrieval. Summary of the Invention
[0009] In view of the above, the purpose of the present invention is to provide a multi-granularity fusion video clip retrieval method based on audio importance perception, aiming to achieve a more flexible and accurate semantic reasoning mechanism between audio, visual, and text modalities, thereby improving the robustness and retrieval accuracy of the model in multi-modal complex scenarios.
[0010] To achieve the above invention purpose, a multi-granularity fusion video clip retrieval method based on audio importance perception provided by an embodiment includes the following steps: Construct a learning framework, which includes an input unit and a segment retrieval prediction unit. The input unit is used to input video and audio segments as well as query text. The segment retrieval prediction unit includes a visual branch, a fusion branch, and an audio branch. The visual branch extracts visual-text fusion features based on video frames and query text and then predicts a first video segment retrieval result. The audio branch extracts audio-text fusion features based on audio segments and query text and then predicts a second video segment retrieval result. After the fusion branch predicts an audio importance score based on the visual-text fusion features and the audio-text fusion features, it performs multi-granularity fusion on the two fusion features based on the audio importance score to obtain a total fusion feature and then predicts a third video segment retrieval result; Construct a retrieval loss based on each video segment retrieval result, construct a pseudo label according to the retrieval losses corresponding to the visual branch and the audio branch, construct an audio importance prediction loss based on the pseudo label and the predicted audio importance score, construct a knowledge distillation loss between the fusion branch and the visual branch and the audio branch respectively, and introduce a saliency contrast loss among the three fusion features; After training the learning framework using all the loss functions and optimizing the parameters of the learning framework, perform video segment retrieval based on at least one branch in the input unit and the segment retrieval prediction unit.
[0011] Optionally, the visual branch includes a visual encoder, a text encoder, a visual-text fusion module, and a first segment location predictor. Extracting visual-text fusion features based on video frames and query text and then predicting a first video segment retrieval result includes: After using the visual encoder and the text encoder to extract visual features and semantic features from the video frames and the query text respectively, use the visual-text fusion module to perform semantic interaction on the visual features and the semantic features, and adopt a context query attention mechanism to extract the context features that are activated by the queried semantics and are the most relevant as the visual-text fusion features. Use the first segment location predictor to predict the start and end positions of the video segment corresponding to the query semantics based on the visual-text fusion features as the first video segment retrieval result.
[0012] Optionally, the audio branch includes an audio encoder, a text encoder, an audio-text fusion module, and a second segment location predictor. Extracting audio-text fusion features based on audio segments and query text and then predicting a second video segment retrieval result includes: After extracting audio features and semantic features from an audio clip and a query text using an audio encoder and a text encoder respectively, an audio-text fusion module is used to perform semantic interaction on the audio features and semantic features, and a context query attention mechanism is adopted to extract the most relevant context features activated by the queried semantics as the audio-text fusion features. A second clip localization predictor is used to predict the start and end positions of the video clip corresponding to the query semantics based on the audio-text fusion features as the second video clip retrieval result.
[0013] Optionally, the fusion branch includes an importance-aware multi-granularity fusion module and a third clip localization predictor. The importance-aware multi-granularity fusion module includes an audio importance predictor and a multi-granularity fusion sub-module. After predicting the audio importance score based on the vision-text fusion features and the audio-text fusion features, a multi-granularity fusion is performed on the two fusion features based on the audio importance score to obtain the total fusion features, and then the third video clip retrieval result is predicted, including: Using the audio importance predictor to predict the audio importance score based on the vision-text fusion features and the audio-text fusion features, using the multi-granularity fusion sub-module to perform weighted summation at the local level, event level, and global level on the vision-text fusion features and the audio-text fusion features respectively to achieve multi-granularity fusion and obtain the total fusion features, and using the third clip localization predictor to predict the start and end positions of the target video clip based on the total fusion features as the third video clip retrieval result.
[0014] Optionally, the audio importance predictor includes a global feature aggregation sub-module and a multi-layer perceptron. Predicting the audio importance score based on the vision-text fusion features and the audio-text fusion features includes: Using the global feature aggregation sub-module to perform global feature aggregation on the vision-text fusion features and the audio-text fusion features respectively using attention pooling to obtain two global semantic features, and then concatenating the two global semantic features and inputting them into the multi-layer perceptron for cross-modal feature interaction and predicting the audio importance score.
[0015] Optionally, performing weighted summation at the local level, event level, and global level on the vision-text fusion features and the audio-text fusion features respectively based on the audio importance score to achieve multi-granularity fusion and obtain the total fusion features includes: For local-level fusion, the vision-text fusion features and the audio-text fusion features respectively pass through a multi-core convolutional neural network to extract their respective multiple local features, and then are concatenated and input into their respective multi-layer perceptrons to obtain locally enhanced video features and audio features. Then, the locally enhanced video features and audio features are weighted element-wise fused using the audio importance score as the weight to obtain the local perception fusion features; For event-level fusion, the visual-text fusion feature and the audio-text fusion feature respectively extract their respective event features through the slot attention sub-module. After the respective event features and the original fusion feature extract features through their respective cross-attention sub-modules, the features output by the cross-attention sub-module are weighted element-wise fused with the audio importance score as the weight to obtain the event-aware fusion feature; For global-level fusion, after the visual-text fusion feature and the audio-text fusion feature respectively perform attention summation operations to obtain their respective global features, each element in the respective global features and the original fusion feature is concatenated and then passed through their respective multi-layer perceptrons to obtain the globally enhanced video feature and audio feature. Then, the globally enhanced video feature and audio feature are weighted element-wise fused with the audio importance score as the weight to obtain the global-aware fusion feature; For multi-granularity fusion, for the local-aware fusion feature, the event-aware fusion feature, and the global-aware fusion feature, a group of Bi-GRUs are introduced to pairwise combine the fusion features at each level to reconstruct the cross-awareness relationship between them. Then, the fusion results are concatenated and passed through a multi-layer perceptron to obtain the total fusion feature.
[0016] Optionally, construct pseudo-labels according to the retrieval losses corresponding to the visual branch and the audio branch, including: ; Among them, and respectively represent the retrieval losses of the audio branch and the visual branch, represents the temperature hyperparameter, represents the initial pseudo-label, represents the output pseudo-label after processing, is the lower threshold, is the upper threshold.
[0017] Optionally, the audio importance prediction loss constructed based on the pseudo-label and the predicted audio importance score uses binary cross-entropy loss.
[0018] Optionally, the knowledge distillation loss constructed between the fusion branch and the visual branch and the audio branch respectively uses the KL divergence between the third video segment retrieval result of the fusion branch and the first video segment retrieval result and the second video segment retrieval result of the visual branch and the audio branch; Introduce a significance contrast loss inside and outside the true value interval for the three fusion features, denoted as : ; Among them, is the feature sequence after the fusion feature is compressed in dimension to 1 by the linear layer, Represents a randomly selected feature within the true value interval, Represents a randomly selected feature outside the true value interval, loss Forced feature And the feature Need to pull apart a Distance.
[0019] Compared with the prior art, the beneficial effects of the present invention at least include: Based on obtaining visual-text fusion features and audio-text fusion features, after predicting the audio importance score, the present invention performs multi-granularity fusion on the two fusion features based on the importance score to obtain the total fusion feature, and then predicts the retrieval result of the third video segment. During training, retrieval loss, audio importance prediction loss, knowledge distillation loss between branches, and saliency contrast loss between fusion features are introduced. Such training enables each branch after training to significantly improve the retrieval accuracy, adapt to the uncertainty of the audio modality, enhance the robustness of the machine clue, and enhance the single-modal performance, thereby improving the practicality of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 Is a flowchart of the multi-granularity fusion video segment retrieval method based on audio importance perception provided by the embodiment; Figure 2 Is a structural and training schematic diagram of the learning framework provided by the embodiment; Figure 3 Is a structural schematic diagram of the multi-granularity fusion sub-module provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0023] The inventive concept of the present invention is as follows: Aiming at the following problems commonly existing in the existing video clip retrieval technology: First, the existing methods mostly focus on the joint modeling of visual and text modalities. Although cross-modal semantic alignment is achieved to a certain extent, the rich semantic information contained in the audio modality in the video is generally ignored. Especially in scenes where it is difficult to distinguish visually, such as speaking and laughing, the audio often plays a key role. The lack of full utilization of the audio modality will lead to misjudgment or fuzzy recognition of the model in semantic understanding and segment localization. Second, although some methods have tried to introduce the audio modality for trimodal fusion, there is generally a problem of treating multimodal features equally, and it is not possible to dynamically judge according to the semantic value of the audio information in different scenes. Due to the lack of audio importance supervision labels available for training, it is difficult for the existing technology to effectively identify whether the audio has key semantics in model design, resulting in limited performance improvement and even introducing noise interference in some scenes.
[0024] The embodiment of the present invention provides a video clip retrieval solution that can perceive the semantic importance of audio and achieve dynamic regulation of the fusion of audio and visual information, aiming to realize a more flexible and accurate effective joint inference mechanism for semantics among audio, visual, and text modalities, so as to improve the robustness and accuracy of the model in video clip retrieval in multimodal complex scenes.
[0025] A multi-granularity fusion video clip retrieval method based on audio importance perception provided by the embodiment aims to retrieve the start and end frame pairs { , } of a specific clip that semantically matches the natural language query Q from the uncropped video V. In addition, for each frame in the video, the corresponding audio clip can be extracted as a supplementary modality to provide context information to enhance the retrieval effect. As Figure 1 shown, it includes the following steps: S1, construct a learning framework, which includes an input unit and a clip retrieval prediction unit. The input unit is used to input video, audio clips, and query text. The clip retrieval prediction unit includes a visual branch, a fusion branch, and an audio branch, all of which are used to retrieve video clips to obtain video clip retrieval results.
[0026] In the embodiment, the learning framework is as Figure 2 shown. The input unit is used to input video, audio clips, and query text. Among them, video frames are captured from the video as the input of the visual modality, the query text is input in natural language as the input of the text modality, and the audio clip is used as the input of the audio modality.
[0027] In the embodiment, the visual branch included in the clip retrieval prediction unit extracts visual-text fusion features based on the video frame and the query text and then predicts the first video clip retrieval result. As Figure 2As shown in the figure, the visual branch includes a visual encoder, a text encoder, a visual-text fusion module, and a first segment localization predictor. The process of predicting the first video segment retrieval result is as follows: After extracting visual features and semantic features from video frames and query texts using the visual encoder and the text encoder respectively, the visual-text fusion module performs semantic interaction on the visual features and semantic features, and uses the context query attention mechanism to extract the most relevant context features activated by the queried semantics as the visual-text fusion features. The first segment localization predictor predicts the start and end positions of the video segment corresponding to the query semantics based on the visual-text fusion features as the first video segment retrieval result.
[0028] In the embodiment, the audio branch included in the segment retrieval prediction unit predicts the second video segment retrieval result after extracting audio-text fusion features based on the audio segment and the query text. As Figure 2 shown in the figure, the audio branch includes an audio encoder, a text encoder, an audio-text fusion module, and a second segment localization predictor. The process of predicting the second video segment retrieval result is as follows: After extracting audio features and semantic features from the audio segment and the query text using the audio encoder and the text encoder respectively, the audio-text fusion module performs semantic interaction on the audio features and semantic features, and uses the context query attention mechanism to extract the most relevant context features activated by the queried semantics as the audio-text fusion features. The second segment localization predictor predicts the start and end positions of the video segment corresponding to the query semantics based on the audio-text fusion features as the second video segment retrieval result.
[0029] Specifically, for the visual modality, a pre-trained visual CNN model is used to extract the original visual features of the video frames , and is enhanced by a visual encoder composed of a feed-forward network (FFN), a convolutional layer, and a Transformer layer to obtain visual features . For the audio modality, a pre-trained audio perception CNN model is used to extract the original audio features of the audio segment , and is enhanced by an audio encoder with the same structure as the visual encoder to obtain audio features . For the text modality, the query text is directly initialized using GloVe word vectors. Since the query text may have different alignment methods semantically with the visual modality and the audio modality, two independent text encoders (with the same structure as the audio encoder) are further used to encode the query text respectively, so as to obtain modality-specific enhanced semantic features, which are and . To highlight the key semantic alignment between the visual modality and the given query text, in the visual-text fusion module, the visual features and the semantic features Apply the context-query attention mechanism to obtain visual-text fusion features Meanwhile, in the audio-text fusion module, for the audio features and semantic features apply the context-query attention mechanism to obtain audio-text fusion features .
[0030] In the embodiment, after the fusion branch included in the segment retrieval prediction unit predicts the audio importance score based on the visual-text fusion features and the audio-text fusion features, a multi-granularity fusion of the two fusion features is performed based on the importance score to obtain the total fusion feature, and then the third video segment retrieval result is predicted. As Figure 2 shown, the fusion branch includes an importance-aware multi-granularity fusion module and a third segment location predictor. Among them, the importance-aware multi-granularity fusion module includes an audio importance predictor and a multi-granularity fusion sub-module. The importance-aware multi-granularity fusion module is a multi-granularity fusion structure guided by the audio importance predictor. This predictor is trained to identify and emphasize semantically relevant audio cues, enabling the fusion module to selectively fuse effective audio signals with visual features. Under the guidance of the predicted audio importance score, the multi-granularity fusion process can effectively filter out irrelevant or noisy audio content, while aggregating meaningful cross-modal information at multiple time scales, thereby improving the retrieval performance.
[0031] Specifically, the process of the fusion branch predicting the third video segment retrieval result includes: using the audio importance predictor to predict the audio importance score based on the visual-text fusion features and the audio-text fusion features, and using the multi-granularity fusion sub-module to perform weighted summation of the visual-text fusion features and the audio-text fusion features at the local level (video frame-audio segment alignment), event level (semantic event aggregation), and global level (overall semantic abstraction) respectively based on the audio importance score, fully capturing the complementary information between audio and vision. The features at each level are reconstructed and integrated into a unified multi-modal fusion representation by a Bi-GRU to achieve multi-granularity fusion to obtain the total fusion feature, and using the third segment location predictor to predict the start and end positions of the target video segment based on the total fusion feature as the third video segment retrieval result.
[0032] Among them, the audio importance predictor adopts a lightweight design and is used to dynamically estimate the importance of the audio in each video-query pair. Specifically, the audio importance predictor includes a global feature aggregation sub-module and a multi-layer perceptron, and predicts the audio importance score based on the visual-text fusion features and the audio-text fusion features, including: using the global feature aggregation sub-module to perform attention pooling operations on the visual-text fusion features And audio-text fusion features Global feature aggregation is performed to obtain two global semantic features And After that, the two global semantic features And Are concatenated and input into a multi-layer perceptron (MLP) for cross-modal feature interaction, enabling it to reason about the relative importance of audio based on visual context and predict an audio importance score through activation functions such as sigmoid This audio importance score 𝑝 is used to guide the subsequent multi-modal fusion process, achieving dynamic fusion control by adjusting the contribution of the audio modality
[0033] Given that the audio modality itself has stronger noise and variability compared to visual signals, a simple fusion strategy may not be sufficient to fully exploit the complementarity between audio and visual. Therefore, the present invention proposes a multi-granularity fusion sub-module that performs hierarchical fusion from three perspectives: local level, event level, and global level, and is guided by the dynamically predicted audio importance score for weighted fusion to achieve multi-granularity fusion and obtain the total fusion feature
[0034] For local-level fusion, as shown in (a) of Figure 3 To achieve frame-by-frame alignment and fine-grained fusion of visual frames and audio segments, a symmetric multi-core one-dimensional convolutional network is constructed to more deeply perceive the local relationship between video frames and audio segments. Specifically, the visual-text fusion features And audio-text fusion features Respectively pass through the multi-core convolutional neural network to extract their respective multiple local features And Then they are concatenated and input into their respective multi-layer perceptrons (MLPs) to obtain the locally enhanced video features And audio features And then, through the audio importance score p As the weight, the locally enhanced video features And audio features Are weighted element-wise fused to obtain the local perception fusion feature ; ; ; ; Among them, Represents the kernel size of the convolutional network, Represents the convolutional kernel The corresponding output feature dimension, Conv1 represents one-dimensional convolution, nIndicates the total number of convolutional kernels, LN Indicates layer normalization.
[0035] For event-level fusion, in order to capture the event-semantic matching relationship between vision and audio for activity understanding, as Figure 3 shown in (b) of [reference], first, the vision-text fusion feature and the audio-text fusion feature respectively pass through the slot attention sub-module. The slot attention mechanism with a group of learnable event slots aggregates similar visual / audio segments into multiple events and extracts their respective event features and . Their respective event features and and the original fusion features and pass through their respective cross-attention sub-modules to extract features and . After that, through the audio importance score p as the weight, the features and output by the cross-attention sub-module are fused element-wise with weights to obtain the event-aware fusion feature : ; Among them, SlotAttn represents the slot attention mechanism. In the cross-attention sub-module, the original feature / feature serves as the query, while the extracted event features and serve as the key and value. When performing weighted element-wise fusion on the features and , the same method as local-level fusion is adopted.
[0036] For global-level fusion, in order to match the vision and audio context from a global perspective, as Figure 3 shown in (c) of [reference], the vision-text fusion feature and the audio-text fusion feature respectively obtain their respective global features after the attention pooling operation. Then, each element in their respective global features and the original fusion features is concatenated and then passed through their respective multi-layer perceptrons (MLPs) to obtain the globally enhanced video feature and the audio feature . Then, through the audio importance score p as the weight, the globally enhanced video feature and audio features perform weighted element-wise fusion to obtain global perception fusion features . Here, the weighted element-wise fusion still adopts the same method as the local-level fusion.
[0037] For multi-granularity fusion, since there are different perceptual correlations between the fusion features from different granularity levels, for the local perception fusion features , event perception fusion features , and global perception fusion features , a group of Bi-GRUs (Bidirectional Gated Recurrent Units) are introduced to pairwise combine the fusion features at each level to reconstruct the cross-perception relationship between them, and then the fusion results are concatenated and mapped to the original feature space dimension d through a multi-layer perceptron to obtain the total fusion features F .
[0038] In the embodiment, the first segment localization predictor adopted by the visual branch, the second segment localization predictor adopted by the visual branch, and the third segment localization predictor adopted by the fusion branch all adopt the same structure, which is composed of a convolutional layer, a Transformer layer, and a linear layer. In each segment localization predictor, logits for predicting the start position and end position of the video segment are respectively based on the input features, so as to accurately locate the target segment in the video. For the first segment localization predictor, its logits for predicting the start position and end position of the video segment are based on the visual-text fusion features and , and then the subscript of the maximum value is selected through softmax as the finally predicted start position or end position and used as the first video segment retrieval result; for the second segment localization predictor, its logits for predicting the start position and end position of the video segment are based on the audio-text fusion features and , and similarly, the subscript of the maximum value is selected through softmax as the finally predicted start position or end position and used as the second video segment retrieval result; for the third segment localization predictor, its logits for predicting the start position and end position of the video segment F are based on the total fusion features and , and similarly, the subscript of the maximum value is selected through softmax as the finally predicted start position or end position and used as the third video segment retrieval result.
[0039] S2. Based on the retrieval results of each video segment, construct a retrieval loss, construct pseudo-labels according to the retrieval losses corresponding to the visual branch and the audio branch, construct an audio importance prediction loss based on the pseudo-labels and the predicted audio importance scores, construct a knowledge distillation loss between the fusion branch and the visual branch and the audio branch respectively, and introduce a significance contrast loss for the three fusion features both within and outside the true value interval.
[0040] In the embodiment, a retrieval loss is constructed for the retrieval results of each video segment of each branch as the core loss to ensure that each branch has the video segment retrieval ability. Specifically, each retrieval loss is the cross-entropy between the predicted start position of the video segment in the video segment retrieval result and the true start position, and the cross-entropy between the predicted end position of the video segment and the true end position. Taking the retrieval loss of the fusion branch as an example, the specific retrieval loss is expressed as: ; where CE represents the cross-entropy loss function, and are 0-1 sequences provided during training, which are 1 only at the start / end positions, representing the true positions.
[0041] In the embodiment, since the audio importance predictor lacks true importance annotations, a loss-aware pseudo-importance label generator is designed to construct pseudo-labels as supervision signals during training. Inspired by the observation that neural networks tend to preferentially learn from "easy" samples during training, and such samples usually exhibit smaller training losses. Based on this, compare the retrieval losses of each video-query pair under the audio branch and the visual branch. The modality with the smaller loss is considered to provide more relevant information, so a higher pseudo-importance score should be assigned. Then, the pseudo-labels constructed based on the retrieval losses corresponding to the visual branch and the audio branch include: ; where and represent the retrieval losses of the audio branch and the visual branch respectively, represents the temperature hyperparameter, represents the pseudo-importance score of the audio modality, represents the processed output pseudo-importance score and serves as the pseudo-label, is the lower threshold, is the upper threshold. When is lower than the threshold , the audio is regarded as an uninformative modality and its contribution will be suppressed. On the contrary, if is higher than the threshold , it indicates that audio plays a dominant role in the retrieval.
[0042] Then, based on the pseudo-labels and the predicted audio importance scores construct the audio importance prediction loss , and adopt binary cross-entropy loss: ; where represents the batch size, represents the index of the -th sample. The audio importance score serves as a key control parameter in the subsequent multi-granularity fusion stage, guiding the selective fusion of audio and visual features. To prevent unstable prediction results from misleading the fusion process at the initial stage of training, the initial fusion weight is set to the neutral value 0.5, and the influence of the audio importance score is gradually increased as the training progresses. This "curriculum-like" strategy helps to build a robust multi-modal interaction at the initial stage while mitigating the impact of noise in the early importance estimation.
[0043] In the embodiment, the fusion branch naturally captures richer and more comprehensive semantic representations by jointly modeling audio and visual cues. However, in practical applications, the audio signal may be missing, damaged, or unavailable during the inference stage. To ensure that the single-modal branches, especially the visual branch, can still maintain a strong retrieval ability under such conditions, a cross-modal knowledge distillation strategy is introduced to transfer the joint semantic knowledge learned in the fusion branch to the single-modal branches, enhancing their ability to work independently during the inference stage, especially to maintain good performance in scenarios where audio is missing. Specifically, we regard the fusion branch as the teacher network, and distill its knowledge into the student networks, namely the visual branch and the audio branch, especially the visual branch, so that it can inherit the modality complementary information contained in the fusion branch and achieve good retrieval results even with only visual input. To this end, minimize the Kullback-Leibler (KL) divergence between the retrieval results of the third video segment of the fusion branch and the retrieval results of the first and second video segments output by the visual branch and the audio branch to construct the knowledge distillation loss : ; where and represent the logits of the start and end positions in the video segment retrieval results predicted by the student network (visual branch or audio branch), and The logits of the start position and the end position in the video segment retrieval result predicted by the teacher network (fusion branch). is the temperature coefficient. is the softmax function. Integrating the distillation processes of the visual branch and the audio branch, the final knowledge distillation loss for the visual branch and the audio branch is the sum of them.
[0044] In the embodiment, significance contrast losses within and outside the ground truth interval are respectively introduced for three fusion features , and F This loss enhances the model's attention to key information by enlarging the feature gap between inside and outside the ground truth interval, and is specifically expressed as :
[0045] Among them, is the feature sequence after the fusion feature (being , or F ) is compressed in dimension to 1 by a linear layer. represents a randomly selected feature within the ground truth interval, represents a randomly selected feature outside the ground truth interval, and the loss forces the feature and the feature to have a distance of .
[0046] Then the total loss function of the entire learning framework is:
[0047] Among them, , and are the balance coefficients of each loss term.
[0048] S3. After training the learning framework using all loss functions and optimizing the parameters of the learning framework, video segment retrieval is performed based on at least one branch of the input unit and the segment retrieval prediction unit.
[0049] In the embodiment, the learning framework is trained using the total loss function and the parameters of the learning framework are optimized. After the optimization is completed, video segment retrieval is performed based on at least one branch of the input unit and the segment retrieval prediction unit, that is, the system can select the fusion branch or the unimodal branch for retrieval according to requirements, and has good flexibility and generalization ability, and the visual-audio fusion branch is usually the first choice.
[0050] Through the above technical solutions, the present invention realizes the automatic perception and dynamic weighting of the contribution of the audio modality in the video, effectively solves the problem of insufficient utilization of audio information or even introduction of noise in the existing methods, and improves the robustness and accuracy of video clip retrieval in multi-modal complex scenarios. Specifically, the technical effects demonstrated by experiments are as follows: 1. Significantly improve the retrieval accuracy By introducing an audio importance prediction module, the present invention effectively avoids the interference of low-quality audio on the model performance, and enhances the complementarity between different modalities through a multi-granularity semantic fusion strategy. According to the experimental results on two public benchmark datasets, Charades-STA and ActivityNet Captions, the proposed method outperforms existing mainstream methods in multiple evaluation metrics. For the Charades-STA dataset, the indicators of R1@5 / R1@7 / mIOU are improved by 0.87% / 9.76% / 5.33% respectively, while for the ActivityNet Captions dataset, the indicators of R1@5 / R1@7 / mIOU are improved by 8.84% / 12.01% / 6.81% respectively. Compared with not introducing audio, the present invention improves the indicators of R1@5 / R1@7 / mIOU on the Charades-STA dataset by 9.78% / 11.92% / 5.42% respectively, and on the ActivityNet Captions dataset, the indicators of R1@5 / R1@7 / mIOU are improved by 8.84% / 8.55% / 4.58% respectively. Among them, R1@5 represents the ratio of the test samples whose IOU (Intersection over Union) between the first-ranked interval returned by the retrieval and the ground-truth interval is greater than 0.5 to all test samples, while R1@7 represents the ratio of the test samples whose IOU is greater than 0.7 to all test samples, and mIOU represents the average IOU of all test samples.
[0051] 2. Adapt to the uncertainty of the audio modality and improve the model robustness The audio importance prediction mechanism proposed by the present invention can dynamically evaluate the actual contribution of the audio modality in different scenarios, so as to automatically adjust the degree of audio participation in the inference stage. Experiments show that in the face of scenarios with strong noise interference or audio information loss, the present invention can still maintain a high retrieval accuracy. When the audio noise level increases, the performance of the model using the audio importance perception module drops much less than that of the version without this module, verifying the robustness and practical adaptability of the system.
[0052] 3. Achieve modular design, with good compatibility and scalability The audio importance predictor and multi-granularity fusion sub-module proposed by the present invention have good pluggability and can be integrated as general modules into other multi-modal retrieval frameworks. After being combined with existing methods such as EMB and EAMAT, their performance has been significantly improved, indicating that the present invention has good generality and transferability.
[0053] 4. Enhance the performance of single modality and improve the practicality of the system To address the situation of missing or unavailable audio in real applications, the present invention introduces a cross-modal knowledge distillation mechanism to transfer the joint knowledge in the fusion branch to the visual and audio branches, significantly improving the independent retrieval ability of the visual branch under the condition of no audio input. Experiments have shown that when only the visual modality is used in the inference stage, the decline of the method of the present invention is very small. The R1@5 index only drops by 1.46%, the R1@7 index only drops by 1.79%, and the mIoU index drops by 1.17%, demonstrating excellent deployment flexibility and practical value.
[0054] In summary, the present invention not only achieves a performance breakthrough in technical indicators, but also enhances the adaptability and expansion ability of the system through modular design and general fusion mechanism. At the same time, it demonstrates excellent robustness and applicability in actual scenarios, and has high application value and promotion prospects.
[0055] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-granularity fusion video clip retrieval method based on audio importance perception, characterized in that The steps include: Construct a learning framework, which includes an input unit and a segment retrieval prediction unit. The input unit is used to input video, audio segments, and query text. The segment retrieval prediction unit includes a visual branch, a fusion branch, and an audio branch. The visual branch extracts visual-text fusion features based on video frames and query text and then predicts the first video segment retrieval result. The audio branch extracts audio-text fusion features based on audio segments and query text and then predicts the second video segment retrieval result. After the fusion branch predicts the audio importance score based on the visual-text fusion features and the audio-text fusion features, it performs multi-granularity fusion on the two fusion features based on the audio importance score to obtain the total fusion feature and then predicts the third video segment retrieval result; Construct a retrieval loss based on each video segment retrieval result, construct pseudo-labels according to the retrieval losses corresponding to the visual branch and the audio branch, construct an audio importance prediction loss based on the pseudo-labels and the predicted audio importance score, construct a knowledge distillation loss between the fusion branch and the visual branch and the audio branch respectively, and at the same time introduce a significance contrast loss inside and outside the true value interval for the three fusion features; After training the learning framework using all loss functions and optimizing the parameters of the learning framework, perform video segment retrieval based on at least one branch in the input unit and the segment retrieval prediction unit.
2. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 1, wherein, The visual branch includes a visual encoder, a text encoder, a visual-text fusion module, and a first segment location predictor. Extracting visual-text fusion features based on video frames and query text and then predicting the first video segment retrieval result includes: After using the visual encoder and the text encoder to extract visual features and semantic features from the video frames and query text respectively, use the visual-text fusion module to perform semantic interaction on the visual features and semantic features, and adopt a context query attention mechanism to extract the context features activated by the queried semantics and most relevant as the visual-text fusion features. Use the first segment location predictor to predict the start and end positions of the video segment corresponding to the query semantics based on the visual-text fusion features as the first video segment retrieval result.
3. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 1, wherein The audio branch includes an audio encoder, a text encoder, an audio-text fusion module, and a second segment location predictor. Extracting audio-text fusion features based on audio segments and query text and then predicting the second video segment retrieval result includes: After using the audio encoder and the text encoder to extract audio features and semantic features from the audio segments and query text respectively, use the audio-text fusion module to perform semantic interaction on the audio features and semantic features, and adopt a context query attention mechanism to extract the context features activated by the queried semantics and most relevant as the audio-text fusion features. Use the second segment location predictor to predict the start and end positions of the video segment corresponding to the query semantics based on the audio-text fusion features as the second video segment retrieval result.
4. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 1, characterized in that The fusion branch includes an importance-aware multi-granularity fusion module and a third segment location predictor. Among them, the importance-aware multi-granularity fusion module includes an audio importance predictor and a multi-granularity fusion sub-module. After predicting the audio importance score based on the visual-text fusion feature and the audio-text fusion feature, the two fusion features are multi-granularity fused based on the audio importance score to obtain the total fusion feature, and then the retrieval result of the third video segment is predicted, including: Using the audio importance predictor to predict the audio importance score based on the visual-text fusion feature and the audio-text fusion feature, using the multi-granularity fusion sub-module to perform weighted summation at the local level, event level, and global level on the visual-text fusion feature and the audio-text fusion feature respectively to achieve multi-granularity fusion to obtain the total fusion feature, and using the third segment location predictor to predict the start and end positions of the target video segment as the retrieval result of the third video segment.
5. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 4, characterized in that, The audio importance predictor includes a global feature aggregation sub-module and a multi-layer perceptron, and predicts the audio importance score based on the visual-text fusion feature and the audio-text fusion feature, including: Using the global feature aggregation sub-module to perform global feature aggregation on the visual-text fusion feature and the audio-text fusion feature respectively through attention pooling operation to obtain two global semantic features, and then splicing the two global semantic features and inputting them into the multi-layer perceptron for cross-modal feature interaction and predicting the audio importance score.
6. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 4, wherein Performing weighted summation at the local level, event level, and global level on the visual-text fusion feature and the audio-text fusion feature respectively based on the audio importance score to achieve multi-granularity fusion to obtain the total fusion feature, including: For local-level fusion, the visual-text fusion feature and the audio-text fusion feature are respectively passed through a multi-core convolutional neural network to extract their respective multiple local features, and then spliced and input into their respective multi-layer perceptrons to obtain locally enhanced video features and audio features, and then the locally enhanced video features and audio features are weighted element-wise fused using the audio importance score as the weight to obtain the local perception fusion feature; For event-level fusion, the visual-text fusion feature and the audio-text fusion feature are respectively passed through a slot attention sub-module to extract their respective event features, and after the respective event features and the original fusion features are passed through their respective cross-attention sub-modules to extract features, the features output by the cross-attention sub-module are weighted element-wise fused using the audio importance score as the weight to obtain the event perception fusion feature; For global-level fusion, the visual-text fusion feature and the audio-text fusion feature are respectively passed through an attention pooling operation to obtain their respective global features, and then each element in their respective global features and the original fusion features are spliced and then passed through their respective multi-layer perceptrons to obtain globally enhanced video features and audio features, and then the globally enhanced video features and audio features are weighted element-wise fused using the audio importance score as the weight to obtain the global perception fusion feature; For multi-granularity fusion, for the local perception fusion feature, event perception fusion feature, and global perception fusion feature, a group of Bi-GRUs is introduced to pairwise combine the fusion features at each level to reconstruct the cross-perception relationship between them, and then the fusion results are concatenated and passed through a multi-layer perceptron to obtain the total fusion feature.
7. The method for retrieving multi-granularity fusion video segments based on audio importance perception according to claim 1, wherein For each branch, the retrieval loss is constructed based on the retrieval results of each video segment as follows: the cross-entropy between the predicted start position of the video segment in the video segment retrieval result and the true start position, and the cross-entropy between the predicted end position of the video segment and the true end position.
8. The method for retrieving multi-granularity fusion video segments based on audio importance perception according to claim 1, wherein Pseudo-labels are constructed according to the retrieval losses corresponding to the visual branch and the audio branch, including: ; Among them, and respectively represent the retrieval losses of the audio branch and the visual branch, represents the temperature hyperparameter, represents the initial pseudo-label, represents the pseudo-label output after processing, is the lower threshold, is the upper threshold.
9. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 1, characterized in that The audio importance prediction loss constructed based on the pseudo-labels and the predicted audio importance scores, using binary cross-entropy loss.
10. The multi-granularity fusion video clip retrieval method based on audio importance perception according to claim 1, wherein, The knowledge distillation loss constructed between the fusion branch and the visual branch and the audio branch respectively, using the KL divergence between the third video segment retrieval result of the fusion branch and the first and second video segment retrieval results of the visual branch and the audio branch respectively; Introduce the significance contrast loss inside and outside the true value interval for the three fusion features, denoted as : ; Among them, is the feature sequence after the fused features are compressed in dimension to 1 by a linear layer, represents a randomly selected feature within the true value interval, represents a randomly selected feature outside the true value interval, and the loss forces the feature and the feature to be separated by a distance.
Citation Information
Patent Citations
Video classification method based on knowledge distillation and multi-modal fusion
CN115147641A
Video clip retrieval method based on fine-grained modal relationship sensing network
CN118520140A
Training method of video retrieval model and video retrieval method and device
CN119622029A
Multi-modal sentiment analysis method and system based on knowledge distillation and dynamic fusion mechanism
CN120046695A
Audio-visual fusion with cross-modal attention for video action recognition
WO2021184026A1
Cited By
A multi-granularity zentropy-based visual saliency prediction method and system
CN122530753A
A multi-granularity zentropy-based visual saliency prediction method and system
CN122530753B