Weak supervision combined moment retrieval method for multi-granularity semantics and dynamic Gaussian modeling
Through the multi-grained semantics and dynamic Gaussian modeling methods, the problem of insufficient diversified query text expression and action modeling in weak-supervised moment retrieval is solved, and efficient and accurate video moment retrieval effect is achieved.
Patent Information
- Application Number
- CN202510365652.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The existing weak supervision moment search method is difficult to model compound semantic expressions and dynamic action patterns, and cannot effectively deal with the insufficient modeling of diversified query text expressions and proposal internal frame weights.
The multi-grained semantics and dynamic Gaussian modeling method is adopted to fuse visual and text features through the slot attention mechanism, dynamically adjust the frame weight distribution, combine mask reconstruction to select the optimal proposal, and perform multimodal feature fusion and dynamic proposal integration.
It significantly improves the accuracy and robustness of weak supervision moment search under complex text expression, can better adapt to the needs of diversified query texts, and achieve efficient and accurate video moment search.
Smart Images

Figure CN120277235A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision, image, and video processing, and particularly to a weakly supervised combined moment retrieval method with multi-granularity semantics and dynamic Gaussian modeling. Background Art
[0002] With the rapid development of mobile Internet and multimedia technologies, video data has shown an exponential growth trend. Against this background, the moment retrieval task, that is, accurately locating specific segments in a video according to a text query, has become a research hotspot in the field of cross-modal understanding. Although fully supervised methods have made significant progress in this task, their characteristic of relying on accurate timestamp annotations faces two major practical dilemmas:
[0003] 1) The cost of manual annotation is high and there are subjective biases. The determination differences of time boundaries for the same semantics by different annotators can reach 15%-20%;
[0004] 2) The singularity of text expressions in the annotated data makes it difficult for the model to adapt to the language diversity in real scenarios.
[0005] Therefore, the weakly supervised moment retrieval paradigm has emerged. It only requires video-text pair-level annotations to complete model training, significantly improving the practicality of the method.
[0006] In recent years, weakly supervised moment retrieval technologies have continued to evolve along two main directions:
[0007] In terms of cross-modal alignment, the CNM framework proposed by Zheng et al. in the literature "Weakly supervised video moment localization with contrastive negative sample mining" alleviates the problem of local semantic confusion through a contrastive negative sampling strategy. Its improved CPL method (Weakly Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning) further introduces Gaussian correlation learning inside the proposal, and combines a learnable boundary prediction module to improve the prediction accuracy;
[0008] At the feature modeling level, the Lv team found that the traditional mask reconstruction mechanism has the problem of spurious correlation in the literature "Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment Localization", and proposed a counterfactual reasoning method to effectively reduce the error matching rate; Zhou et al. in the literature "Query-aware multi-scale proposal network for weakly supervised temporal sentence grounding in videos" aimed at the problem of insufficient proposal diversity, designed a multi-scale mapping network and implemented text-aware proposal frame weight reconstruction modeling. It is worth noting that Kim et al. proposed the PPS method in the literature "Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding", and for the first time tried to use a mixture of Gaussian distributions to replace the single distribution to model the frame weights. However, due to the symmetry assumption of the distribution and the manual control of the multi-modal combination method, it is still difficult to accurately depict the action evolution process.
[0009] Despite the significant progress made in the above work, two major challenges in practical applications remain unsolved:
[0010] First, existing methods rely on a single global feature map to generate candidate segments, making it difficult to model composite semantic expressions (such as "a player breaks through and then shoots and scores" containing a temporal action chain). Taking the CPL method as an example, when faced with synonymous words not seen in the training set, the model retrieval metrics Recall@1, IoU = 0.5 drop sharply, with a decline of 10.59%.
[0011] Second, there is an essential conflict between the static frame weight assumption and the dynamic action pattern.
[0012] 1. Existing weakly supervised methods have significant limitations in modeling diverse query text expressions. Existing technologies generally adopt a single global feature map paradigm, generating an overall representation of video segments through cross-modal fusion, but failing to effectively decouple the multi-level structure in text semantics;
[0013] 2. There is an essential contradiction between the existing proposal internal frame weight reconstruction mechanism and the dynamic action pattern. Although traditional methods (such as PPS) use Gaussian distributions to model the importance of frames within segments, the combination method of its multi-modal distribution is limited by the manually designed weight assignment strategy and lacks the ability to model cross-frame semantic coherence. Summary of the Invention
[0014] This application provides a weakly supervised combined moment retrieval method based on multi-granularity semantics and dynamic Gaussian modeling, which can solve the technical problems that existing weakly supervised methods are difficult to handle diverse query text expressions and the internal core action modeling of proposals is insufficient.
[0015] In a first aspect, this application provides a weakly supervised combined moment retrieval method based on multi-granularity semantics and dynamic Gaussian modeling, including the following steps:
[0016] Extract visual features from the video frame sequence and extract text features from diverse query texts;
[0017] Based on the multi-granularity semantic modeling method of the slot attention mechanism, perform multi-modal feature fusion on the extracted visual features and text features, generate proposals with global and local text semantic injection, and perform dynamic proposal integration to generate a diverse set of proposals;
[0018] Based on the dynamic Gaussian modeling method, dynamically adjust the frame weight distribution within the diverse set of proposals to generate a diverse set of proposals with dynamically adjusted frame weight distribution;
[0019] Based on the proposal selection mechanism of mask reconstruction, select the optimal proposal from the diverse set of proposals with dynamically adjusted frame weight distribution to obtain the moment retrieval result;
[0020] Using the extracted visual features and text features as model inputs, the training process converges through contrastive learning based on the reconstruction loss of the cross-entropy of positive proposals and the reconstruction loss of the frame weights of negative proposals. During inference, the positive proposal with the minimum reconstruction loss is used as the combined moment retrieval result, and the parameters of the entire network are optimized through backpropagation to construct a weakly supervised combined moment retrieval model, and based on this model, perform combined moment retrieval of diverse query texts on video data.
[0021] Combined with the first aspect, in an implementation, the extraction of visual features from the video frame sequence and the extraction of text features from diverse query texts specifically include the following steps:
[0022] Use a pre-trained I3D network to extract visual features from the video frame sequence;
[0023] Use the GLoVe framework to extract text features from diverse query texts;
[0024] Map the extracted video features and text features to the same hidden dimension through linear mapping to obtain video features and text features in the same hidden dimension.
[0025] In combination with the first aspect, in one implementation, the multi-granularity semantic modeling method based on the slot attention mechanism performs multi-modal feature fusion on the extracted visual features and text features, generates proposals for global and local text semantic injection, and dynamically integrates the proposals to obtain a diverse set of proposals. Specifically, it includes the following steps:
[0026] Interactively fuse video features and text features through a Transformer decoder to generate multi-modal features;
[0027] Extract global semantic information from the multi-modal features through a learnable global vector and generate a set of proposals for global text semantic injection using linear mapping to obtain a global set of proposals;
[0028] Introduce the Hungarian matching algorithm and a learnable fusion coefficient to dynamically integrate the global set of proposals and the local set of proposals to generate a diverse set of proposals.
[0029] In combination with the first aspect, in one implementation, the step of extracting global semantic information from the multi-modal features through a learnable global vector and generating a global set of proposals for global text semantic injection specifically includes the following steps;
[0030] Extract global semantic information from the multi-modal features through a learnable global vector and generate a set of proposals for global text semantic injection using linear mapping to obtain a global set of proposals;
[0031] Iteratively decode the multi-modal features with local text semantic injection through the slot attention mechanism to generate a set of proposals with local text semantic injection and obtain a local set of proposals.
[0032] In combination with the first aspect, in one implementation, the step of introducing the Hungarian matching algorithm and a learnable fusion coefficient to dynamically integrate the global set of proposals and the local set of proposals to generate a diverse set of proposals specifically includes the following steps:
[0033] Calculate the Euclidean distance matrix of the global set of proposals and the local set of proposals;
[0034] Find the optimal bijective mapping in the Euclidean distance matrix through the Hungarian algorithm to obtain the optimal matching relationship between the global set of proposals and the local set of proposals;
[0035] Perform weighted fusion on the generation parameters and local parameters of the optimal matching relationship to obtain a diverse set of proposals.
[0036] In combination with the first aspect, in one implementation, the step of dynamically adjusting the internal frame weight distribution of the diverse set of proposals based on the dynamic Gaussian modeling method to obtain the diverse set of proposals with dynamically adjusted frame weight distribution specifically includes the following steps:
[0037] Based on the timestamps of the proposals, use the traditional Gaussian distribution to perform proposal frame weight mapping on the diverse set of proposals to obtain the proposal frame weight distribution;
[0038] Introduce a learnable width ratio parameter to dynamically adjust the proposal frame weight distribution and generate a diverse proposal set with dynamically adjusted weights.
[0039] Combined with the first aspect, in one implementation, the proposal selection mechanism based on mask reconstruction selects the optimal proposal from the diverse proposal set with dynamically adjusted frame weight distribution to obtain the moment retrieval result, specifically including the following steps:
[0040] Mask out some words in the diverse query text to generate a masked text;
[0041] Use a learnable action-aware Gaussian distribution module to aggregate the proposal features of the diverse proposal set with dynamically adjusted frame weights, and use the aggregated proposal features to restore the masked text to obtain a restored text;
[0042] Evaluate the quality of the selected proposals according to the matching degree between the restored text and the original text, and select the optimal proposal from the diverse proposal set as the retrieval result.
[0043] Combined with the first aspect, in one implementation, the restoration loss of the cross-entropy is shown as follows:
[0044]
[0045] In the formula, P p is the probability distribution according to the vocabulary based on the proposal mask, L is the number of words, t i+1 is the (i + 1)-th target word to be predicted, is the visual feature, is the masked text from the 1st word to the i-th word.
[0046] In the second aspect, the present application provides a weakly supervised moment retrieval system with multi-granularity semantics and dynamic Gaussian modeling, including:
[0047] A multi-modal feature acquisition module for extracting visual features from a video frame sequence and extracting text features from a diverse query text;
[0048] A diverse proposal set acquisition module, connected to the multi-modal feature acquisition module, for performing multi-modal feature fusion, global and local text semantic injection-based proposal generation, and dynamic proposal integration on the extracted visual features and text features based on the multi-granularity semantic modeling method of the slot attention mechanism to obtain a diverse proposal set;
[0049] A frame weight distribution dynamic adjustment module, communicatively connected to the diverse proposal set acquisition module, for dynamically adjusting the frame weight distribution inside the diverse proposal set based on the dynamic Gaussian modeling method to generate the frame weights of the diverse proposal set with dynamically adjusted frame weight distribution;
[0050] The optimal proposal selection module is communicatively connected to the frame weight distribution dynamic adjustment module, and is used to select the optimal proposal from the diversified proposal set after the dynamic adjustment of the frame weight distribution based on the proposal selection mechanism reconstructed by the mask, so as to obtain the moment retrieval result;
[0051] The combined moment retrieval module is communicatively connected to the optimal proposal selection module. Using the extracted visual features and text features as the model input, the training process converges through contrastive learning based on the reduction loss of the cross-entropy of the positive proposal and the reduction loss of the frame weight of the negative proposal. During inference, the positive proposal with the minimum reduction loss is used as the combined moment retrieval result, and the parameters of the entire network are optimized through backpropagation to construct a weakly supervised combined moment retrieval model. Based on the constructed weakly supervised combined time retrieval model, the combined moment retrieval of diversified query texts is performed on the video data.
[0052] Combined with the first aspect, in one implementation, the frame weight distribution dynamic adjustment module includes:
[0053] The proposal frame weight distribution acquisition unit is used to perform proposal frame weight mapping on the diversified proposal set using the traditional Gaussian distribution based on the timestamps of the proposals, so as to obtain the proposal frame weight distribution;
[0054] The frame weight distribution dynamic adjustment unit is communicatively connected to the proposal frame weight distribution acquisition unit, and is used to introduce a learnable width ratio parameter to dynamically adjust the proposal frame weight distribution, so as to obtain a diversified proposal set after dynamic weight adjustment.
[0055] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0056] Through technical means such as multimodal feature fusion, dynamic frame weight adjustment, optimal proposal selection, and weakly supervised contrastive learning, the weakly supervised moment retrieval accuracy and robustness under complex text expressions are significantly improved, better adapting to the needs of diversified query texts, and achieving the effect of efficient and accurate video moment retrieval. Description of the Drawings
[0057] Figure 1 It is a schematic flowchart of the weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling provided by the embodiments of the present application;
[0058] Figure 2 It is another schematic flowchart of the weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling provided by the embodiments of the present application;
[0059] Figure 3 It is a functional module block diagram of the weakly supervised combined moment retrieval system of multi-granularity semantics and dynamic Gaussian modeling provided by the embodiments of the present application. Detailed Embodiments
[0060] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0061] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The descriptions such as "first", "second", and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second", and "third" are different types.
[0062] In the description of the embodiments of this application, "exemplary", "for example", or "for instance" etc. are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present the relevant concepts in a specific manner.
[0063] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; the "and / or" in the text is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.
[0064] In some processes described in the embodiments of this application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0065] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0066] In a first aspect, as Figure 1 and Figure 2 shown, the embodiments of this application provide a weakly supervised combined moment retrieval method for multi-granularity semantics and dynamic Gaussian modeling, which specifically includes the following steps:
[0067] Step S1: Extract visual features from the video frame sequence and extract text features from diverse query texts;
[0068] Step S2: Based on the multi-granularity semantic modeling method of the slot attention mechanism, perform multi-modal feature fusion on the extracted visual features and text features, generate proposals for global and local text semantic injection, and perform dynamic proposal integration to obtain a diverse set of proposals;
[0069] Step S3: Dynamically adjust the internal frame weight distribution of the diverse set of proposals based on the dynamic Gaussian modeling method to obtain a diverse set of proposals with dynamically adjusted frame weight distribution;
[0070] Step S4: Based on the proposal selection mechanism of mask reconstruction, select the optimal proposal from the diverse set of proposals with dynamically adjusted frame weight distribution, where the optimal proposal is a positive proposal with the minimum reconstruction loss;
[0071] Step S5: Use the extracted visual features and text features as model inputs. During the training process, perform contrastive learning convergence based on the reconstruction loss of the cross-entropy of positive proposals and the reconstruction loss of negative proposal frame weights. During inference, use the positive proposal with the minimum reconstruction loss as the combined moment retrieval result, backpropagate to optimize the parameters of the entire network, construct a weakly supervised combined moment retrieval model, and perform combined moment retrieval of diverse query texts on video data based on the constructed weakly supervised combined time retrieval model.
[0072] This application significantly improves the weakly supervised moment retrieval accuracy and robustness under complex text expressions, better adapts to the needs of diverse query texts, and achieves the effect of efficient and accurate video moment retrieval through technical means such as multi-modal feature fusion, dynamic frame weight adjustment, optimal proposal selection, and weakly supervised contrastive learning.
[0073] In one embodiment, the step S1: Extract visual features from the video frame sequence and extract text features from diverse query texts specifically includes the following steps:
[0074] Step S11: Use a pre-trained I3D network to extract visual features from the video frame sequence and capture the spatio-temporal information in the video; specifically, use an I3D model pre-trained on Kinetic-400 to extract video features Among them, L is the number of video segments, d v is the hidden dimension of the video feature extractor;
[0075] Step S12: Use the GLoVe framework to extract text features from diverse query texts and convert the query texts into semantic vector representations; specifically, for text features, use a GLoVe encoder to extract word-level text features from diverse query texts Among them, N is the number of words in the visual text, d t is the corresponding hidden dimension of the feature extractor;
[0076] Step S13: Map the multimodal features containing video features and text features to the same hidden dimension d through linear mapping to obtain video features and text features in the same hidden dimension, facilitating subsequent multimodal feature interaction and fusion.
[0077] Based on Step S11 and Step S12, provide the basic features for subsequent proposal generation as input.
[0078] In one embodiment, the step S2: A multi-granularity semantic modeling method based on the slot attention mechanism performs multimodal feature fusion, global and local text semantic injection for proposal generation, and dynamic proposal integration on the extracted visual features and text features to obtain a diverse set of proposals. Specifically, it includes the following steps:
[0079] Step S21: As shown in the following formula, perform interactive fusion on video features and text features to generate multimodal features The generated multimodal features contain both the spatio-temporal information of the video and the semantic information of diverse query texts, providing input for generating the global proposal set and the local proposal set:
[0080]
[0081] Among them, is the query vector mapped from the video feature ; and are respectively the transposed key vector and the value vector mapped from the text feature ; is the multimodal feature for text local semantic injection;
[0082] Step S22: Extract multimodal features through a learnable global vector to generate a global proposal set, and at the same time use the slot attention mechanism to iteratively decode the multimodal features with local text semantic injection to generate a local proposal set; further, it includes the following steps:
[0083] Step S221: Extract global semantic information from multimodal features through a learnable global vector, and generate a set of proposals for global text semantic injection using linear mapping to obtain the global proposal set P g , as shown in the following formula:
[0084]
[0085] where P g ∈R M×2 is the set of proposals for global text semantic injection, simply referred to as the global proposal set; Linear(·) is a linear mapping;
[0086] Step S222: Iteratively decode the multimodal features of local text semantic injection through the slot attention mechanism to generate a set of proposals for local text semantic injection, and obtain the local proposal set P l , as shown in the following formula:
[0087]
[0088] where P l ∈R M×2 is the set of proposals for local text semantic injection, simply referred to as the local proposal set; Slot Attention(·) is the slot attention mechanism; F ∈ R M×d is a learnable slot used to iteratively extract diverse semantic features from the given text for facilitating proposal mapping; is the multimodal input feature;
[0089] Step S23: In order to make full use of the semantics of diverse query texts, introduce the Hungarian matching algorithm and learnable fusion coefficients to dynamically integrate the global proposal set and the local proposal set to generate a diverse proposal set, and improve the model's adaptability to complex text expressions; further, it specifically includes the following steps:
[0090] Step S231: For the generated global proposal set P g ∈R M×2 (each row represents the center and width ) of each global proposal and the local proposal set P l ∈R M×2 (each row represents the center and width ) of each local proposal, calculate and obtain the Euclidean distance matrix C ∈ R M×M , as shown in the following formula:
[0091]
[0092] Step S232: Use the Hungarian algorithm to find the optimal bijection σ: {1, …, M} → {1, …, M} in the Euclidean distance matrix to minimize the total matching cost, and obtain the optimal matching relationship between the global proposal set and the local proposal set:
[0093]
[0094] Step S243: According to the optimal matching relationship, fuse the global proposal set and the local proposal set with weights. Let β = sigmoid(α) ∈ [0, 1] (α is a learnable parameter), then as shown in the following formula, the fused proposal P ∈ R M×2 is:
[0095]
[0096] In the formula, c i ′ represents the center of the i-th proposal, and w i ′ is the width of the i-th proposal.
[0097] In one embodiment, to optimize the frame weight reconstruction model inside the proposal, a learnable dynamic Gaussian distribution is constructed. Step S3: Dynamically adjust the frame weight distribution inside the diverse proposal set based on the dynamic Gaussian modeling method to obtain the diverse proposal set with dynamically adjusted frame weight distribution, which specifically includes the following steps:
[0098] Step S31: Based on the timestamps of the proposals, use the traditional Gaussian distribution to map the frame weights of the diverse proposal set to obtain the frame weight distribution of the proposals, providing an initial distribution for subsequent dynamic weight adjustment;
[0099] Step S32: Introduce a learnable width ratio parameter to dynamically adjust the frame weight distribution of the proposals, obtaining the diverse proposal set with dynamically adjusted weights, enabling the model to directly perceive the core action region of the proposals and avoiding over-reliance on a single frame in the middle of the proposals; Define a learnable width scaling factor to scale the maximum and minimum value regions in the original Gaussian distribution, allowing the model to fully perceive the frames where the core actions are located and optimizing the proposal distribution; For a certain proposal P j , the internal frame weight construction process is as follows:
[0100] First, construct the core action maximum and minimum value region mask M j of the proposal P i :
[0101]
[0102] where ρ is a learnable platform range coefficient; ∈ = 10 -3 is a constant controlling the steepness of the mask transition; N is the number of video frames, i ∈ [1, N]; σ is a predefined Gaussian distribution coefficient; M iThe core action maximum value region mask represents the mask value on video frame i;
[0103] Then, for proposal P j The intra-frame weights are divided into the original Gaussian term and the core frame sequence term, and the intra-frame weight distribution within the proposal is dynamically adjusted as shown in the following formula:
[0104]
[0105] Therefore, for M proposals, the dynamic Gaussian distribution is: D ∈ R M×N , where D is the dynamic Gaussian distribution, M is the number of proposals, N is the number of video frames, and R is the set of real numbers.
[0106] In one embodiment, after obtaining diverse proposals, the optimal proposal needs to be selected from them as the retrieval result to facilitate the convergence of contrastive learning; the semantics of the optimal proposal must be the closest to the semantics of the given text. Therefore, a part of the words in the original text is masked, and then the proposal features are used to restore it. The proposal with the restoration result closest to the original text is the optimal proposal. Based on this common sense, a proposal selection scheme based on mask reconstruction is proposed. Step S4: The proposal selection mechanism based on mask reconstruction selects the optimal proposal from the diverse proposal set with dynamically adjusted frame weight distribution, which specifically includes the following steps:
[0107] Step S41: Mask part of the words in the diverse query text to generate a masked text;
[0108] Step S42: Use the learnable action-aware Gaussian distribution module to aggregate the proposal features of the diverse proposal set with dynamically adjusted frame weights, and use the aggregated proposal features to restore the masked text to obtain a restored text;
[0109] Step S43: Evaluate the proposal quality according to the matching degree between the restored text and the original text, and select the optimal proposal from the diverse proposal set as the retrieval result.
[0110] In one embodiment, first, similar to the training process of BERT, one-third of the words in the original text T are masked to obtain
[0111] Then, the left-to-right restoration is performed according to the autoregressive process, and the restoration loss based on cross-entropy obtained is as follows:
[0112]
[0113] Among them, P p is the probability distribution according to the vocabulary based on the proposal mask, and L is the number of words.
[0114] In addition, similar to previous work, negative proposal frame weights are also introduced for text reconstruction for contrastive learning to obtain the reconstruction loss. Indirectly optimize the proposal quality.
[0115] The present application has the following remarkable advantages over the existing weakly supervised moment retrieval methods:
[0116] It can handle diverse user text expressions in practical scenarios: Through the innovative slot attention mechanism, the present application makes full use of the text information provided by users, captures both local text semantics and global multimodal semantics simultaneously, and generates a diverse set of proposals. Compared with existing methods that only rely on global features, this method can deconstruct text semantics more finely and effectively meet the expression requirements of complex text combinations. Especially when facing out-of-domain texts, this method can deeply mine their semantic information and generate proposals highly relevant to the query, thus significantly improving the accuracy of retrieval results. Experiments show that this method performs excellently in the combined moment retrieval task and can fully adapt to the challenges of diverse text expressions.
[0117] It can fully capture the internal action evolution process of proposals: Before aggregating proposal features, the present application introduces a learnable action-aware Gaussian distribution module to dynamically adjust the internal frame weight distribution of proposals through learnable parameters, breaking through the limitation of traditional methods that only emphasize the single middle frame. This design enables the model to accurately perceive the phased features such as the start, development, and end of actions, optimize the proposal feature expression, and provide more reliable feature support for subsequent query reconstruction and proposal selection. On the general dataset for combined moment retrieval, this method has achieved significant performance improvement. Especially when facing new text combinations, Recall@IoU reaches 39.39%, exceeding the previous best model by about 2%, fully verifying the superiority of this method.
[0118] In summary, through the innovative design of the slot attention mechanism and the learnable Gaussian distribution module, the present application not only solves the modeling problem of diverse text expressions but also optimizes the perception ability of the internal action evolution process of proposals, providing a more efficient and accurate solution for the weakly supervised combined moment retrieval task. The effects of the present application can be further illustrated by the following experiments.
[0119] Experimental conditions
[0120] This application conducts simulations by building a model using the PyTorch deep learning framework with a central processing unit of Gen Intel(R) Core(TM) i7, a graphics processing unit of GeForce GTX 4090, 64G of memory, and a Linux operating system. To comprehensively evaluate the performance of this application, the Charades dataset is used for training and testing. The Charades dataset is a video dataset focused on the understanding of daily indoor activities. The original version contains 9,848 videos, with each video being on average accompanied by 2.82 natural language query annotations, covering 157 action categories and 46 object category labels, and providing video-level action timestamps and object detection annotations. Charades-CG is a re-structured data partitioning version based on this. By merging the original training set and test set, removing easily predictable instances, and splitting it according to the combined generalization objective, it is divided into a training set (3,555 videos / 8,281 queries), a new combined test set (2,480 videos / 3,442 queries), a new vocabulary test set (588 videos / 703 queries), and a test set that retains common combinations (1,689 videos / 3,096 queries). Its core innovation lies in controlling data segmentation through a syntactic component combination statistical table, ensuring that the training set covers all basic language components, strictly separating the test scenarios for new combinations and new vocabulary, and avoiding overlap of video content between the training and test sets to more rigorously evaluate the language combination generalization ability of the model.
[0121] 1. Experimental Content
[0122] To fully verify the effectiveness of the algorithm of this application, 3 weakly supervised moment retrieval methods based on deep learning are separately selected from the Charades dataset for performance comparison. The comparison algorithms are from the following literatures respectively:
[0123] [1] Zheng M, Huang Y, Chen Q, et al. Weakly Supervised Temporal Sentence Grounding With Gaussian-Based Contrastive Proposal Learning[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022:15555 - 15564.
[0124] [2]Lv Z,Su B,Wen J R.Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment Localization[C] / / Proceedings of the ACM International Conference on Multimedia.2023:6539-6547.
[0125] [3]Kim S,Cho J,Yu J,et al.Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding[J].Proceedings of the AAAI Conference on Artificial Intelligence,2024,38(3):2795-2803.
[0126] This application uses the Rn@m metric and the mIoU metric to quantify the performance of each method. The Rn@m metric refers to the Recall@n metric, where IoU = m; the mIoU metric refers to the average IoU value between the predicted result and the annotated timestamp. The comparison of the metrics of the method of this application and the comparative algorithms is shown in Table 1 below:
[0127] Table 1 Performance comparison table on the Charades-CG dataset
[0128]
[0129] As can be seen from Table 1, the performance of the weakly supervised combined moment retrieval method of this application is the best. In addition, it can be noted that on the two subsets of combined moment retrieval ("new combination" and "new word"), the proposed invention performs even more outstandingly, verifying that the proposed multi-granularity semantic modeling module based on the slot attention mechanism can deeply mine the diverse text semantics given by different users in reality and generate reliable reference proposals. In addition, on the training set similar to the training set style, the proposed method can also achieve remarkable results, thanks to the construction of the learnable dynamic Gaussian distribution component, which is convenient for aggregating high-quality proposal features, reconstructing the masked text, and then selecting the optimal proposal. Generally speaking, the method of this application has better performance, and the advanced effectiveness of the method of this application is further proved through the comparative experiments on the public dataset.
[0130] Second aspect, the present application provides a weakly supervised moment retrieval system with multi-granularity semantics and dynamic Gaussian modeling, including a multi-modal feature acquisition module 100, a diverse proposal set acquisition module 200, a frame weight distribution dynamic adjustment module 300, an optimal proposal selection module 400, and a combined moment retrieval module 500; the multi-modal feature acquisition module 100 is used to extract visual features from a video frame sequence and extract text features from diverse query texts; the diverse proposal set acquisition module 200 is communicatively connected to the multi-modal feature acquisition module 100, and is used for a multi-modal feature fusion, global and local text semantic injection-based proposal generation, and dynamic proposal integration of the extracted visual features and text features based on a multi-granularity semantic modeling method with a slot attention mechanism to obtain a diverse proposal set; the frame weight distribution dynamic adjustment module 300 is communicatively connected to the diverse proposal set acquisition module 200, and is used for dynamically adjusting the frame weight distribution inside the diverse proposal set based on a dynamic Gaussian modeling method to generate the frame weights of the diverse proposal set after the frame weight distribution is dynamically adjusted; the optimal proposal selection module 400 is communicatively connected to the frame weight distribution dynamic adjustment module 300, and is used for selecting the optimal proposal from the diverse proposal set after the frame weight distribution is dynamically adjusted based on a proposal selection mechanism based on mask reconstruction; the combined moment retrieval module 500 is communicatively connected to the optimal proposal selection module 400, uses the extracted visual features and text features as model inputs, and performs contrastive learning convergence based on the reconstruction loss of the cross-entropy of positive proposals and the reconstruction loss of the frame weights of negative proposals during the training process. During inference, the positive proposal with the minimum reconstruction loss is used as the combined moment retrieval result, and the parameters of the entire network are optimized by backpropagation to construct a weakly supervised combined moment retrieval model. Based on the constructed weakly supervised combined time retrieval model, combined moment retrieval of diverse query texts is performed on video data.
[0131] The frame weight distribution dynamic adjustment module includes:
[0132] A proposal frame weight distribution acquisition unit, which is used to perform proposal frame weight mapping on the diverse proposal set using a traditional Gaussian distribution based on the timestamps of the proposals to obtain a proposal frame weight distribution;
[0133] A frame weight distribution dynamic adjustment unit, which is communicatively connected to the proposal frame weight distribution acquisition unit, and is used to introduce a learnable width ratio parameter to dynamically adjust the proposal frame weight distribution to obtain a diverse proposal set after dynamic weight adjustment.
[0134] Among them, the function implementations of the various modules in the above-mentioned weakly supervised moment retrieval system with multi-granularity semantics and dynamic Gaussian modeling correspond to the respective steps in the above-mentioned embodiment of the weakly supervised combined moment retrieval method with multi-granularity semantics and dynamic Gaussian modeling, and their functions and implementation processes will not be elaborated here one by one.
[0135] In a third aspect, an embodiment of the present application provides a weakly supervised moment retrieval device for multi-granularity semantics and dynamic Gaussian modeling. The weakly supervised moment retrieval device for multi-granularity semantics and dynamic Gaussian modeling can be a device with data processing functions such as a personal computer (PC), a laptop, a server, etc.
[0136] The communication interface includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of components inside the weakly supervised moment retrieval device for multi-granularity semantics and dynamic Gaussian modeling, as well as interfaces for implementing the interconnection between the weakly supervised moment retrieval device for multi-granularity semantics and dynamic Gaussian modeling and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.
[0137] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0138] The processor can be a general-purpose processor. The general-purpose processor can call the weakly supervised moment retrieval program stored in the memory and execute the weakly supervised combined moment retrieval method provided by the embodiment of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the weakly supervised moment retrieval program for multi-granularity semantics and dynamic Gaussian modeling is called can refer to the various embodiments of the weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling in the present application, which will not be elaborated here.
[0139] In a fourth aspect, an embodiment of the present application further provides a readable storage medium.
[0140] The weakly supervised moment retrieval program for multi-granularity semantics and dynamic Gaussian modeling is stored on the readable storage medium of the present application. When the weakly supervised moment retrieval program for multi-granularity semantics and dynamic Gaussian modeling is executed by a processor, the steps of the weakly supervised combined moment retrieval method as described above are implemented.
[0141] Among them, for the method implemented when the weak supervision moment retrieval program of multi-granularity semantics and dynamic Gaussian modeling is executed, reference can be made to each embodiment of the weak supervision combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling in this application, which will not be elaborated here.
[0142] It should be noted that the serial numbers of the above embodiments of this application are only for description and do not represent the superiority or inferiority of the embodiments.
[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in each embodiment of this application.
[0144] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
Claims
1. A weakly supervised combined moment retrieval method based on multi-granularity semantics and dynamic Gaussian modeling, characterized in that It includes the following steps: Extract visual features from the video frame sequence and extract text features from diverse query texts; Based on the multi-granularity semantic modeling method of the slot attention mechanism, perform multi-modal feature fusion on the extracted visual features and text features, generate proposals through global and local text semantic injection, and perform dynamic proposal integration to generate a diverse set of proposals; Based on the dynamic Gaussian modeling method, dynamically adjust the frame weight distribution within the diverse set of proposals to generate the frame weights of the diverse set of proposals with dynamically adjusted frame weight distribution; Based on the proposal selection mechanism of mask reconstruction, select the optimal proposal from the diverse set of proposals with dynamically adjusted frame weight distribution to obtain the moment retrieval result; Using the extracted visual features and text features as model inputs, the training process converges through contrastive learning based on the reconstruction loss of the cross-entropy of positive proposals and the reconstruction loss of negative proposal frame weights. During inference, the positive proposal with the minimum reconstruction loss is used as the combined moment retrieval result, and the parameters of the entire network are optimized through backpropagation to construct a weakly supervised combined moment retrieval model for performing combined moment retrieval of diverse query texts on video data.
2. The weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling according to claim 1, characterized in that The steps of extracting visual features from the video frame sequence and extracting text features from diverse query texts specifically include the following steps: Use the pre-trained I3D network to extract visual features from the video frame sequence; Use the GLoVe framework to extract text features from diverse query texts; Map the extracted video features and text features to the same hidden dimension through linear mapping to obtain video features and text features in the same hidden dimension.
3. The weakly supervised combined moment retrieval method with multi-granularity semantics and dynamic Gaussian modeling according to claim 1, characterized in that The multi-granularity semantic modeling method based on the slot attention mechanism performs multi-modal feature fusion on the extracted visual features and text features, generates proposals through global and local text semantic injection, and performs dynamic proposal integration to obtain a diverse set of proposals. Specifically, it includes the following steps: Interactively fuse video features and text features through a Transformer decoder to generate multi-modal features; Extract multi-modal features through a learnable global vector to generate a global set of proposals; Introduce the Hungarian matching algorithm and learnable fusion coefficients to dynamically integrate the global set of proposals and the local set of proposals to generate a diverse set of proposals.
4. The weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling according to claim 3, characterized in that The steps of extracting multi-modal features through a learnable global vector to generate a global set of proposals specifically include the following steps; Extract global semantic information from multi-modal features through a learnable global vector and use linear mapping to generate a set of proposals with global text semantic injection to obtain the global set of proposals; Iteratively decode the multi-modal features with local text semantic injection through the slot attention mechanism to generate a set of proposals with local text semantic injection to obtain the local set of proposals.
5. The weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling according to claim 3, characterized in that The steps of introducing the Hungarian matching algorithm and learnable fusion coefficients to dynamically integrate the global set of proposals and the local set of proposals to generate a diverse set of proposals specifically include the following steps: Calculate the Euclidean distance matrix of the global set of proposals and the local set of proposals; Find the optimal bijective mapping in the Euclidean distance matrix through the Hungarian algorithm to obtain the optimal matching relationship between the global set of proposals and the local set of proposals; Perform weighted fusion on the generation parameters and local parameters of the optimal matching relationship to obtain a diverse set of proposals.
6. The weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling according to claim 1, characterized in that, Dynamically adjust the internal frame weight distribution of the diverse proposal set based on the dynamic Gaussian modeling method, and obtain the diverse proposal set with the dynamically adjusted frame weight distribution, which specifically includes the following steps: Based on the timestamp of the proposal, use the traditional Gaussian distribution to perform proposal frame weight mapping on the diverse proposal set to obtain the proposal frame weight distribution; Introduce a learnable width ratio parameter to dynamically adjust the proposal frame weight distribution and generate a diverse proposal set with dynamically adjusted weights.
7. The weakly supervised combined moment retrieval method based on multi-granularity semantics and dynamic Gaussian modeling according to claim 1, characterized in that The proposal selection mechanism based on mask reconstruction selects the optimal proposal from the diverse proposal set with the dynamically adjusted frame weight distribution to obtain the moment retrieval result, which specifically includes the following steps: Mask out some words in the diverse query text to generate a masked text; Use a learnable action-aware Gaussian distribution module to aggregate the proposal features of the diverse proposal set with dynamically adjusted frame weights, and use the aggregated proposal features to restore the masked text to obtain the restored text; Evaluate the quality of the selected proposal according to the matching degree between the restored text and the original text, and select the optimal proposal from the diverse proposal set as the retrieval result.
8. The weakly supervised combined moment retrieval method of multi-granularity semantics and dynamic Gaussian modeling according to claim 1, characterized in that The reduction loss of the cross entropy As shown in the following formula: where P p is the probability distribution according to the vocabulary based on the proposed mask, L is the number of words, t i+1 is the (i + 1)-th target word to be predicted, is the visual feature, is the masked text from the 1st word to the i-th word.
9. A weakly supervised moment retrieval system with multi-granularity semantics and dynamic Gaussian modeling, characterized in that, Include: A multi-modal feature acquisition module for extracting visual features from the video frame sequence and text features from the diverse query text; A diverse proposal set acquisition module, connected to the multi-modal feature acquisition module, for performing multi-modal feature fusion, global and local text semantic injection-based proposal generation, and dynamic proposal integration on the extracted visual and text features using the multi-granularity semantic modeling method based on the slot attention mechanism to obtain a diverse proposal set; A frame weight distribution dynamic adjustment module, communicatively connected to the diverse proposal set acquisition module, for dynamically adjusting the internal frame weight distribution of the diverse proposal set based on the dynamic Gaussian modeling method to generate the frame weight of the diverse proposal set with the dynamically adjusted frame weight distribution; An optimal proposal selection module, communicatively connected to the frame weight distribution dynamic adjustment module, for selecting the optimal proposal from the diverse proposal set with the dynamically adjusted frame weight distribution based on the proposal selection mechanism based on mask reconstruction to obtain the moment retrieval result; A combined moment retrieval module, communicatively connected to the optimal proposal selection module, using the extracted visual and text features as the model input. During the training process, contrastive learning convergence is performed based on the reconstruction loss of the positive proposal's cross-entropy and the reconstruction loss of the negative proposal's frame weight. During inference, the positive proposal with the minimum reconstruction loss is used as the combined moment retrieval result, and the parameters of the entire network are optimized by backpropagation to construct a weakly supervised combined moment retrieval model. Based on the constructed weakly supervised combined time retrieval model, combined moment retrieval of diverse query texts is performed on the video data.
10. The weakly supervised moment retrieval system with multi-granularity semantics and dynamic Gaussian modeling according to claim 9, characterized in that, The frame weight distribution dynamic adjustment module includes: A proposal frame weight distribution acquisition unit for performing proposal frame weight mapping on the diverse proposal set using the traditional Gaussian distribution based on the timestamp of the proposal to obtain the proposal frame weight distribution; A frame weight distribution dynamic adjustment unit, communicatively connected to the proposal frame weight distribution acquisition unit, for introducing a learnable width ratio parameter to dynamically adjust the proposal frame weight distribution to obtain a diverse proposal set with dynamically adjusted weights.