Local Supervision Long Video Temporal Text Retrieval Method and System
By introducing local supervision settings and multi-scale Transformer feature comparison learning in long video timing text retrieval, the problem of low retrieval performance caused by high labeling costs and low quality labeling in the existing technology is solved, and efficient and economical video timing retrieval effect is achieved.
Patent Information
- Application Number
- CN202211581256.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-12-09
AI Technical Summary
In the search for long video timing texts, strong supervision settings require expensive labeling and cleaning costs, which limits the scale of labeling, while weak supervision settings have low retrieval performance and limited practical value due to low quality labeling.
提出局部监督设定,通过在长视频中随机选择连续的若干秒作为事件局部时序标注,结合多尺度Transformer和特征对比学习,优化多模态特征融合,并通过边界细化学习模块进一步微调检索结果。
While maintaining low labeling costs, it provides accurate search position anchors to improve model performance; improves the search effect by aggregating events, backgrounds and text features; refines the search results from coarse to fine learning strategies to obtain more accurate and complete video timing retrieval effects.
Smart Images

Figure CN115809352B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and specifically, to a local supervised long video temporal text retrieval method and system. Background Art
[0002] In recent years, multimedia videos have shown an explosive growth, giving rise to a huge demand for video content understanding and key segment retrieval. In daily applications, natural language in the form of text is the main medium for humans to understand and process long video content. Therefore, using text composed of several sentences in the temporal dimension to retrieve segments from long videos: long video temporal text retrieval, constitutes the key to video editing, recommendation, and secondary creation.
[0003] Regarding an unedited long video and a retrieval text as a group of samples, for a large number of samples, they are first divided into two groups, one of which is used as a training set, and event location annotations for text descriptions are made in the temporal dimension, while the other is used as a test set without any annotations. The goal of long video temporal text retrieval is to use the training set to learn a generalizable model that can process samples in the test set, that is, given a text description, it can retrieve the occurrence and end moments of the corresponding event in the long video completely and accurately.
[0004] More specifically, for the training set annotation, there are two existing settings for long video temporal text retrieval: strong supervision and weak supervision. The strong supervision setting requires marking the exact start and end moments of events in the long video for each text description in the training set. The weak supervision setting only needs to verify that the text descriptions in the training set have corresponding events in the long video.
[0005] Patent document CN112015947A discloses a language description-guided video temporal localization method and system. It adopts the strong supervision setting, designs a temporal boundary adaptive optimization framework, and introduces a reinforcement learning paradigm to refine the detection results. However, the above patent ignores that the precise annotation in strong supervision is a double-edged sword. While bringing excellent retrieval performance, it also requires expensive annotation and cleaning costs, and to a certain extent limits the annotation scale. Therefore, it is difficult to adapt to large-scale applications in actual scenarios.
[0006] Patent document CN112685597A discloses a weak supervision video segment retrieval method and system based on an erasure mechanism. It adopts the weak supervision setting, designs a dynamic erasure mechanism, and realizes retrieval through the cooperation of a double-branch shared network. Although the annotation cost is low and the scale of the training set can be easily expanded, the above patent ignores the low quality of weak supervision annotations, and these low-quality annotations seriously limit the retrieval performance and have very low practical value.
[0007] Therefore, it is necessary to propose a new supervision setting (annotation method) and design a corresponding reasonable technical solution to improve the above technical problems. Summary of the Invention
[0008] Aiming at the defects in the prior art, the purpose of the present invention is to provide a local supervision long video temporal text retrieval method and system.
[0009] A local supervision long video temporal text retrieval method provided by the present invention includes:
[0010] Text feature extraction step: using a text feature encoding network composed of fully connected layers to extract text initial features of a preset dimension from the input retrieval text;
[0011] Video feature extraction step: using a visual feature encoding network composed of 3D depth convolution to extract video initial features of a preset dimension from the input long video;
[0012] Cross-modal feature fusion step: using a cross-modal feature fusion network composed of multi-scale Transformer to deeply fuse the input text initial features and video initial features, and output a text feature map and a video feature map;
[0013] Event detection step: using an event detection network composed of a multi-layer perceptron to map the video feature map into an event proposal described by text, the event proposal is composed of an event center and an event length, and then calculating an event temporal position mask based on the event proposal to represent the time position of the video frames containing the event, as the output;
[0014] Local position label supervision step: constraining the event proposal to completely contain local temporal labels;
[0015] Event feature aggregation step: applying the event temporal position mask to the video feature map to generate an event feature map, and aggregating to generate event-level visual features;
[0016] Multi-modal feature contrast step: using feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space;
[0017] Rough retrieval result generation step: based on the event proposal, using a threshold method to generate a rough retrieval temporal result;
[0018] Accurate boundary refinement step: using the rough retrieval temporal result to train an accurate boundary refinement network, designing a multi-scale fine-tuning paradigm for the event boundary, and providing an explicit optimization constraint for the boundary at the same time, and outputting an accurate boundary retrieval result map;
[0019] Precise retrieval result generation step: Based on the precise boundary retrieval result graph, use the threshold method to generate the prediction result.
[0020] Preferably, in the text feature extraction step, the retrieval text is tokenized and encoded based on a tokenizer, and a text feature encoding network composed of fully connected layers is used to map the tokenized encoding to a text initial feature map F of dimension W*D q ; where W represents the number of words in the retrieval text, and D q represents the initial feature dimension of the text; q
[0021] In the video feature extraction step, a visual feature encoding network composed of 3D depth convolution is used to map the video RGB frames to a video initial feature F of dimension T*D v ; where T represents the time length of the video, and D v represents the initial feature dimension of the video; v
[0022] The cross-modal feature fusion step: Use a cross-modal feature fusion network Φ composed of multi-scale Transformers fus , to deeply fuse the input text initial feature F q ' and video initial feature F v ', and output the text feature map F q and video feature map F v , and the calculation formula is as follows:
[0023] F v ,F q =Φ fus (F' v ,F' q )
[0024] where the dimension of F q is W*D, the dimension of F v is T*D, and D represents the fusion feature dimension.
[0025] Preferably, in the event detection step, an event detection network Φ composed of a multi-layer perceptron is used det , to map the video feature map F v to an event proposal P=(C,L) of the text description, and the calculation formula is as follows:
[0026] P=(C,L)=Φ det (F v )
[0027] Where C represents the predicted event center and L represents the predicted event length. Furthermore, the plateau shape transformation is used to map the event proposal P to the event temporal position mask M as the output. M represents the temporal positions of the video frames containing the event, and the dimension is T*1.
[0028] Preferably, in the local position label supervision step, according to the provided local temporal label Y, it is constrained that the event proposal P completely contains the local temporal label, and the constraint calculation formula is as follows:
[0029]
[0030] Where θ det are the parameters of the event detection network, (F v W , Y W ) represents the distribution of the video feature map and the local temporal label, f v i represents an instance of the video feature map, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function.
[0031] Preferably, in the event feature aggregation step, the event temporal position mask M is applied to the video feature map to generate an event feature map, and an average is taken in the temporal dimension to aggregate into the event-level visual feature f e . The inverse of the event temporal position mask, 1-M, is applied to the video feature map to generate a background feature map, and an average is taken in the temporal dimension to aggregate into the background-level visual feature f b . An average is taken of the video feature map in the temporal dimension to aggregate into the video-level visual feature f a , and the dimensions of all three are 1*D. The calculation formula is as follows:
[0032] f e = AVG(F v *M) f b = AVG(F v *(1 - M)) f a = AVG(F v )
[0033] Where AVG is the average function in the temporal dimension;
[0034] In the multi-modal feature contrast step: An average is taken of the text feature map F q in the word dimension to generate the text feature f q, with a dimension of 1*D, calculates the cosine similarity for all text features in the dataset, clusters all samples into several semantically similar clusters, and uses feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, while event-text features and event-event features with dissimilar semantics are dispersed in the representation space. The feature contrast loss function is calculated as follows:
[0035]
[0036]
[0037]
[0038]
[0039] where θ det are the parameters of the event detection network, and θ fus are the parameters of the cross-modal feature fusion network. represents an event-level visual feature instance, represents a video-level visual feature instance, represents a background-level visual feature instance, f e + represents a set of event visual features with similar semantics, and f q + represents a set of text features with similar semantics, and f e - represents a set of event visual features with dissimilar semantics, and f q - represents a set of text features with dissimilar semantics. S(·,·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrast learning function, and τ represents the temperature coefficient.
[0040] For the rough retrieval result generation step, for the event proposal P, a rough retrieval time series result H with a dimension of T*1 is generated using the threshold method.
[0041] For the precise boundary refinement step, based on the rough retrieval time series result H, the precise boundary refinement network Φ pre is trained to design a multi-scale fine-tuning paradigm for the event boundary and provide explicit optimization constraints for the boundary at the same time.
[0042] The constraint loss function is as follows:
[0043]
[0044] where θ pre are the parameters of the precise boundary refinement network, (F v W,F q W ,HW ) represents the distributions of video feature maps, text feature maps, and retrieval temporal pseudo-labels, represents an instance of a video feature map, represents an instance of a text feature map, H i is its temporal pseudo-label, U represents the cross-entropy function. The precise boundary refinement network Φ pre outputs a precise boundary retrieval result map K, whose dimension is T*1;
[0045] The precise retrieval result generation step is based on the precise boundary retrieval result map K and uses a threshold method to generate a prediction result.
[0046] A local supervised long video temporal text retrieval system provided by the present invention includes:
[0047] Text feature extraction module: Using a text feature encoding network composed of fully connected layers, it extracts text initial features of a preset dimension from the input retrieval text;
[0048] Video feature extraction module: Using a visual feature encoding network composed of 3D depth convolution, it extracts video initial features of a preset dimension from the input long video;
[0049] Cross-modal feature fusion module: Using a cross-modal feature fusion network composed of multi-scale Transformers, it deeply fuses the input text initial features and video initial features, and outputs text feature maps and video feature maps;
[0050] Event detection module: Using an event detection network composed of a multi-layer perceptron, it maps the video feature map to an event proposal described by text. This event proposal consists of an event center and an event length, and then calculates an event temporal position mask based on the event proposal, representing the time positions of the video frames containing the event, as the output;
[0051] Local position label supervision module: Constraining the event proposal to completely contain local temporal labels;
[0052] Event feature aggregation module: Acting the event temporal position mask on the video feature map to generate an event feature map, and aggregating to generate event-level visual features;
[0053] Multi-modal feature contrast module: Using feature contrast learning, encouraging event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space;
[0054] Rough retrieval result generation module: Based on the event proposal, using a threshold method to generate a rough retrieval temporal result;
[0055] Accurate Boundary Refinement Module: Using the rough retrieval time series results, train the accurate boundary refinement network, design a multi-scale fine-tuning paradigm for event boundaries, and at the same time provide explicit optimization constraints for the boundaries, and output the accurate boundary retrieval result map;
[0056] Accurate Retrieval Result Generation Module: Based on the accurate boundary retrieval result map, use the threshold method to generate the prediction results.
[0057] Preferably, the text feature extraction module tokenizes and encodes the retrieval text based on the tokenizer, and uses the text feature encoding network composed of fully connected layers to map the tokenized encoding to the text initial feature map F of W*D q dimensions; where W represents the number of words in the retrieval text, and D q represents the initial feature dimension of the text; q
[0058] The video feature extraction module uses the visual feature encoding network composed of 3D depth convolution to map the video RGB frames to the video initial feature F' of T*D v dimensions; where T represents the time length of the video, and D v represents the initial feature dimension of the video; v
[0059] The cross-modal feature fusion module: uses the cross-modal feature fusion network Φ composed of multi-scale Transformer fus , and performs deep fusion on the input text initial feature F q ' and video initial feature F v ', and outputs the text feature map F q and the video feature map F v , and the calculation formula is as follows:
[0060] F v ,F q =Φ fus (F' v ,F' q )
[0061] where the dimension of F q is W*D, the dimension of F v is T*D, and D represents the fusion feature dimension.
[0062] Preferably, the event detection module uses the event detection network Φ composed of multi-layer perceptrons det , and maps the video feature map F v to the event proposal P=(C,L) of the text description, and the calculation formula is as follows:
[0063] P=(C,L)=Φ det (F v )
[0064] Where C represents the predicted event center, L represents the predicted event length, and then the plateau shape transformation is used to map the event proposal P to the event temporal position mask M as the output. M represents the temporal position of the video frames containing the event, and the dimension is T*1.
[0065] Preferably, the local position label supervision module constrains the event proposal P to completely contain the local temporal label according to the provided local temporal label Y, and the constraint calculation formula is as follows:
[0066]
[0067] Where θ det is the parameter of the event detection network, (F v W , Y W ) represents the distribution of the video feature map and the local temporal label, represents the video feature map instance, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function.
[0068] Preferably, the event feature aggregation module acts the event temporal position mask M on the video feature map to generate an event feature map, and averages it in the temporal dimension to aggregate it into an event-level visual feature f e . Acting the inverse 1-M of the event temporal position mask on the video feature map generates a background feature map, and averages it in the temporal dimension to aggregate it into a background-level visual feature f b . Averaging the video feature map in the temporal dimension aggregates it into a video-level visual feature f a , and the dimensions of all three are 1*D. The calculation formula is as follows:
[0069] f e = AVG(F v *M) f b = AVG(F v *(1 - M)) f a = AVG(F v )
[0070] Where AVG is the average function in the temporal dimension;
[0071] The multimodal feature contrast module: averages the text feature map F q in the word dimension to generate the text feature f q, with a dimension of 1*D, calculates the cosine similarity for all text features in the dataset, clusters all samples into several semantically similar clusters, and uses feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, while event-text features and event-event features with dissimilar semantics are dispersed in the representation space. The feature contrast loss function is calculated as follows:
[0072]
[0073]
[0074]
[0075]
[0076] where θ det are the parameters of the event detection network, and θ fus are the parameters of the cross-modal feature fusion network. represents the event-level visual feature instance, represents the video-level visual feature instance, represents the background-level visual feature instance, and f e + represents the set of event visual features with similar semantics, and f q + represents the set of text features with similar semantics, and f e - represents the set of event visual features with dissimilar semantics, and f q - represents the set of text features with dissimilar semantics. S(·,·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrast learning function, and τ represents the temperature coefficient.
[0077] The rough retrieval result generation module generates a rough retrieval time series result H with a dimension of T*1 for the event proposal P using the threshold method.
[0078] The precise boundary refinement module trains the precise boundary refinement network Φ pre based on the rough retrieval time series result H, designs a multi-scale fine-tuning paradigm for the event boundary, and provides explicit optimization constraints for the boundary at the same time.
[0079] The constraint loss function is as follows:
[0080]
[0081] where θ pre are the parameters of the precise boundary refinement network, and (F v W ,Fq W ,H W ) represent the distributions of video feature maps, text feature maps, and retrieval temporal pseudo-labels. represent video feature map instances, represent text feature map instances, H i is its temporal pseudo-label, and U represents the cross-entropy function. The precise boundary refinement network Φ pre outputs an accurate boundary retrieval result map K, whose dimension is T*1;
[0082] The precise retrieval result generation module generates a prediction result using a threshold method based on the precise boundary retrieval result map K.
[0083] Compared with the prior art, the present invention has the following beneficial effects:
[0084] 1. The present application proposes a local supervision setting to balance the annotation cost and model performance. As an intermediate setting between strong and weak supervision, local supervision provides an accurate retrieval position anchor while maintaining a low annotation cost, laying a solid foundation for strong performance. Compared with strong supervision, local annotation does not require searching for boundaries, so the annotation speed is fast, the annotation cost is low, and it is easy to scale up. Compared with weak supervision, this local annotation provides limited but clear event anchors, facilitating the rapid training and convergence of the model, thus obtaining more powerful retrieval performance;
[0085] 2. The present application makes a special design for the local supervision setting and proposes a quadruple contrastive learning strategy to fully optimize multi-modal features. By aggregating event features, background features, and text features, the present application encourages semantic consistency within samples, between samples, within video modalities, and between video modalities, that is, similar semantics are aggregated in the feature space and dissimilar semantics are dispersed. Therefore, video events and text are aligned and fused in the feature space, improving the retrieval effect;
[0086] 3. Considering the incompleteness of local supervision, the present application designs a learning strategy from coarse to fine to refine the retrieval results. After generating rough retrieval results, a boundary refinement learning module is additionally introduced to further fine-tune the retrieval results, thereby obtaining a more accurate and complete video temporal retrieval effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent:
[0088] Figure 1 is the method flow chart in the embodiment of the present invention;
[0089] Figure 2 is the system schematic diagram in the embodiment of the present invention. Detailed Implementation Manner
[0090] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.
[0091] Example 1:
[0092] As Figure 1 shown, it is a flowchart of an embodiment of a method for local supervised long video temporal text retrieval based on a contrast-unified strategy of the present invention. Aiming at the defects of the existing supervision setting, the present invention first proposes a local supervision setting to balance the annotation cost and the model performance, that is, for each text description in the training set, during the event occurrence stage in the long video, several consecutive seconds (even one frame) are randomly selected as the event local temporal annotation. As an intermediate setting between strong and weak supervision, local supervision not only maintains a low annotation cost but also provides an accurate retrieval position anchor, laying a strong performance foundation. For the specific setting of local supervision, the present invention provides a method and system for local supervised long video temporal text retrieval based on a contrast-unified strategy. Specifically, a quadruple contrast learning strategy is carefully designed to fully optimize the multi-modal features. By aggregating event features, background features, and text features, the present invention encourages semantic consistency within samples, between samples, within video modalities, and between video modalities, that is, similar semantics are aggregated in the feature space and dissimilar semantics are dispersed. Therefore, video events and texts are aligned and fused in the feature space, improving the retrieval effect. Considering the incompleteness of local supervision, the present invention designs a learning strategy from coarse to fine to refine the retrieval results. After generating a rough retrieval result, a boundary refinement learning module is additionally introduced to further fine-tune the retrieval result, thereby obtaining a more accurate and complete video temporal retrieval effect.
[0093] According to a local supervised long video temporal text retrieval method provided by the present invention, the method includes the following steps:
[0094] Text feature extraction step: Using a text feature encoding network composed of a fully connected layer, extract the text initial features of a preset dimension from the input retrieval text;
[0095] Video feature extraction step: Using a visual feature encoding network composed of 3D depth convolution, extract the video initial features of a preset dimension from the input long video;
[0096] Cross-modal feature fusion step: Use a cross-modal feature fusion network composed of multi-scale Transformers to deeply fuse the input text initial features and video initial features, and output a text feature map and a video feature map;
[0097] Event detection step: Use an event detection network composed of a multi-layer perceptron to map the video feature map into an event proposal described by text, which is composed of an event center and an event length. Furthermore, calculate an event temporal position mask based on the event proposal, representing the temporal positions of the video frames containing the event, as the output;
[0098] Local position label supervision step: Constrain the event proposal to completely contain the local temporal label;
[0099] Event feature aggregation step: Apply the event temporal position mask to the video feature map to generate an event feature map, and generate event-level visual features through aggregation;
[0100] Multi-modal feature contrast step: Use feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space;
[0101] Rough retrieval result generation step: Based on the event proposal, use a threshold method to generate a rough retrieval temporal result;
[0102] Precise boundary refinement step: Use the rough retrieval temporal result to train a precise boundary refinement network, design a multi-scale fine-tuning paradigm for the event boundary, and at the same time provide explicit optimization constraints for the boundary, and output a precise boundary retrieval result map;
[0103] Precise retrieval result generation step: Based on the precise boundary retrieval result map, use a threshold method to generate a prediction result.
[0104] Specifically, the text feature extraction step includes: Tokenize and encode the retrieved text based on a tokenizer, and use a text feature encoding network composed of a fully connected layer to map the tokenized encoding into a text initial feature map F' of W*D q dimensions. Where W represents the number of words in the retrieved text, and D q . represents the initial feature dimension of the text; q represents the initial feature dimension of the text;
[0105] Specifically, the video feature extraction step includes: Use a visual feature encoding network composed of 3D depth convolution to map the input long video RGB frames into a video initial feature F' of T*D v dimensions; where T represents the time length of the video, and D v . represents the initial feature dimension of the video; v represents the initial feature dimension of the video;
[0106] Specifically, the cross-modal feature fusion step includes: using a cross-modal feature fusion network Φ composed of a multi-scale Transformer architecture fus , to perform deep alignment and fusion on the input text initial feature F' q and video initial feature F' v , and output the text feature map F q and video feature map F v . The calculation formula is as follows:
[0107] F v , F q = Φ fus (F' v , F' q ) (1)
[0108] where the dimension of F q is W*D, the dimension of F v is T*D, and D represents the fusion feature dimension;
[0109] Specifically, the event detection step includes: using an event detection network Φ composed of a multi-layer perceptron det , to map the video feature map F v to the event proposal P=(C, L) of the text description, which represents the time interval corresponding to the event described in the long video. The calculation formula is as follows:
[0110] P=(C, L = Φ det (F v ) (2)
[0111] where C represents the predicted event center and L represents the predicted event length. Then, use the plateau shape transformation to map the event proposal P to the event temporal position mask M as the output. M represents the temporal positions of the video frames containing the event, and the dimension is T*1. Compared with predicting the frame-level event probability, predicting the event proposal only uses two variables, the center and the length, to characterize the event instance, so it has a smaller solution space. In addition, the frame-level event probability within the constraint interval of the event proposal is 1, and the probability outside the interval is 0, so it naturally has good probability smoothness.
[0112] Specifically, the local position label supervision step includes: according to the provided local temporal label Y, supervising that the event proposal P completely contains the local temporal label, calculating the loss function to train the event detection network until the loss function converges;
[0113] The calculation formula of the loss function is as follows:
[0114]
[0115] where θdet are the parameters of the event detection network, (F v W , Y W ) represents the distribution of video feature maps and local temporal labels. represents an instance of the video feature map, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function. Under the constraint of the loss function, the model retrieves events with local annotations as anchor frames to avoid false positive results. Sufficient annotated data encourages the model to expand the anchor frames to improve the retrieval results.
[0116] Specifically, the event feature aggregation step includes: applying the event temporal position mask M to the video feature map to generate an event feature map, and averaging in the time dimension to aggregate into event-level visual features f e . Applying the inverse of the event temporal position mask, 1 - M, to the video feature map to generate a background feature map, and averaging in the time dimension to aggregate into background-level visual features f b . Averaging the video feature map in the time dimension to aggregate into video-level visual features f a , and the dimensions of all three are 1 * D. The calculation formulas are as follows:
[0117] f e = AVG(F v * M) f b = AVG(F v *(1 - M)) f a = AVG(F v ) (4)
[0118] where AVG is the average function in the time dimension.
[0119] Specifically, the multi-modal feature contrast step includes: averaging the text feature map F q in the word dimension to generate text features f q , whose dimension is 1 * D. Calculate the cosine similarity for all text features in the dataset, and cluster all samples into several semantically similar clusters. Using feature contrast learning, encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space. The feature contrast loss function, the calculation formula is as follows:
[0120]
[0121]
[0122]
[0123]
[0124] where θ det is a parameter of the event detection network, and θ fus is a parameter of the cross-modal feature fusion network. represents an event-level visual feature instance, represents a video-level visual feature instance, represents a background-level visual feature instance, f e + represents a set of event visual features with similar semantics, f q + represents a set of text features with similar semantics, f e - represents a set of event visual features with dissimilar semantics, f q - represents a set of text features with dissimilar semantics. S(·,·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrast learning function, and τ represents the temperature coefficient. Under the constraint of the contrast loss function, video events and texts are aligned and fused in the feature space to improve the retrieval effect.
[0125] Specifically, the step of generating the rough retrieval result includes: for the event proposal P, using the threshold method to generate a rough retrieval time series result H, whose dimension is also T*1;
[0126] Specifically, the step of refining the precise boundary includes: based on the rough retrieval time series result H, training the precise boundary refinement network Φ pre , designing a multi-scale fine-tuning paradigm for the event boundary, and at the same time providing an explicit optimization constraint for the boundary.
[0127] The constraint loss function is as follows:
[0128]
[0129] where θ pre is a parameter of the precise boundary refinement network, (F v W , F q W , H W ) represents the distributions of the video feature map, the text feature map, and the retrieval time series pseudo-label, represents a video feature map instance, represents a text feature map instance, H i is its time series pseudo-label, and U represents the cross-entropy function. The precise boundary refinement network Φ pre outputs an accurate boundary retrieval result map K, whose dimension is T*1;
[0130] Specifically, the precise retrieval result generation step includes: generating a prediction result using a threshold method based on the precise boundary retrieval result graph K.
[0131] Example 2
[0132] The present invention also provides a local supervised long video temporal text retrieval system, which can be implemented by executing the process steps of the local supervised long video temporal text retrieval method. That is, those skilled in the art can understand the local supervised long video temporal text retrieval method as a preferred implementation manner of the local supervised long video temporal text retrieval system.
[0133] A local supervised long video temporal text retrieval system provided by the present invention includes:
[0134] Text feature extraction module: Using a text feature encoding network composed of fully connected layers, extract the initial text features of a preset dimension from the input retrieval text.
[0135] Video feature extraction module: Using a visual feature encoding network composed of 3D depth convolution, extract the initial video features of a preset dimension from the input long video.
[0136] Cross-modal feature fusion module: Using a cross-modal feature fusion network composed of multi-scale Transformers, perform deep fusion on the input initial text features and initial video features, and output a text feature map and a video feature map.
[0137] Event detection module: Using an event detection network composed of a multi-layer perceptron, map the video feature map to an event proposal described by text, which is composed of an event center and an event length. Furthermore, calculate an event temporal position mask based on the event proposal, representing the time position of the video frames containing the event, as the output.
[0138] Local position label supervision module: Constrain the event proposal to completely contain the local temporal label.
[0139] Event feature aggregation module: Apply the event temporal position mask to the video feature map to generate an event feature map, and generate event-level visual features through aggregation.
[0140] Multi-modal feature contrast module: Using feature contrast learning, encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space.
[0141] Coarse retrieval result generation module: Based on the event proposal, generate a coarse retrieval temporal result using a threshold method.
[0142] Accurate boundary refinement module: Using the rough retrieval time series results, train an accurate boundary refinement network, design a multi-scale fine-tuning paradigm for event boundaries, and at the same time provide explicit optimization constraints for the boundaries, and output an accurate boundary retrieval result map;
[0143] Accurate retrieval result generation module: Based on the accurate boundary retrieval result map, use the threshold method to generate prediction results.
[0144] Specifically, the text feature extraction module includes: tokenizing and encoding the retrieval text based on a tokenizer, and using a text feature encoding network composed of fully connected layers to map the tokenized encoding to a text initial feature map F' of dimension W*D q where W represents the number of words in the retrieval text, and D q represents the initial feature dimension of the text; q
[0145] Specifically, the video feature extraction module includes: using a visual feature encoding network composed of 3D depth convolution to map the input long video RGB frames to a video initial feature F' of dimension T*D v where T represents the time length of the video, and D v represents the initial feature dimension of the video; v
[0146] Specifically, the cross-modal feature fusion module includes: using a cross-modal feature fusion network Φ composed of a multi-scale Transformer architecture fus , to perform deep alignment and fusion on the input text initial feature F' q and video initial feature F' v , and output a text feature map F q and a video feature map F v . The calculation formula is as follows:
[0147] F v , F q = Φ fus (F' v , F' q ) (1)
[0148] where the dimension of F q is W*D, the dimension of F v is T*D, and D represents the fusion feature dimension;
[0149] Specifically, the event detection module includes: using an event detection network Φ composed of a multi-layer perceptron det , to map the video feature map F v to an event proposal P=(C, L) of text description, representing the time interval corresponding to the event described in the text in the long video. The calculation formula is as follows:
[0150] P = (C, L) = Φ det (F v ) (2)
[0151] Among them, C represents the predicted event center, and L represents the predicted event length. Furthermore, the plateau shape transformation is used to map the event proposal P to the event temporal position mask M as the output. M represents the temporal positions of the video frames containing the event, and the dimension is T * 1. Compared with predicting the frame-level event probability, predicting the event proposal only uses two variables, the center and the length, to characterize the event instance, so it has a smaller solution space. In addition, the frame-level event probability within the constraint interval of the event proposal is 1, and the probability outside the interval is 0, so it naturally has good probability smoothness.
[0152] Specifically, the local position label supervision module includes: according to the provided local temporal label Y, supervising that the event proposal P completely contains the local temporal label, calculating the loss function to train the event detection network until the loss function converges;
[0153] The calculation formula of the loss function is as follows:
[0154]
[0155] where θ det is the parameter of the event detection network, (F v W , Y W ) represents the distribution of the video feature map and the local temporal label, represents the video feature map instance, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function. Under the constraint of the loss function, the model retrieves events with the local annotation as the anchor frame to avoid false positive results. Sufficient annotation data encourages the model to expand the anchor frame to improve the retrieval results.
[0156] Specifically, the event feature aggregation module includes: applying the event temporal position mask M to the video feature map to generate an event feature map, and averaging in the temporal dimension to aggregate into the event-level visual feature f e . Applying the inverse of the event temporal position mask, 1 - M, to the video feature map to generate a background feature map, and averaging in the temporal dimension to aggregate into the background-level visual feature f b . Averaging the video feature map in the temporal dimension to aggregate into the video-level visual feature f a , and the dimensions of all three are 1 * D. The calculation formula is as follows:
[0157] f e = AVG(Fv *M)f b = AVG(F v *(1 - M))f a = AVG(F v ) (4)
[0158] Where AVG is the average function in the time dimension.
[0159] Specifically, the multimodal feature comparison module includes: averaging the text feature map F q in the word dimension to generate the text feature f q , whose dimension is 1*D. Calculate the cosine similarity for all text features in the dataset, and cluster all samples into several semantically similar clusters. Using feature contrast learning, encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space. The feature contrast loss function has the following calculation formula:
[0160]
[0161]
[0162]
[0163]
[0164] Where θ det are the parameters of the event detection network, and θ fus are the parameters of the cross-modal feature fusion network. represents the event-level visual feature instance, represents the video-level visual feature instance, represents the background-level visual feature instance, f e + represents the set of event visual features with similar semantics, f q + represents the set of text features with similar semantics, f e - represents the set of event visual features with dissimilar semantics, f q - represents the set of text features with dissimilar semantics, S(·,·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrast learning function, and τ represents the temperature coefficient. Under the constraint of the contrast loss function, video events and text are aligned and fused in the feature space to improve the retrieval effect.
[0165] Specifically, the rough retrieval result generation module includes: for the event proposal P, using the threshold method to generate a rough retrieval time series result H, whose dimension is also T*1;
[0166] Specifically, the precise boundary refinement module includes: based on the rough retrieval time series result H, training the precise boundary refinement network Φ pre , designing a multi-scale fine-tuning paradigm for the event boundary, and providing explicit optimization constraints for the boundary at the same time.
[0167] The constraint loss function is as follows:
[0168]
[0169] where θ pre are the parameters of the precise boundary refinement network, (F v W , F q W , H W ) represent the distributions of the video feature map, text feature map and retrieval time series pseudo-label, represents an instance of the video feature map, represents an instance of the text feature map, H i is its time series pseudo-label, and U represents the cross-entropy function. The precise boundary refinement network Φ pre outputs the precise boundary retrieval result map K, whose dimension is T*1;
[0170] Specifically, the precise retrieval result generation module includes: based on the precise boundary retrieval result map K, using the threshold method to generate a prediction result.
[0171] Example 3
[0172] Example 3 is a variant of Example 1.
[0173] Text feature extraction module: using a text feature encoding network composed of fully connected layers to extract the text initial features of a preset dimension from the input retrieval text;
[0174] Video feature extraction module: using a visual feature encoding network composed of 3D depth convolution to extract the video initial features of a preset dimension from the input long video;
[0175] Cross-modal feature fusion module: using a cross-modal feature fusion network composed of multi-scale Transformer to deeply fuse the input text initial features and video initial features, and output the text feature map and video feature map;
[0176] Event Detection Module: An event detection network composed of a multi-layer perceptron is used to map the video feature map into an event proposal described in text. The proposal consists of an event center and an event length. Furthermore, an event temporal position mask is calculated based on the event proposal to represent the temporal positions of the video frames containing the event, which is used as the output;
[0177] Local Position Label Supervision Module: Constrains the event proposal to completely contain the local temporal label;
[0178] Event Feature Aggregation Module: Applies the event temporal position mask to the video feature map to generate an event feature map, and generates event-level visual features through aggregation;
[0179] Multi-modal Feature Contrast Module: Uses feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space;
[0180] Rough Retrieval Result Generation Module: Based on the event proposal, uses a threshold method to generate a rough retrieval temporal result;
[0181] Precise Boundary Refinement Module: Uses the rough retrieval temporal result to train a precise boundary refinement network, designs a multi-scale fine-tuning paradigm for the event boundary, and provides explicit optimization constraints for the boundary, and outputs a precise boundary retrieval result map;
[0182] Precise Retrieval Result Generation Module: Based on the precise boundary retrieval result map, uses a threshold method to generate a prediction result.
[0183] Specifically, the local supervised long video temporal text retrieval framework composed of the text feature extraction module, video feature extraction module, cross-modal feature fusion module, event detection module, local position label supervision module, event feature aggregation module, multi-modal feature contrast module, rough retrieval result generation module, precise boundary refinement module, and precise retrieval result generation module is as Figure 2 shown.
[0184] In the system framework of the embodiment as Figure 2 shown, on the one hand, the retrieval text is input into the text feature extraction module, and the text initial feature map F q ' is output, whose dimension is W*D q , where W represents the number of words in the retrieval text, and D q represents the initial feature dimension of the text. On the other hand, the RGB corresponding to the video to be detected is input into the video feature extraction module, and the video initial feature F v ' is output, whose dimension is T*D v , where T represents the time length of the video, and D vRepresents the initial feature dimension of the video. The text feature extraction module is an encoding network composed of a series of fully connected layers, and existing network structures such as Bert can be used. The video feature extraction module is a downsampling encoding network composed of a series of 3D convolutional layers (+ batchnorm layer + relu layer), and existing network structures such as two-stream I3D, C3D, etc. can be used. The initial text feature F q ' and the initial video feature F v ' will be input into the cross-modal feature fusion module for deep alignment and fusion, and are mapped to the text feature map F q and the video feature map F v , whose dimensions are W*D and T*D respectively, and D represents the fusion feature dimension.
[0185] The video feature map F v will be further input into the event detection module and mapped to the event proposal P=(C, L) described in the text, which represents the time interval corresponding to the event described in the text in the long video. Where C represents the predicted event center and L represents the predicted event length. Then, the plateau shape transformation is used to map the event proposal P to the event temporal position mask M as the output. M represents the time positions of the video frames containing the event, and the dimension is T*1. Compared with predicting the frame-level event probability, predicting the event proposal only uses two variables, the center and the length, to characterize the event instance, so it has a smaller solution space. In addition, the frame-level event probability within the constraint interval of the event proposal is 1, and the probability outside the interval is 0, so it naturally has good probability smoothness. In order to supervise that the event proposal P completely contains the local temporal label, according to the provided local temporal label Y, the local position label supervision module calculates the loss function to train the event detection network until the loss function converges. The calculation formula of the loss function is as follows:
[0186]
[0187] Where θ det is the parameter of the event detection network, (F v W , Y W ) represents the distribution of the video feature map and the local temporal label, represents the video feature map instance, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function. Under the constraint of the loss function, the model performs event retrieval with the local annotation as the anchor frame to avoid false positive results. Sufficient annotation data encourages the model to expand the anchor frame to improve the retrieval results.
[0188] To encourage the aggregation of event-text features and event-event features with similar semantics in the representation space, and the dispersion of event-text features and event-event features with dissimilar semantics in the representation space, the event feature aggregation module calculates multimodal features, and the multimodal feature contrast module constructs a contrast loss function. Apply the event temporal position mask M to the video feature map to generate an event feature map, and average it in the time dimension to aggregate into event-level visual feature f e Apply the inverse of the event temporal position mask, 1-M, to the video feature map to generate a background feature map, and average it in the time dimension to aggregate into background-level visual feature f b Average the video feature map in the time dimension to aggregate into video-level visual feature f a , and the dimensions of all three are 1*D. For the text feature map F q Average it in the word dimension to generate text feature f q , whose dimension is 1*D. Calculate the cosine similarity for all text features in the dataset, and cluster all samples into several semantically similar clusters. Use feature contrast learning. The loss function is calculated as follows:
[0189]
[0190]
[0191]
[0192]
[0193] where θ det are the parameters of the event detection network, and θ fus are the parameters of the cross-modal feature fusion network. represents an event-level visual feature instance, represents a video-level visual feature instance, represents a background-level visual feature instance, f e + represents a set of event visual features with similar semantics, f q + represents a set of text features with similar semantics, f e - represents a set of event visual features with dissimilar semantics, f q - represents a set of text features with dissimilar semantics, S(·,·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrast learning function, and τ represents the temperature coefficient. Under the constraint of the contrast loss function, video events and texts are aligned and fused in the feature space to improve the retrieval effect.
[0194] Based on event proposal P, a rough retrieval time series result H with dimension T*1 can be generated using the threshold method.
[0195] Considering the incompleteness of local supervision, in order to refine the retrieval results, a boundary refinement learning module is additionally introduced to further fine-tune the retrieval results, so as to obtain a more accurate and complete video time series retrieval effect. The constraint loss function of the boundary refinement learning module is as follows:
[0196]
[0197] where θ pre are the parameters of the precise boundary refinement network, (F v W , F q W , H W ) represents the distributions of the video feature map, text feature map, and retrieval time series pseudo-label. represents an instance of the video feature map. represents an instance of the text feature map, and H i is its time series pseudo-label. U represents the cross-entropy function. The precise boundary refinement network Φ pre outputs the precise boundary retrieval result map K with dimension T*1.
[0198] After the overall framework training is completed, for the predicted precise boundary retrieval result map K, the threshold method is used to generate the final prediction result.
[0199] In summary, aiming at the defects of the existing supervision setting, the present invention first proposes a local supervision setting to balance the annotation cost and model performance, that is, for each text description in the training set, during the event occurrence stage in the long video, several consecutive seconds (even one frame) are randomly selected as the event local time series annotation. As an intermediate setting between strong and weak supervision, local supervision not only maintains a low annotation cost but also provides an accurate retrieval position anchor, laying a solid performance foundation. For the specific setting of local supervision, the present invention provides a method and system for local supervision long video time series text retrieval based on a contrast-unified strategy. Specifically, a quadruple contrast learning strategy is carefully designed to fully optimize the multi-modal features. By aggregating event features, background features, and text features, the present invention encourages semantic consistency within samples, between samples, within video modalities, and between video modalities, that is, similar semantics are aggregated in the feature space and dissimilar semantics are dispersed. Therefore, video events and texts are aligned and fused in the feature space to improve the retrieval effect. Considering the incompleteness of local supervision, the present invention designs a learning strategy from coarse to fine to refine the retrieval results. After generating the rough retrieval results, a boundary refinement learning module is additionally introduced to further fine-tune the retrieval results, so as to obtain a more accurate and complete video time series retrieval effect.
[0200] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1 and Embodiment 2.
[0201] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a kind of hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as software modules for implementing the method and the structures within the hardware component.
[0202] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A local supervised long video temporal text retrieval method, characterized in that Including: Text feature extraction step: Using a text feature encoding network composed of fully connected layers, extract text initial features of a preset dimension from the input retrieval text; Video feature extraction step: Using a visual feature encoding network composed of 3D depth convolution, extract video initial features of a preset dimension from the input long video; Cross-modal feature fusion step: Using a cross-modal feature fusion network composed of multi-scale Transformers, deeply fuse the input text initial features and video initial features, and output a text feature map and a video feature map; Event detection step: Using an event detection network composed of a multi-layer perceptron, map the video feature map to an event proposal described in text. This event proposal consists of an event center and an event length. Furthermore, calculate an event temporal position mask based on the event proposal to represent the time positions of the video frames containing the event, and use it as the output; Local position label supervision step: Constrain the event proposal to completely contain the local temporal label; Event feature aggregation step: Apply the event temporal position mask to the video feature map to generate an event feature map, and generate event-level visual features through aggregation; Multi-modal feature contrast step: Use feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space; Rough retrieval result generation step: Based on the event proposal, use a threshold method to generate a rough retrieval temporal result; Precise boundary refinement step: Use the rough retrieval temporal result to train a precise boundary refinement network, design a multi-scale fine-tuning paradigm for the event boundary, and at the same time provide an explicit optimization constraint for the boundary, and output a precise boundary retrieval result map; Precise retrieval result generation step: Based on the precise boundary retrieval result map, use a threshold method to generate a prediction result; The local position label supervision step constrains the event proposal P to completely contain the local temporal label Y according to the provided local temporal label Y. The constraint calculation formula is as follows: where θ det is a parameter of the event detection network, representing the distribution of the video feature map and the local temporal label, representing the video feature map instance, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function.
2. The method for local supervised long video sequential text retrieval according to claim 1, wherein The text feature extraction step tokenizes and encodes the retrieved text based on a tokenizer, and uses a text feature encoding network composed of fully connected layers to map the tokenized encoding to a text initial feature map F' of dimension W*D q where W represents the number of words in the retrieved text, and D q represents the initial feature dimension of the text; q The video feature extraction step uses a visual feature encoding network composed of 3D depth convolution to map the video RGB frames into the video initial feature F' of T*D v dimensions; v where T represents the time length of the video, and D v represents the initial feature dimension of the video. The cross-modal feature fusion step: Use the cross-modal feature fusion network Φ composed of multi-scale Transformers fus , to perform deep fusion on the input text initial feature F' q and the video initial feature F' v , and output the text feature map F q and the video feature map F v , and the calculation formula is as follows: F v ,F q = Φ fus (F' v ,F' q ) Among them, F q has a dimension of W * D, and F v has a dimension of T * D, where D represents the fused feature dimension.
3. The local supervised long video sequential text retrieval method according to claim 1, wherein The event detection step uses an event detection network Φ composed of a multi-layer perceptron det , which maps the video feature map F v to an event proposal P=(C, L) of text description, and the calculation formula is as follows: P = (C, L) = Φ det (F v ) Where C represents the predicted event center, L represents the predicted event length, and then use a plateau shape transformation to map the event proposal P to an event temporal position mask M as the output. M represents the time positions of the video frames containing the event, and the dimension is T*1.
4. The method for local supervised long video sequential text retrieval according to claim 1, wherein The event feature aggregation step applies the event temporal position mask M to the video feature map to generate an event feature map, and averages it in the time dimension to aggregate it into the event-level visual feature f e applies the inverse of the event temporal position mask, 1-M, to the video feature map to generate a background feature map, and averages it in the time dimension to aggregate it into the background-level visual feature f b averages the video feature map in the time dimension to aggregate it into the video-level visual feature f a The dimensions of all three are 1*D; The calculation formula is as follows: f e = AVG(F v * M)f b = AVG(F v *(1 - M))f a = AVG(F v ) Where AVG is the average function in the time dimension; The multi-modal feature comparison step: Average the text feature map F q in the word dimension to generate the text feature f q , whose dimension is 1*D. Calculate the cosine similarity for all text features in the dataset, cluster all samples into several semantically similar clusters, and use feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, while event-text features and event-event features with dissimilar semantics are dispersed in the representation space. The feature contrast loss function is calculated as follows: where θ det is a parameter of the event detection network, and θ fus is a parameter of the cross-modal feature fusion network, represents an event-level visual feature instance, represents a video-level visual feature instance, represents a background-level visual feature instance, represents a set of event visual features with similar semantics, represents a set of text features with similar semantics, represents a set of event visual features with dissimilar semantics, represents a set of text features with dissimilar semantics, S(·,·) represents the cosine similarity function, MAX represents the function of taking the maximum value, G represents the contrastive learning function, and τ represents the temperature coefficient; The rough retrieval result generation step generates a rough retrieval temporal result H with a dimension of T*1 for the event proposal P using a threshold method; The precise boundary refinement step is based on the rough retrieval time series result H to train the precise boundary refinement network Φ pre and design a multi-scale fine-tuning paradigm for the event boundary, while providing explicit optimization constraints for the boundary; The constraint loss function is as follows: where θ pre is a parameter of the precise boundary refinement network, represents the distributions of the video feature map, the text feature map, and the retrieval temporal pseudo-label, represents an instance of the video feature map, represents an instance of the text feature map, H i is its temporal pseudo-label, U represents the cross-entropy function, and the precise boundary refinement network Φ pre outputs the precise boundary retrieval result map K, whose dimension is T*1; The precise retrieval result generation step generates a prediction result based on the precise boundary retrieval result map K using a threshold method.
5. A local supervised long video temporal text retrieval system, characterized in that, Including: Text feature extraction module: Using a text feature encoding network composed of fully connected layers, extract text initial features of a preset dimension from the input retrieval text; Video feature extraction module: Using a visual feature encoding network composed of 3D depth convolution, extract video initial features of a preset dimension from the input long video; Cross-modal Feature Fusion Module: A cross-modal feature fusion network composed of multi-scale Transformers is used to deeply fuse the input initial text features and initial video features, and output text feature maps and video feature maps; Event Detection Module: An event detection network composed of multi-layer perceptrons is used to map the video feature maps into event proposals described in text. The event proposals consist of an event center and an event length. Furthermore, an event temporal position mask is calculated based on the event proposals to represent the temporal positions of the video frames containing the event, and used as the output; Local Position Label Supervision Module: Constrains the event proposals to completely contain local temporal labels; Event Feature Aggregation Module: Applies the event temporal position mask to the video feature maps to generate event feature maps, and aggregates them to generate event-level visual features; Multi-modal Feature Contrast Module: Uses feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space; Rough Retrieval Result Generation Module: Based on the event proposals, uses a threshold method to generate rough retrieval temporal results; Precise Boundary Refinement Module: Uses the rough retrieval temporal results to train a precise boundary refinement network, designs a multi-scale fine-tuning paradigm for the event boundaries, and provides explicit optimization constraints for the boundaries, and outputs a precise boundary retrieval result map; Precise Retrieval Result Generation Module: Based on the precise boundary retrieval result map, uses a threshold method to generate prediction results; The Local Position Label Supervision Module constrains the event proposal P to completely contain the local temporal label Y according to the provided local temporal label Y. The constraint calculation formula is as follows: where θ det is a parameter of the event detection network, representing the distribution of the video feature map and the local temporal label, representing the video feature map instance, y i is its local temporal label, ||·|| represents the absolute value function, Φ det represents the event detection network, and MAX represents the maximum value function.
6. The local supervised long video temporal text retrieval system according to claim 5, wherein The text feature extraction module tokenizes and encodes the retrieved text based on a tokenizer, and uses a text feature encoding network composed of fully connected layers to map the tokenized encoding to a text initial feature map F' of dimension W*D q ; where W represents the number of words in the retrieved text, and D q represents the initial feature dimension of the text; q The video feature extraction module uses a visual feature encoding network composed of 3D depth convolution to map the video RGB frames into the initial video features F' of T*D v dimensions v ; where T represents the time length of the video, and D v represents the dimension of the initial video features; The cross-modal feature fusion module: uses a cross-modal feature fusion network Φ composed of multi-scale Transformers fus , to deeply fuse the input text initial feature F' q and the video initial feature F' v and outputs the text feature map F q and the video feature map F v , and the calculation formula is as follows: F v , F q = Φ fus (F' v , F' q ) Among them, F q has a dimension of W*D, and F v has a dimension of T*D, where D represents the dimension of the fused feature.
7. The local supervised long video temporal text retrieval system according to claim 5, characterized in that, The event detection module uses an event detection network Φ composed of a multi-layer perceptron det , which maps the video feature map F v to an event proposal P=(C, L) of a text description, and the calculation formula is as follows: P = (C, L) = Φ det (F v ) Where C represents the predicted event center, L represents the predicted event length, and then the event proposal P is mapped into an event temporal position mask M as the output using a plateau shape transformation. M represents the temporal positions of the video frames containing the event, and the dimension is T*1.
8. The local supervised long video sequential text retrieval system according to claim 5, wherein The event feature aggregation module applies the event temporal position mask M to the video feature map to generate an event feature map, and averages it in the time dimension to aggregate it into an event-level visual feature f e applies the inverse of the event temporal position mask, 1-M, to the video feature map to generate a background feature map, and averages it in the time dimension to aggregate it into a background-level visual feature f b averages the video feature map in the time dimension to aggregate it into a video-level visual feature f a The dimensions of all three are 1*D, and the calculation formula is as follows: f e = AVG(F v * M)f b = AVG(F v *(1 - M))f a = AVG(F v ) Where AVG is the average function in the time dimension; The multimodal feature comparison module: averages the text feature map F q in the word dimension to generate the text feature f q , whose dimension is 1*D. Calculate the cosine similarity for all text features in the dataset, cluster all samples into several semantically similar clusters, and use feature contrast learning to encourage event-text features and event-event features with similar semantics to aggregate in the representation space, and event-text features and event-event features with dissimilar semantics to disperse in the representation space. The feature contrast loss function has the following calculation formula: where θ det is the parameter of the event detection network, and θ fus is the parameter of the cross-modal feature fusion network, represents the event-level visual feature instance, represents the video-level visual feature instance, represents the background-level visual feature instance, represents the set of event visual features with similar semantics, represents the set of text features with similar semantics, represents the set of event visual features with dissimilar semantics, represents the set of text features with dissimilar semantics, S(·, ·) represents the cosine similarity function, MAX represents the maximum value function, G represents the contrastive learning function, and τ represents the temperature coefficient; The Rough Retrieval Result Generation Module uses a threshold method to generate rough retrieval temporal results H for the event proposal P, and its dimension is T*1; The precise boundary refinement module trains a precise boundary refinement network Φ based on the rough retrieval time series result H pre , designs a multi-scale fine-tuning paradigm for event boundaries, and provides explicit optimization constraints for the boundaries at the same time; The constraint loss function is as follows: where θ pre is a parameter of the precise boundary refinement network, representing the distributions of the video feature map, the text feature map, and the retrieval temporal pseudo-label, representing an instance of the video feature map, representing an instance of the text feature map, H i is its temporal pseudo-label, U represents the cross-entropy function, and the precise boundary refinement network Φ pre outputs the precise boundary retrieval result map K, whose dimension is T*1; The Precise Retrieval Result Generation Module generates prediction results based on the precise boundary retrieval result map K using a threshold method.
Citation Information
Patent Citations
Language description guided video timing sequence positioning method and system
CN112015947A
Weak supervision video clip retrieval method and system based on erasure mechanism
CN112685597A
Semi-supervised video paragraph positioning method based on average teacher model
CN114155477A
Efficient and fine-grained video retrieval
US20200302294A1