Video Moment Retrieval Method Based on Fine-Grained Information of Video Content

Through object detection network and part-of-speech annotation technology, a cross-modal feature fusion module is built, which solves the problem of failing to make full use of fine-grained video and text information in the existing technology, and achieves higher video moment retrieval accuracy and speed.

CN116450883BActive Publication Date: 2025-07-04XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310448759.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-04
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

Existing cross-modal video moment retrieval technology fails to make full use of fine-grained information in video and text, resulting in a decrease in retrieval accuracy and speed, making it difficult to accurately determine the relevant moments of video content and query statements.

Method used

The fine-grained information of the video is extracted through the object detection network, combined with the pre-trained word embedding model and part-of-speech annotation, a cross-modal feature fusion module and a word meaning matching module are constructed, correlation weights are generated, and timely search is guided.

Benefits of technology

It improves the accuracy and speed of video time search, reduces the search time, and significantly improves the detection accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116450883B_ABST
    Figure CN116450883B_ABST
Patent Text Reader

Abstract

A video moment retrieval method based on fine-grained information of video content includes the following steps: Step 1, construct a training set and a test set, and select the original video; Step 2, perform feature pre-extraction on the original video to obtain key frame features and in-frame objects; Step 3, construct a text feature extraction module, use a pre-trained word embedding model to map the query statement into the embedding space, complete feature extraction, and obtain text features; Step 4, construct a text part-of-speech tagging module to tag the nouns in the query statement; Step 5, construct a cross-modal feature fusion module to obtain cross-modal fine-grained content features; Step 6, construct a semantic matching module to generate relevance weights through semantic matching; Step 7, construct a moment retrieval guidance module to calculate the relevance content fine-grained features corresponding to the entire video. The present invention extracts fine-grained information in the video through an object detection network, constructs a cross-modal retrieval model, and improves the accuracy of video moment retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network retrieval, and particularly relates to a video moment retrieval method based on fine-grained information of video content. Background Art

[0002] In recent years, multi-modal data such as text, images, and videos has grown rapidly. It is difficult for users to effectively search for information of interest, and various search technologies have emerged. Traditional search technologies mainly perform retrieval within a single modality. For example, keyword-based retrieval mainly performs similarity searches on the content of a single modality. With the development of Internet technology and the popularity of smart phones, users' requirements for cross-modal data retrieval are getting higher and higher. Cross-modal video retrieval technology is one of the key technologies, which determines the start and end times of the time segment that best matches the description statement in a complete video through a query statement described in natural language. Cross-modal video retrieval not only needs to mine rich visual, text, and speech information contained in the video, but also needs to determine the content similarity between different modalities. The current cross-modal video retrieval technologies can be mainly divided into two categories: ranking-based methods and localization-based methods.

[0003] The core of the ranking-based method lies in ranking candidate segments. Its characteristics are simple implementation, easy to interpret and understand. Further, according to the process of generating candidate segments, it can be divided into a method of presetting candidate segments and a method of generating candidate segments in a guided manner. The former manually segments the video to generate candidate segments without query statement information, and then ranks them according to their relevance to the query statement. The latter is guided by the query statement or the video itself. First, the model is used to exclude most irrelevant candidate segments, and then the generated candidate segments are ranked. Most of the methods of generating candidate segments in a guided manner use weak supervised learning or reinforcement learning. This type of localization-based method does not take candidate video segments as the processing unit, but takes the entire video as the processing unit, and directly uses the segment time point as the prediction target. Due to the particularity and complexity of this task, there are still great deficiencies in the current cross-modal video moment retrieval technology, and the returned results are often not accurate enough, and the accuracy still cannot satisfy users.

[0004] The patent application with the publication number CN202011575231 and the name "Cross-modal Video Moment Retrieval Method Based on Cross-modal Dynamic Convolution Network" discloses a cross-modal video moment retrieval method based on a cross-modal dynamic convolution network. This method first constructs a network structure of a hierarchical video feature extraction module and a text feature extraction module based on the attention mechanism to extract the features of videos and texts respectively. Then, a cross-modal fusion mechanism is used to fuse the features of the two modalities. Finally, a moment positioning module based on a cross-modal convolutional neural network is used to complete the moment retrieval. This method uses the fused features and text features to dynamically generate convolution kernels and uses the moment positioning module based on a cross-modal convolutional neural network to complete the moment retrieval. However, the disadvantage of this method is that it does not fully extract the fine-grained information in videos and texts, and at the same time, it is unable to match the fine-grained information in videos and texts, resulting in a decrease in the accuracy and speed of retrieval.

[0005] When manually retrieving video moments, the most intuitive way to determine the video content is often to distinguish the objects in the video, match them with the objects in the query statement, and then determine whether the relevant actions in the video are associated with the query statement, so as to roughly determine the query moment position. This shows that the fine-grained information in the query data, such as which objects exist in the video and which objects are in the statement description, plays a key role in video moment retrieval. However, many existing video moment retrieval methods have defects in dealing with fine-grained content and often do not make good use of text information to help identify the objects and actions in the video. For a description statement of a video, it may contain some keywords that can help determine the objects and actions in the video and the fine-grained information. The lack of utilization of this information will cause the video moment retrieval model to not be able to better distinguish the information in the video content. Summary of the Invention

[0006] In order to overcome the above problems existing in the prior art, the purpose of the present invention is to provide a video moment retrieval method based on the fine-grained information of video content. By extracting the fine-grained information in the video through an object detection network and constructing a cross-modal retrieval model, the accuracy of video moment retrieval can be improved.

[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0008] A video moment retrieval method based on the fine-grained information of video content includes the following steps;

[0009] Step 1, select the Charades-STA dataset to construct a training set and a test set, and select the original video V;

[0010] Step 2, construct a video fine-grained information extraction module, use the YOLOv5 object detection network to pre-extract features from the original video V, and obtain the key frame features F of the original video V C and the objects O within the frame C ;

[0011] Step 3, construct a text feature extraction module, use a pre-trained word embedding model to map the query statement S into the embedding space, complete feature extraction, and obtain the text feature Q:

[0012] Step 4, construct a text part-of-speech tagging module to tag the nouns H in the query statement S;

[0013] Step 5, construct a cross-modal feature fusion module to fuse the video key frame F C features in Step 2 and the text feature Q in Step 3 to obtain the cross-modal fine-grained content feature F a ;

[0014] Step 6, construct a word meaning matching module, through the objects O within the frame in Step 2 c and the nouns H extracted from the query statement in Step 4, generate the correlation weight Y through word meaning matching;

[0015] Step 7, construct a moment retrieval guidance module through the cross-modal content fine-grained feature F a and the correlation weight Y to calculate the correlation content fine-grained feature F corresponding to the entire video A .

[0016] In the above Step 1, the Charades-STA dataset is constructed by time annotation based on the Charades dataset. The Charades dataset includes action categories, videos, and "query, video segment" pairs; in some videos, structured complex queries need to be made, that is, each query contains at least two clauses, and the time span of the "query, video segment" pair is less than half of the video length.

[0017] The specific content of Step 2 is as follows:

[0018] Step 2.1, sample the original video at equal intervals with an interval of τ frames. The total number of frames of the video is T, and the extracted key frame pictures are where n c is the total number of frames extracted;

[0019] Step 2.2, use the YOLOv5 object detection network to extract the key frame features F C and the objects O within the frame C .

[0020] Further, in step 2.2.1, the key-frame image C is fed into the YOLOv5 object detection network. The backbone network adopts CSPNet. By dividing the convolution into two stages and using cross-stage feature reuse and information fusion, the number of model parameters and computational complexity are reduced, and the speed and accuracy of the model are improved. Through the backbone network, a feature map M1 of size 19×19 is obtained;

[0021] In step 2.2.2, the feature map M1 is fed into the top-down feature pyramid structure to extract strong semantic features, and the feature map M2 is output through upsampling;

[0022] In step 2.2.3, the feature map M2 passes through the bottom-up feature pyramid structure to extract strong localization features, and the feature map M3 is output;

[0023] In step 2.2.4, the feature map M3 is used as the detection head of the three-layer convolutional block, and the object detection task is carried out by operating on features of three different scales; the objects within the frame contained in the network output frame g is the number of objects within the frame. At the same time, multi-scale key-frame features F1 C , F2 C , and F3 C are obtained at the output of the spatial pyramid pooling.

[0024] The specific content of step 3 is as follows:

[0025] In step 3.1, the GloVe pre-trained word embedding model is used to map the query statement S into the embedding space to complete the extraction of the text feature Q. The process of text feature Q extraction is as follows:

[0026]

[0027] where m is the number of words in the sentence, d q is the dimension of the extracted text feature, Q is the text feature, s is the specific query statement, and q is the specific text feature.

[0028] The specific content of step 4 is as follows:

[0029] In step 4.1, the NLTK is used to split the query statement S into individual words;

[0030] In step 4.2, the NLTK is used to construct a hidden Markov model. Through lemmatization and word sense disambiguation, the part of speech of each word is labeled, and the nouns H={H1,...,H u} in the query statement are extracted as the keywords for matching with the video content, where u is the number of nouns in the statement.

[0031] The specific content of step 5 is as follows:

[0032] The three content features F1 output by the YOLOv5 model in step 2 C , F2 C , and F3 C have sizes of 80×80×256, 40×40×512, and 20×20×1024 respectively, and the size of the text feature Q is m×d q , where m is the number of words in the query statement and d q is the text feature dimension;

[0033] Step 5.1, pad and align the text feature Q in the first dimension to change the size of Q to m'×d q , where m' > m;

[0034] Step 5.2, add a dimension to the text feature Q and replicate and expand it in the second dimension to transform the size of the text feature Q into m'×m'×d q text feature This process can be expressed by the following formula:

[0035]

[0036] Step 5.3, for the multi-scale key frame feature F1 c , use a pooling layer to transform the size of the content feature into feature This process formula is as follows:

[0037]

[0038] Step 5.4, for the expanded text feature respectively use three fully connected layers with an input size of d q and an output size of to perform dimensional transformation, turning the text feature Q into three features Q with sizes of i ';

[0039]

[0040] where FC() is the fully connected layer operation;

[0041] Step 5.5, fuse the content feature and the text feature using the Hadamard product;

[0042]

[0043] Step 5.6, concatenate the three fused features, and perform feature extraction and dimensional transformation on the concatenated features through a fully connected layer to further enhance the expression ability of the features, making the features more discriminative, and finally obtaining a feature with a length of d vA cross-modal fine-grained content feature F a , d v is the dimension corresponding to the fusion vector in the moment retrieval network. The above process formula is expressed as follows:

[0044]

[0045] where FC() is the fully connected layer operation.

[0046] Specifically, step 6 is as follows:

[0047] Step 6.1, calculate the cosine similarity between pairwise word vectors;

[0048]

[0049] where w1 and w2 are any two word vectors, and similarity(w1, w2) is the similarity.

[0050] Step 6.2, for the object phrases within the frame and the nouns in the sentence where g is the number of objects within the frame and u is the number of nouns in the sentence, calculate the correlation weight Y of the two groups of word keys by calculating the average similarity. The specific calculation formula is as follows:

[0051]

[0052] Specifically, step 7 is as follows:

[0053] Step 7.1, multiply the cross-modal content fine-grained feature F a calculated in step 5 by the correlation weight Y to obtain the correlation content fine-grained feature for guiding the moment retrieval network in the current i-th frame

[0054]

[0055] where n c is the number of key frames in the video.

[0056] Step 7.2, splice the correlation content fine-grained features of all key frames in the video to obtain the correlation content fine-grained feature F corresponding to the entire video A :

[0057]

[0058] where n c is the number of key frames in the video;

[0059] Step 7.3, the content fine-grained feature F AThrough the bidirectional gated recurrent unit, the starting position T of the moment positioning is obtained begin and the end position T of the moment positioning end .

[0060] An electronic device comprises a processor, a memory and a communication bus, wherein the processor and the memory communicate with each other via the communication bus;

[0061] Memory, used to store computer programs;

[0062] The processor is used to implement the above-mentioned video moment retrieval method based on fine-grained information of video content when executing the program stored in the memory.

[0063] Beneficial effects of the present invention:

[0064] The present invention fully extracts fine-grained features in the video and cross-modally matches key frames with query statements through part-of-speech tagging. A video moment retrieval model based on fine-grained information of video content is constructed. The object detection network is used to extract fine-grained features of the video, and a cross-modal information matching method is used to match query statement part-of-speech tags with video key frame objects.

[0065] The present invention extracts fine-grained information from a video, and reduces the retrieval time through key frame matching and similarity calculation, thereby achieving higher video moment retrieval accuracy.

[0066] Moreover, this model is highly portable and can significantly improve the detection accuracy of the model by fusing it with the existing model based on the anchor-free method. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0068] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0069] like Figure 1 As shown in the figure: the invention is implemented by a video moment retrieval model based on video fine-grained information. The video moment retrieval model based on video fine-grained information includes a video fine-grained information extraction module, a text feature extraction module, a text part-of-speech tagging module, a feature fusion module, a word meaning matching module and a moment retrieval guidance module. Figure 1 The present invention is described in further detail.

[0070] Step 1: Build training and test sets based on the video set and query dataset:

[0071] In this example, two datasets that are common and challenging in the field of video moment retrieval, the Charades-STA dataset, are selected. Among them, 70% of the dataset is used as the training set, and 30% of the dataset is used for the test set, and the data is guaranteed to be randomly assigned.

[0072] The Charades-STA dataset is constructed by time annotation based on the Charades dataset. The Charades dataset includes action categories, videos, and "query, video segment" pairs; in some videos, structured complex queries need to be made, that is, each query contains at least two clauses, and the time span of the "query, video segment" pair is less than half of the video length.

[0073] Step 2, construct a video fine-grained information extraction module, use a target detection network to pre-extract features from the original video V, and obtain the key frame features F C and in-frame objects O C :

[0074] In this example, the YOLOv5 target detection network is selected for in-frame feature extraction and object detection. This method is a single-stage target detection model, which can achieve efficient and accurate real-time target detection. The network consists of a backbone network, a feature extraction network, and a detection head, and is one of the current mainstream target detection algorithms.

[0075] Step 2.1, for the original video V, the total number of frames of the video is T, and the video is sampled at equal intervals according to the interval of τ frames. The key frames taken out are where n c is the total number of frames taken out.

[0076] Step 2.2, use the YOLOv5 network to extract the key frame features F C and in-frame objects O C .

[0077] Step 2.2.1 Send the key frame C into the target detection network. The backbone network uses CSPNet. By dividing the convolution into two stages and using cross-stage feature reuse and information fusion, the number of parameters and computational complexity of the model are reduced, and the speed and accuracy of the model are improved.

[0078] Step 2.2.2 The feature extraction network adopts a spatial pyramid pooling structure, which can extract features of different scales, so as to better adapt to target objects of different sizes.

[0079] Step 2.2.3 On the output of the feature extraction network, a detection head with a three-layer convolutional block is adopted to perform the target detection task by operating on features of three different scales.

[0080] Step 2.2.4 The network outputs the objects contained in the frame g is the number of in-frame objects. At the same time, multi-scale key frame features F1 are obtained from the output of spatial pyramid pooling C , F2 C , and F3 C .

[0081] Step 3: Construct a text feature extraction module. Use a pre-trained word embedding model to map the query statement S into the embedding space, complete feature extraction, and obtain the text feature Q:

[0082] In this example, the GloVe pre-trained word embedding model is selected. The GloVe model is a word vector representation model based on global word frequency statistics. This method needs to first construct a co-occurrence matrix, then obtain the approximate relationship between the word vector and the co-occurrence matrix, and finally construct a loss function according to the error of the word vector representation for learning. By learning the word vector, the GloVe model can capture the semantic relationship between words and extract the text feature Q corresponding to the query statement S.

[0083] Step 3.1: Use the GloVe pre-trained word embedding model to map the query statement into the embedding space and complete feature extraction. The process of text feature extraction is as follows:

[0084]

[0085] where m is the number of words in the sentence, d q is the dimension of the extracted text feature, and Q is the text feature.

[0086] Step 4: Construct a text part-of-speech tagging module to tag the nouns in the query statement;

[0087] In this example, nouns are the most meaningful part of the query statement. They can describe the object and content of the query, while other types of words, such as verbs and adjectives, etc., describe the attributes and behaviors of the object more. In addition, there are many irrelevant words, which increase the computational complexity and reduce the retrieval efficiency. By extracting the nouns in the query statement, the video moment retrieval model can be more accurately guided to locate the video segments related to the query statement.

[0088] Step 4.1: Use NLTK to split the query statement S into individual words;

[0089] Step 4.2: Use NLTK to construct a hidden Markov model, and through lemmatization and word sense disambiguation, tag the part of speech of each word. Extract the nouns H = {H1,..., H u} in the query statement as the keywords for matching with the video content, where u is the number of nouns in the statement;

[0090] Step 5, construct a cross-modal feature fusion module to fuse the video key-frame features in Step 2 and the text features in Step 3.

[0091] Here, features of different scales are fused with the text, which can not only improve the model's understanding of video content, but also make the model more general and adaptable to video moment retrieval tasks in different scenarios. Among the three content features F1 C , F2 C and F3 c output by the YOLOv5 model in Step 2, their sizes are 80×80×256, 40×40×512, and 20×20×1024 respectively, and the size of the text feature Q is m×d q , where m is the number of words in the query sentence, and d q is the text feature dimension.

[0092] Step 5.1, pad and align the text feature Q in the first dimension to change the size of Q to m'×d q , where m' > m.

[0093] Step 5.2, add a dimension to the text feature Q and replicate and expand it in the second dimension to transform the size of the text feature Q into m'×m'×d q . This process can be represented by the following formula:

[0094]

[0095] Step 5.3, for the content feature F i c , use a pooling layer to transform the size of the content feature into The formula for this process is shown in the following figure:

[0096]

[0097] Step 5.4, for the expanded text feature , respectively use three fully connected layers with an input size of d q and an output size of for dimensional transformation. The text feature Q is transformed into three features Q of size i ′.

[0098]

[0099] Step 5.5, fuse the content feature and the text feature using the Hadamard product.

[0100]

[0101] Step 5.6, for Obtain the pooled features through an adaptive average pooling layer with a dimension size of 32

[0102] Step 5.7, Concatenate the three fused features, and perform feature extraction and dimensional transformation on the concatenated features through a fully connected layer to further enhance the expression ability of the features and make the features more discriminative. Finally, obtain a cross-modal fine-grained content feature F of length d v where d a is the dimension corresponding to the fused vector in the moment retrieval network. The above process is expressed by the following formula: v

[0103]

[0104] Step 6, Construct a semantic matching module, and generate a correlation weight Y through semantic matching using the in-frame object O c in Step 2 and the nouns H extracted from the query statement in Step 4

[0105] In this example, the method used for semantic matching is the word vector similarity calculation method using the GloVe model in the gensim natural language processing library

[0106] Step 6.1, Calculate the cosine similarity between word vectors pairwise

[0107]

[0108] Step 6.2, For the in-frame object phrase and the nouns in the statement (where g is the number of in-frame objects and u is the number of nouns in the statement). In this example, the correlation weight Y between the two groups of words is calculated by calculating the average similarity. The specific calculation formula is as follows:

[0109]

[0110] Step 7, Construct a moment retrieval guidance module to calculate the correlation content fine-grained feature F corresponding to the entire video A :

[0111] In this example, the greater the correlation weight, the greater the contribution of the key frame to the entire video. By weighted calculation of the cross-modal content fine-grained feature and the correlation weight of the key frame, some irrelevant noise information can be suppressed, and the accuracy and robustness of the retrieval result can also be improved

[0112] Step 7.1, Multiply the cross-modal content fine-grained feature F a calculated in Step 5 by the correlation weight Y to obtain the correlation content fine-grained feature for guiding the moment retrieval network of the current frame:​

[0113] Step 7.2, splice the fine-grained features of the relevant content of all key frames in the video to obtain the fine-grained features F of the relevant content corresponding to the entire video A :

[0114]

[0115] Step 8, construct a model experiment verification module to verify the moment retrieval guidance effect of the model and the ablation experiment of the model

[0116] In this example, the IoU metric is used as the evaluation metric to calculate the intersection over union of the predicted time and the true event. Specifically, it is expressed as R@n, IoU@m, where n = 1, m ∈ {0.3, 0.5, 0.7}. To verify the generality and effectiveness of this method, this method is migrated to the mainstream video moment retrieval method to verify the improvement ability of the network performance. Specifically, the fine-grained features of the relevant content are fused with the original model at the end of the model, and the guidance of the original model can be completed

[0117] Step 8.1, select the DRN, TMLGA, and VSLNet models for experimental verification. The parameters in the specific network are all kept the same as the original method, and the experimental results of the original model and the corresponding migration and fusion of this model (original model Pro) are compared. The experimental results are shown in the following table. It can be seen that the accuracy after fusing the models is better than the original models

[0118]

[0119] Step 8.2, to verify the effectiveness and necessity of the model operation, an ablation experiment is carried out on Charades-STA. And it is stipulated that w / o WM is to remove the word sense matching part, w / o FF is to remove the feature fusion part, w / o TC is not to add text features, w / o pool is to remove the pooling operation of the key frame features, and directly fuse after extending the text features. W / add FF is to fuse directly using addition instead of dot product. The results of the ablation experiment are as follows

[0120]

[0121]

Claims

1. A video moment retrieval method based on fine-grained information of video content, characterized in that, It includes the following steps; Step 1, select the Charades-STA dataset to construct the training set and the test set, and select the original video V; Step 2: Construct a video fine-grained information extraction module, and use the YOLOv5 object detection network to pre-extract features from the original video V to obtain the key frame features F of the original video V C and the in-frame objects O C ; Step 3, construct a text feature extraction module, use a pre-trained word embedding model to map the query statement S into the embedding space, complete feature extraction, and obtain the text feature Q: Step 4, construct a text part-of-speech tagging module to tag the nouns H in the query statement S; Step 5, construct a cross-modal feature fusion module to fuse the video key frame features F C in Step 2 and the text features Q in Step 3 to obtain cross-modal fine-grained content features F a ; Step 6, construct a word meaning matching module, and through the in-frame object O in Step 2 c and the noun H extracted from the query statement in Step 4, generate a correlation weight Y through word meaning matching; Step 7, through the cross-modal content fine-grained feature F a and the correlation weight Y, construct a moment retrieval guidance module to calculate the correlation content fine-grained feature F corresponding to the entire video A .

2. The video moment retrieval method based on fine-grained information of video content according to claim 1, characterized in that, In the above Step 1, the Charades-STA dataset is constructed by time annotation based on the Charades dataset. The Charades dataset includes action categories, videos, and "query, video segment" pairs; in some videos, structured complex queries need to be made, that is, each query contains at least two clauses, and the time span of the "query, video segment" pair is less than half of the video length.

3. The video moment retrieval method based on fine-grained information of video content according to claim 1, wherein The specific content of Step 2 is as follows: Step 2.1, perform equidistant sampling on the original video at intervals of τ frames. The total number of frames of the video is T, and the extracted key-frame images are where n c is the total number of frames extracted; Step 2.2, use the YOLOv5 object detection network to extract the key frame feature F C and the object O within the frame C .

4. The video moment retrieval method based on fine-grained information of video content according to claim 3, characterized in that Step 2.2.1, send the key-frame image C into the YOLOv5 object detection network. The backbone network adopts CSPNet. By dividing the convolution into two stages and using cross-stage feature reuse and information fusion to reduce the number of model parameters and computational complexity, the speed and accuracy of the model are improved; obtain the feature map M1 through the backbone network; Step 2.2.2, send the feature map M1 into the top-down feature pyramid structure to extract strong semantic features, and output the feature map M2 through upsampling; Step 2.2.3, send the feature map M2 through the bottom-up feature pyramid structure to extract strong localization features, and output the feature map M3; Step 2.2.4, use the feature map M3 as the detection head of the three-layer convolutional block to perform the object detection task by operating on features of three different scales; the objects within the frame contained in the network output frame g is the number of objects within the frame, and at the same time, multi-scale key frame features are obtained at the output of the spatial pyramid pooling and 5. The video moment retrieval method based on fine-grained information of video content according to claim 3, wherein The specific content of Step 3 is as follows: Step 3.1, use the GloVe pre-trained word embedding model to map the query statement S into the embedding space, complete the extraction of the text feature Q, and the process of text feature Q extraction is expressed as follows: where m is the number of words in the sentence, d q is the dimension of the extracted text features, Q is the text feature, s is the specific query statement, and q is the specific text feature.

6. The video moment retrieval method based on fine-grained information of video content according to claim 5, wherein, The specific content of Step 4 is as follows: Step 4.1, use NLTK to split the query statement S into individual words; Step 4.2, use NLTK to build a hidden Markov model, perform lemmatization and word sense disambiguation to label the part of speech of each word, and extract the nouns H = {H1,..., H u} in the query statement as keywords for matching with the video content, where u is the number of nouns in the statement.

7. The video moment retrieval method based on fine-grained information of video content according to claim 6, wherein The specific content of Step 5 is as follows: In step 2, the size of the three content feature text features Q output by the YOLOv5 model is m×d q , where m is the number of words in the query statement, and d q is the text feature dimension; Step 5.1, pad and align the text feature Q in the first dimension to change the size of Q to m'×d q , where m' > m; Step 5.2, increase the dimension of the text feature Q and perform replication expansion in the second dimension to transform the size of the text feature Q into m'×m'×d q of the text feature This process can be expressed by the following formula: Step 5.3, the multi-scale key frame feature F i c , use the pooling layer to convert the size of the content feature into feature of The process formula is as follows: Step 5.4, for the expanded text features respectively use three fully connected layers with an input size of d q , and an output size of to perform dimensional transformation, and transform the text feature Q into three features Q with a size of i '; Among them, FC() is a fully connected layer operation; Step 5.5, fuse the content feature and the text feature using the Hadamard product; Step 5.6, splice the three fused features, and perform feature extraction and dimensionality transformation on the spliced features through a fully connected layer to further enhance the expression ability of the features, making the features more discriminative, and finally obtaining a cross-modal fine-grained content feature F of length d v where d a is the dimension corresponding to the fused vector in the moment retrieval network. The above process is expressed by the following formula: v ​ Among them, FC() is a fully connected layer operation.

8. The video moment retrieval method based on fine-grained information of video content according to claim 7, wherein, The specific content of Step 6 is as follows: Step 6.1, calculate the cosine similarity between pairwise word vectors; where w1 and w2 are any two word vectors, and similarity(w1, w2) is the similarity; Step 6.2, for the intra-frame object phrases and the nouns in the sentence where g is the number of intra-frame objects and u is the number of nouns in the sentence. The correlation weight Y of the two groups of word keys is calculated by calculating the average similarity. The specific calculation formula is as follows:

9. The method for retrieving video moments based on fine-grained information of video content according to claim 8, wherein The specific content of Step 7 is as follows: Step 7.1, multiply the cross-modal content fine-grained feature F obtained in Step 5 a by the correlation weight Y to obtain the correlation content fine-grained feature for guiding the moment retrieval network at the current i-th frame where n c is the number of key frames in the video; Step 7.2, splice the fine-grained features of the relevant content of all key frames in the video to obtain the fine-grained features F of the relevant content corresponding to the entire video A : where n c is the number of key frames in the video; Step 7.3, take the content fine-grained feature F A through a bidirectional gated recurrent unit to obtain the start position T of the moment positioning begin and the end position T of the moment positioning end .

10. An electronic device, characterized in that, It includes a processor, a memory, and a communication bus. Among them, the processor and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor, when executing the programs stored on the memory, implements the video moment retrieval method based on fine-grained video content information described in any one of the above claims 1-9.

Citation Information

Patent Citations

  • A method for cross-modal video time-retrieval based on cross-modal dynamic convolutional networks

    CN112650886B

  • Cross-modal video moment retrieval method based on cross-modal dynamic convolutional network

    CN112650886A

  • Frame-level fine-grained natural language video moment positioning method

    CN115935001A