Multi-granularity Video Retrieval Method and Device

Through the multi-grained video retrieval method, combined with sentence-level text features and video features, the multi-center and multi-scale dual-branch collaborative processing is carried out, which solves the problem of insufficient recall and positioning accuracy in video retrieval, and achieves more efficient video and clip-level retrieval.

CN117194710BActive Publication Date: 2025-07-11UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311228436.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-21
Publication Date
2025-07-11
Estimated Expiration
2043-09-21

AI Technical Summary

Technical Problem

The existing video retrieval methods lack recall effect and target clip positioning accuracy in massive video data, especially in long video retrieval, the semantic alignment of the query text input by the user and the video clip is low, resulting in low retrieval accuracy.

Method used

The multi-grained video retrieval method is used to extract sentence-level features of the query text, and combine coarse and fine-grained features in the video library, and use multi-center and multi-scale dual-branch collaborative feature processing to calculate the similarity between video and text, so as to achieve video-level and fragment-level retrieval.

Benefits of technology

It improves the recall rate of video retrieval and the positioning accuracy of target clips, improves the overall accuracy of search, and can more accurately locate related clips in the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117194710B_ABST
    Figure CN117194710B_ABST
Patent Text Reader

Abstract

An embodiment of the present application proposes a multi-granularity video retrieval method and device, belonging to the field of cross-modal content retrieval. Through this retrieval algorithm, based on the sentence-level text features of the text to be queried, the coarse-grained video features and fine-grained video features of each video data in the video library, multi-center and multi-scale dual-branch collaborative feature processing is performed to obtain the similarity data between the text to be queried and each video data, so as to obtain the retrieval results of the overall-level video corresponding to video-level retrieval and the segment-level video corresponding to segment-level retrieval. The retrieval algorithm adopts a dual-branch collaborative strategy, designs a coarse-grained browsing branch and a fine-grained gazing branch, adopts a collaborative retrieval strategy based on focus guidance for the browsing branch and the gazing branch, and introduces a hybrid collaborative contrast learning strategy, which significantly improves the retrieval recall rate of the complete video under weak supervision conditions and the positioning accuracy of the target segment in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cross-modal content retrieval, and more particularly, to a multi-granularity video retrieval method and apparatus. Background Art

[0002] With the development of Internet technology, videos have gradually become a mainstream information medium, and the generation and consumption of video data have shown explosive growth. Against this background, how to effectively retrieve the content that users are interested in from a large number of videos has become an important and challenging problem.

[0003] Currently, the commonly used video retrieval methods include: (1) text-based video retrieval; (2) focusing on video retrieval and video clip retrieval. The first retrieval method is to retrieve videos that are semantically related to the given natural language text from a video library. The second retrieval method is to implement video retrieval and clip retrieval in stages. However, these methods all have the problem of low upper limit of retrieval recall effect. Summary of the Invention

[0004] In view of this, the purpose of the present application is to provide a multi-granularity video retrieval method and apparatus, which can improve the retrieval recall rate and the positioning accuracy of target clips in videos while realizing video-level retrieval and clip-level retrieval.

[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a multi-granularity video retrieval method, the method comprising:

[0007] Processing the text to be queried to obtain sentence-level text features corresponding to the text to be queried;

[0008] Obtaining feature information of each video data in a video library; wherein, the feature information includes coarse-grained video features and fine-grained video features, and the coarse-grained video features are obtained by downsampling the fine-grained video features;

[0009] Inputting the sentence-level text features into a pre-trained retrieval algorithm;

[0010] Through the retrieval algorithm, based on the sentence-level text features and the feature information, performing multi-center and multi-scale dual-branch collaborative feature processing to obtain similarity data between the text to be queried and each video data; wherein, the similarity data includes coarse-grained similarity and fine-grained similarity;

[0011] Obtaining a retrieval result according to the similarity data; wherein, the retrieval result includes an overall-level video corresponding to video-level retrieval and a clip-level video corresponding to clip-level retrieval.

[0012] In a possible implementation, the retrieval algorithm includes a browsing branch and a gazing branch;

[0013] The step of performing multi-center and multi-scale dual-branch collaborative feature processing on the sentence-level text features and the feature information through the retrieval algorithm to obtain the similarity data between the text to be queried and each video data includes:

[0014] Through the browsing branch, based on a plurality of selected center points and a plurality of scales, a plurality of coarse-grained candidate segments of each video data are constructed, and in combination with the coarse-grained video features and the sentence-level text features, the coarse-grained optimal segment is obtained from the plurality of coarse-grained candidate segments, and the coarse-grained similarity between the text to be queried and each video data is calculated;

[0015] Through the gazing branch, according to the center point and a plurality of scales of the coarse-grained optimal segment, a plurality of fine-grained candidate segments of each video data are constructed, and in combination with the fine-grained video features and the sentence-level text features, the fine-grained optimal segment is obtained from the plurality of fine-grained candidate segments, and the fine-grained similarity between the text to be queried and each video data is calculated.

[0016] In a possible implementation, the method further includes the step of training to obtain the retrieval algorithm, including:

[0017] Processing each sample pair in the training data set to obtain the sentence-level text features corresponding to the query text samples in each sample pair, and the frame-level features of the video samples in each sample pair; wherein, the query text sample in the sample pair is a natural language description of a segment of the video sample in the sample pair;

[0018] Selecting a preset number of sample pairs from the training data set as training samples, and inputting each training sample pair into the initial retrieval algorithm; wherein, the initial retrieval algorithm includes an initial Transformer model, an initial browsing branch, and an initial gazing branch, and the training sample pair includes a training video and a training query text;

[0019] Based on the initial Transformer model, processing the frame-level features to obtain fine-grained video sample features and coarse-grained video sample features;

[0020] Through the initial browsing branch, based on a plurality of center points and a plurality of scales, construct a plurality of coarse-grained sample candidate segments of the training video, and combine the coarse-grained video sample features and sentence-level text features of the training sample pair to obtain an optimal coarse-grained sample segment from the plurality of coarse-grained sample candidate segments, and calculate the coarse-grained similarity between the training video and the training query text;

[0021] Through the initial gaze branch, according to the center point and a plurality of scales of the optimal coarse-grained sample segment, construct a plurality of fine-grained sample candidate segments of the training video, and combine the fine-grained video sample features and sentence-level text features of the training sample pair to obtain an optimal fine-grained sample segment from the plurality of fine-grained sample candidate segments, and calculate the fine-grained similarity between the training video and the training query text;

[0022] Based on the coarse-grained similarity and the fine-grained similarity, combine the training videos in all the training samples to calculate a first contrastive learning loss for the coarse-grained and a second contrastive learning loss for the fine-grained;

[0023] Combine the first contrastive learning loss and the second contrastive learning loss to obtain a mixed collaborative contrastive learning loss. Based on the mixed collaborative contrastive learning loss, use an optimization algorithm to update the parameters of the initial retrieval algorithm to obtain a mature retrieval algorithm.

[0024] In a possible implementation manner, the step of combining the coarse-grained video sample features and sentence-level text features of the training sample pair to obtain an optimal coarse-grained sample segment from the plurality of coarse-grained sample candidate segments and calculating the coarse-grained similarity between the training video and the training query text includes:

[0025] For each of the coarse-grained sample candidate segments, perform Gaussian weighted pooling aggregation by combining the coarse-grained video sample features and the center point and width of the coarse-grained sample candidate segment to obtain the segment feature of the coarse-grained sample candidate segment;

[0026] Based on the sentence-level text feature and the segment feature, calculate the cosine similarity between each of the coarse-grained sample candidate segments and the training query text, take the coarse-grained sample candidate segment with the maximum cosine similarity as the optimal coarse-grained sample segment, and take the cosine similarity of the optimal coarse-grained sample segment as the coarse-grained similarity between the training video and the training query text.

[0027] In a possible implementation, the step of obtaining the optimal fine-grained sample segment from the multiple fine-grained sample candidate segments by combining the fine-grained video sample features and the sentence-level text features of the training sample pair, and calculating the fine-grained similarity between the training video and the training query text includes:

[0028] For each of the fine-grained sample candidate segments, perform Gaussian weighted pooling aggregation by combining the fine-grained video sample features and the center point and width of the fine-grained sample candidate segment to obtain the segment feature of the fine-grained sample candidate segment;

[0029] Based on the sentence-level text feature and the segment feature, calculate the cosine similarity between each of the fine-grained sample candidate segments and the training query text, take the fine-grained sample candidate segment with the maximum cosine similarity as the optimal fine-grained sample segment, and take the cosine similarity of the optimal fine-grained sample segment as the fine-grained similarity between the training video and the training query text.

[0030] In a possible implementation, the step of calculating the first contrastive learning loss for the coarse-grained and the second contrastive learning loss for the fine-grained by combining all the training videos in all the training samples based on the coarse-grained similarity and the fine-grained similarity includes:

[0031] According to the coarse-grained similarity, determine the first positive sample and the first negative sample from all the training videos in all the training samples, and jointly calculate the first contrastive learning loss for the coarse-grained by using the triplet loss function and the infoNCE loss function;

[0032] Based on the fine-grained similarity, determine the second positive sample, the second negative sample, the first type of negative sample, and the second type of hard negative sample from all the training videos in all the training samples, and jointly calculate the second contrastive learning loss for the fine-grained by using the triplet loss function and the infoNCE loss function.

[0033] In a possible implementation, the step of determining the second positive sample, the second negative sample, the first type of negative sample, and the second type of hard negative sample from all the training videos in all the training samples based on the fine-grained similarity, and jointly calculating the second contrastive learning loss for the fine-grained by using the triplet loss function and the infoNCE loss function includes:

[0034] Based on the fine-grained similarity, determine the second positive sample and the second negative sample from the multiple fine-grained sample candidate segments;

[0035] From the multiple fine-grained sample candidate segments, select the fine-grained sample candidate segment with the maximum fine-grained similarity as the optimal fine-grained sample segment;

[0036] Use the fine-grained sample candidate segments at both ends of the optimal segment of the fine-grained sample as one type of negative sample, and use the training video as the second type of hard negative sample;

[0037] Adopt a triplet loss function, and calculate the first loss according to the second positive sample and the second negative sample. Adopt an infoNCE loss function, and calculate the second loss and the third loss according to the first type of negative sample and the second type of hard negative sample respectively;

[0038] Combine the first loss, the second loss and the third loss to obtain the second contrastive learning loss.

[0039] In a possible implementation manner, the step of processing the frame-level features based on the initial Transformer model to obtain fine-grained video sample features and coarse-grained video sample features includes:

[0040] Use the initial Transformer model to perform semantic modeling on the video samples in the training samples to obtain fine-grained video sample features;

[0041] Downsample the fine-grained video sample features to obtain coarse-grained video sample features.

[0042] In a possible implementation manner, the step of obtaining retrieval results according to the similarity data includes:

[0043] For each video data in the video library, perform weighted summation on the coarse-grained similarity and the fine-grained similarity corresponding to the video data to obtain the video-level similarity between the video data and the text to be queried;

[0044] Sort all the video data according to the video-level similarity, and select a preset number of video data as the video-level retrieval results according to the sorting result;

[0045] According to the fine-grained similarity, select a preset number of fine-grained candidate segments with the highest cosine similarity from each candidate video as the candidate segments; wherein, the cosine similarity is the cosine similarity between the fine-grained candidate segment and the text to be queried;

[0046] Sort all the candidate segments according to the cosine similarity, and select a preset number of candidate segments as the segment-level retrieval results according to the sorting result.

[0047] In a possible implementation manner, the step of obtaining the feature information of each video data in the video library includes:

[0048] Use a pre-trained visual feature extraction model to extract frame-level features of each video data in the video library;

[0049] Perform semantic modeling on the frame-level features to obtain fine-grained video features, and downsample the fine-grained video representation to obtain coarse-grained video features.

[0050] In a possible implementation manner, the step of processing the text to be queried to obtain the sentence-level text feature corresponding to the text to be queried includes:

[0051] Adopt a pre-trained RoberTa model to extract symbol-level features of the text to be queried, and perform context modeling and feature aggregation on the symbol-level features to obtain sentence-level text features.

[0052] In a second aspect, an embodiment of the present application provides a multi-granularity video retrieval device, including a preprocessing module, a feature acquisition module, an input module, a retrieval processing module, and a result acquisition module;

[0053] The preprocessing module is configured to process the text to be queried to obtain the sentence-level text feature corresponding to the text to be queried;

[0054] The feature acquisition module is configured to acquire feature information of each video data in the video library; wherein, the feature information includes coarse-grained video features and fine-grained video features, and the coarse-grained video features are obtained by downsampling the fine-grained video features;

[0055] The input module is configured to input the sentence-level text feature into a pre-trained retrieval algorithm;

[0056] The retrieval processing module is configured to perform multi-center and multi-scale dual-branch collaborative feature processing based on the sentence-level text feature and the feature information through the retrieval algorithm to obtain similarity data between the text to be queried and each video data; wherein, the similarity data includes coarse-grained similarity and fine-grained similarity;

[0057] The result acquisition module is configured to obtain a retrieval result according to the similarity data; wherein, the retrieval result includes an overall-level video corresponding to video-level retrieval and a segment-level video corresponding to segment-level retrieval.

[0058] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the multi-granularity video retrieval method as described in any possible implementation manner in the first aspect.

[0059] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the air conditioner remote control method described in any possible implementation manner of the first aspect is implemented.

[0060] The multi-granularity video retrieval method and device provided by the embodiments of the present application process the text to be queried to obtain corresponding sentence-level text features, and obtain the coarse-grained video features and fine-grained video features of each video data in the video library. The sentence-level text features are input into a pre-trained retrieval algorithm. Thus, through this retrieval algorithm, based on the sentence-level text features and the coarse-grained video features and fine-grained video features of each video data in the video library, multi-center and multi-scale dual-branch collaborative feature processing is performed to obtain similarity data between the text to be queried and each video data, and a retrieval result including the overall-level video corresponding to video-level retrieval and the segment-level video corresponding to segment-level retrieval is obtained. Video-level retrieval and segment-level retrieval are performed through dual-branch hybrid collaboration in multiple granularities and multiple directions, so as to improve the retrieval recall rate and the positioning accuracy of the target segment in the video, and improve the retrieval accuracy.

[0061] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 The structural schematic diagram of the multi-granularity video retrieval system provided by the embodiments of the present application is shown.

[0064] Figure 2 One of the flow schematic diagrams of the multi-granularity video retrieval method provided by the embodiments of the present application is shown.

[0065] Figure 3 Shows Figure 2 The flow schematic diagram of some sub-steps of step S13 in

[0066] Figure 4 Shows Figure 2 The flow schematic diagram of some sub-steps of step S14 in

[0067] Figure 5 Another flow schematic diagram of the multi-granularity video retrieval method provided by the embodiments of the present application is shown.

[0068] Figure 6 Shows the processing logic diagram of the multi-granularity video retrieval method provided by the embodiments of the present application.

[0069] Figure 7 Shows Figure 5 The flowchart of partial sub-steps of step S27 in.

[0070] Figure 8 Shows Figure 5 The flowchart of partial sub-steps of step S29 in.

[0071] Figure 9 Shows the application result diagram of the multi-granularity video retrieval method provided by the embodiments of the present application.

[0072] Figure 10 Shows the structural schematic diagram of the multi-granularity video retrieval device provided by the embodiments of the present application.

[0073] Figure 11 Shows the structural schematic diagram of the electronic device provided by the embodiments of the present application.

[0074] Explanation of reference numerals: 1000 - multi-granularity video retrieval system; 10 - retrieval device; 20 - client; 30 - training device; 40 - multi-granularity video retrieval device; 401 - preprocessing module; 402 - feature acquisition module; 403 - input module; 404 - retrieval processing module; 405 - result acquisition module; 50 - electronic device. Detailed implementation manners

[0075] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and shown here can be arranged and designed in various different configurations.

[0076] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0077] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.

[0078] Most traditional text-video retrieval tasks are aimed at pre-clipped short videos, which are characterized by highly aligned semantic information between the video and the text. However, in actual scenarios, most videos are long videos that have not been finely clipped. When using query text to retrieve videos, the query text input by the user only has a semantic correlation with some fragments of the target video, that is, some fragments in the video are semantically aligned with the query text. As a result, the accuracy of text-video retrieval is relatively low.

[0079] To address the problem of long-video retrieval, some solutions have been proposed. One type of method defines the text retrieval task as a multi-instance learning problem, divides the video into multiple candidate fragments, maps the text and video fragment features to the same feature subspace, and performs semantic alignment through contrastive learning. However, this method only focuses on the retrieval of the entire video and does not further achieve the localization of specific fragments in the target video. It is a coarse-grained retrieval and thus has limitations in practical applications.

[0080] Another type of method focuses on both video retrieval and video fragment retrieval, but these methods have the following problems: (1) They rely on the annotation of video target fragments in the dataset, which are costly and highly subjective, making it difficult to apply on a large scale; (2) These methods often implement video retrieval and fragment retrieval in stages, without forming a unified ranking framework, and there are sample selection issues between different stages, reducing the upper limit of the retrieval recall effect.

[0081] Currently, there is a cross-modal video retrieval method that encodes the initial feature sequence of a video using a preview branch and a close-reading branch respectively to obtain preview features and precision features, and inputs the preview features and precision features and the text-modal multi-level encoded feature mapping into the corresponding hybrid space respectively. The similarity between the video modality and the text modality is calculated through the hybrid space for modality matching, that is, text-video retrieval is performed. However, this method uses a bidirectional GRU for preview features to generate video feature vectors, and only considers enabling the model to distinguish different videos during model training, without considering the ability to distinguish different videos and different segments of the same video, resulting in low retrieval accuracy.

[0082] In view of the above considerations, an embodiment of the present application provides a multi-granularity video retrieval method, which has the ability to distinguish different videos and different segments of the same video, and can improve the retrieval recall rate and the positioning accuracy of target segments in the video while realizing video-level retrieval and segment-level retrieval.

[0083] The multi-granularity video retrieval method provided by the embodiment of the present application can be applied to Figure 1 the multi-granularity video retrieval system 1000 shown in the figure. The multi-granularity video retrieval system 1000 may include a retrieval device 10, a client 20, and a training device 30. The retrieval device 10 may be communicatively connected to the client 20 through a network. The training device 30 and the retrieval device 10 may be the same device, or two devices that can be communicatively connected through a network or a wired manner.

[0084] The training device 30 is configured to train a retrieval algorithm and deploy the trained retrieval algorithm to the retrieval device 10.

[0085] The client 20 is configured to obtain a text to be queried input by a user and send the text to be queried to the retrieval device 10.

[0086] The retrieval device 10 is configured to implement the multi-granularity video retrieval method provided by the embodiment of the present application based on the deployed retrieval algorithm and the text to be queried.

[0087] The training device 30 and the retrieval device 10 include, but are not limited to: a server cluster, an independent server, a cloud server, a personal computer, etc.

[0088] The client 20 includes, but is not limited to: a personal computer, a laptop, a tablet computer, a mobile terminal, a mobile phone, a wearable portable device, a virtual reality device, a smart terminal, etc.

[0089] In a possible implementation manner, a multi-granularity video retrieval method is provided. Refer to Figure 2, the multi-granularity poem retrieval method can obtain a retrieval algorithm through the following steps. In this embodiment, the manner of obtaining the retrieval algorithm is applied to Figure 1 training device 30 in the following for illustration.

[0090] S10. Process each sample pair in the training dataset to obtain the sentence-level text features corresponding to the query text samples in each sample pair, and the frame-level features of the video samples in each sample pair.

[0091] It should be noted that each sample pair in the training dataset includes a query text sample and a video sample, and the query text sample in the sample pair is a natural language description of each segment of the video sample in the sample pair.

[0092] S11. Select a preset number of sample pairs from the training dataset as training samples, and input each training sample pair into the initial retrieval algorithm.

[0093] In this embodiment, the initial retrieval algorithm includes at least an initial Transformer model, an initial browsing branch, and an initial gazing branch. Each training sample pair includes a training video and a training query text. The preset number can be 100, 200, 50, etc., and can be set flexibly.

[0094] S12. Based on the initial Transformer model, process the frame-level features to obtain fine-grained video sample features and coarse-grained video sample features.

[0095] S13. Through the initial browsing branch, based on multiple center points and multiple scales, construct multiple coarse-grained sample candidate segments of the training video, and combine the coarse-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal coarse-grained sample segment from the multiple coarse-grained sample candidate segments, and calculate the coarse-grained similarity between the training video and the training query text.

[0096] S14. Through the initial gazing branch, according to the center point and multiple scales of the optimal coarse-grained sample segment, construct multiple fine-grained sample candidate segments of the training video, and combine the fine-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal fine-grained sample segment from the multiple fine-grained sample candidate segments, and calculate the fine-grained similarity between the training video and the training query text.

[0097] It should be noted that for each training sample, steps S14 and S13 are executed once to obtain the coarse-grained similarity and fine-grained similarity between the training video and the training query text of the training sample.

[0098] S15. Based on the coarse-grained similarity and the fine-grained similarity, combine the training videos in all training samples, and calculate the first contrastive learning loss for the coarse-grained and the second contrastive learning loss for the fine-grained.

[0099] S16. Combine the first contrastive learning loss and the second contrastive learning loss to obtain a hybrid collaborative contrastive learning loss. Based on the hybrid collaborative contrastive learning loss, use an optimization algorithm to update the parameters of the initial retrieval algorithm.

[0100] S17. Determine whether the iteration end condition is satisfied. If so, end the training and obtain a mature retrieval algorithm. If not, return to execute step S11.

[0101] The iteration end condition can be set flexibly. For example, it can be that the number of iterations reaches a preset number, or the hybrid collaborative contrastive learning loss is less than a preset loss threshold. In this embodiment, it is not specifically limited.

[0102] The method for obtaining the sentence-level text features corresponding to the query text samples in each sample pair obtained in step S10 can also be set flexibly. For example, it can be processed by a machine learning model or processed according to preset rules. In this embodiment, it is not specifically limited.

[0103] In a possible implementation manner, a pre-trained RoberTa model can be used to extract the token-level features of the query text samples, and perform context modeling and feature aggregation on the token-level features to obtain sentence-level text features.

[0104] For context modeling and feature aggregation of the token-level features, a single-layer transformer encoder can be used. Through the multi-head self-attention mechanism, combined with learnable position encoding, the context information of the token-level features is constructed to obtain multiple token-level feature vectors. Furthermore, an additive attention mechanism is used to aggregate the multiple token-level feature vectors into one vector to obtain sentence-level text features.

[0105] The feature aggregation process can be represented by an aggregation formula, and the aggregation formula includes:

[0106]

[0107] where α = Softmax(QW T ), q represents the aggregated sentence-level text features, n q represents the number of token-level feature vectors of the query text (sample), q i represents the i-th token-level feature vector, α represents the self-attention weight matrix, represents the matrix composed of all token-level feature vectors of the sentence, is a learnable vector, WT Denotes the transpose of W.

[0108] In addition, in step S10, a pre-trained visual feature extraction model can be used to extract the frame-level features of the video samples in the sample pair. Among them, the visual feature extraction model can be any model. For example, it can be a pre-trained 2D deep convolutional neural network model or a 3D deep convolutional neural network model.

[0109] In order to improve the retrieval accuracy of the retrieval algorithm, in a possible implementation manner, multiple different pre-trained visual feature extraction models can be used for extraction respectively, and the extracted visual features are concatenated as the final frame-level features. In this way, the frame-level features can include the feature information of all dimensions, thereby improving the retrieval accuracy.

[0110] For step S12, the initial Transformer model can be a single-layer Transformer model, and the Transformer model includes a multi-head self-attention mechanism. Using a single-layer Transformer model, that is, a multi-head self-attention mechanism, with the entire training video as the receptive field, semantic modeling is performed on the frame-level features (the frame-level features contain position encoding) of the training video to obtain fine-grained video sample features. Furthermore, temporal max pooling is used to downsample the fine-grained video sample features, and temporal one-dimensional convolution is used for local semantic modeling on the downsampling result to obtain coarse-grained video sample features.

[0111] The set of fine-grained video sample features can be expressed as: The set of coarse-grained video sample features can be expressed as: where n v represents the number of fine-grained video sample features, and respectively represent the i-th fine-grained video sample feature and the i-th coarse-grained video sample feature, and n c represents the number of coarse-grained video sample features.

[0112] In a possible implementation manner, referring to Figure 3 , step S13 can be further implemented as the following steps.

[0113] S131, Select multiple center points and multiple scales. Taking each center point as the center point of the segment and each scale as the width of the segment, the training video is divided into multiple coarse-grained sample candidate segments.

[0114] Multiple (which can be n pc pieces) center points can be selected at equal intervals between 0 and 1, and multiple (which can be npw ) Scale, and divide the training video into multiple coarse-grained sample candidate segments according to each center point as the center point of the segment and each scale as the width of the segment, where the total number of coarse-grained sample candidate segments is n p = n pc × n pw .

[0115] For example, assume there are 5 center points: A, B, C, D, E, and 5 scales: a, b, c, d, e. Then there are 25 coarse-grained sample candidate segments, namely: five coarse-grained sample candidate segments with A as the center point and widths a, b, c, d, e respectively; five coarse-grained sample candidate segments with B as the center point and widths a, b, c, d, e respectively; five coarse-grained sample candidate segments with C as the center point and widths a, b, c, d, e respectively; five coarse-grained sample candidate segments with D as the center point and widths a, b, c, d, e respectively; and five coarse-grained sample candidate segments with E as the center point and widths a, b, c, d, e respectively.

[0116] S132. For each coarse-grained sample candidate segment, combine the coarse-grained video sample features, the center point and width of this coarse-grained sample candidate segment, and perform Gaussian weighted pooling aggregation to obtain the segment feature of this coarse-grained sample candidate segment.

[0117] The segment feature of the coarse-grained sample candidate segment can be expressed as:

[0118]

[0119] where c j represents the segment feature of the j-th coarse-grained sample candidate segment, n c represents the number of coarse-grained video sample features, represents the i-th coarse-grained video sample feature, represents the center point of the j-th coarse-grained sample candidate segment, represents the width of the j-th coarse-grained sample candidate segment, and σ is the scaling factor.

[0120] S133. Based on the sentence-level text feature and the segment feature, calculate the cosine similarity between each coarse-grained sample candidate segment and the training query text, take the coarse-grained sample candidate segment with the maximum cosine similarity as the optimal coarse-grained sample segment, and take the cosine similarity of this optimal coarse-grained sample segment as the coarse-grained similarity between the training video and the training query text.

[0121] In this embodiment, the coarse-grained similarity can be expressed as: where S c(q, v) represents the coarse-grained similarity, c np represents the segment feature of the np-th coarse-grained sample candidate segment, q T represents the transpose of the sentence-level text feature vector of the training query text.

[0122] Through the above steps S131 to S133, the browsing branch constructs multiple candidate segments with multiple center points and multiple different scales for each center point, and when calculating the segment features of the candidate segments, considering the different semantic representativeness between different frames within the segment, the Gaussian weighted pooling method is used to aggregate the segment features, so as to be able to focus more on the video frames closer to the center of the segment, thereby helping to improve the video retrieval accuracy.

[0123] In a possible implementation manner, referring to Figure 4 , step S14 can be further implemented as the following steps.

[0124] S141, select multiple scales, use the center point of the optimal segment of the coarse-grained sample as the segment center point, and use each scale as the width of the segment to divide the training video into multiple fine-grained sample candidate segments.

[0125] Multiple scales (which can be lb to ub ) can be selected at equal intervals between the preset scale lower limit value and scale upper limit value, and according to using the center point of the optimal segment of the coarse-grained sample as the segment center point and using each scale as the width of the segment, the training video is divided into multiple fine-grained sample candidate segments, where the total number of coarse-grained sample candidate segments is

[0126] For example, assuming that the center point of the optimal segment of the coarse-grained sample is C and there are 5 scales: a, b, c, d, e, then the fine-grained sample candidate segments are 5: five fine-grained sample candidate segments with C as the center point and widths of a, b, c, d, e respectively.

[0127] S142, for each fine-grained sample candidate segment, combine the fine-grained video sample features and the center point and width of this fine-grained sample candidate segment, and perform Gaussian weighted pooling aggregation to obtain the segment feature of this fine-grained sample candidate segment.

[0128] The calculation method of the segment feature of the fine-grained sample candidate segment can refer to the calculation method of the segment feature of the coarse-grained sample candidate segment in S132 above, and will not be elaborated in this implementation manner.

[0129] ​S143. Calculate the cosine similarity between each candidate segment of the fine-grained sample and the training query text based on the sentence-level text features and segment features. Select the candidate segment of the fine-grained sample with the maximum cosine similarity as the optimal segment of the fine-grained sample, and use the cosine similarity of this optimal segment of the fine-grained sample as the fine-grained similarity between the training video and the training query text.

[0130] The calculation formula for the fine-grained similarity can refer to the calculation formula for the coarse-grained similarity in step S133, which will not be elaborated here.

[0131] For step S15, any loss function can be used to calculate the first contrastive learning loss and the second contrastive learning loss. In this embodiment, no specific limitation is made.

[0132] In a possible implementation, to improve the retrieval accuracy of the retrieval algorithm, a joint triplet loss function and an infoNCE loss function are introduced.

[0133] For the calculation of the first contrastive learning loss regarding the coarse-grained level, the first positive sample and the first negative sample can be determined from the training videos in all training samples according to the coarse-grained similarity, and the first contrastive learning loss regarding the coarse-grained level can be calculated by combining the joint triplet loss function and the infoNCE loss function.

[0134] For each training query text of each training sample, the training video corresponding to this training query text can be used as the first positive sample, and the remaining training videos in all training samples can be used as the first negative sample. For each training video of each training sample, the training query text corresponding to this training video can be used as the first positive sample, and the remaining training query texts in all training samples can be used as the first negative sample. Then, according to the first positive sample and the first negative sample, combined with the coarse-grained similarity, the joint triplet loss value and the infoNCE loss value are calculated respectively, and the weighted sum of the joint triplet loss value and the infoNCE loss value is used as the first contrastive learning loss.

[0135] At this time, the joint triplet loss value can be expressed as:

[0136]

[0137] where n represents the number of training samples selected in the current round of training, represents the training sample, v - represents a training video randomly selected from this batch of training samples that does not match the current text q (i.e., the first negative sample), q - represents the training query text that does not match the current video v in this batch of training samples (i.e., the first negative sample), Δ1 represents the margin hyperparameter, S c (q, v) represents the coarse-grained similarity, Sc (q - , v) represents the cosine similarity between the current video v and the training query text that does not match it, S c (q, v - ) represents the cosine similarity between the current text q and the training video that does not match it.

[0138] The infoNCE loss value can be expressed as:

[0139]

[0140] The first contrastive learning loss can be expressed as: where β1 represents the weight hyperparameter, and the value of the weight hyperparameter can be adjusted according to the actual situation. For example, it can be values such as 0.1, 0.01, 0.04, 0.4, etc.

[0141] Regarding the calculation of the second contrastive learning loss for fine-grained, based on the fine-grained similarity, the second positive sample, the second negative sample, the first type of negative sample, and the second type of hard negative sample can be determined from the training videos in all training samples, and by combining the triplet loss function and the infoNCE loss function, the second contrastive learning loss for fine-grained can be calculated.

[0142] For each training query text of each training sample, the training video corresponding to this training query text can be used as the second positive sample, and the remaining training videos in all training samples can be used as the second negative sample. For each training video of each training sample, the training query text corresponding to this training video can be used as the second positive sample, and the remaining training query texts in all training samples can be used as the second negative sample. Then, by using the combined triplet loss function and the infoNCE loss function, according to this second positive sample, the second negative sample, and the fine-grained similarity, the combined triplet loss value and the infoNCE loss value are calculated respectively, and the weighted sum of the combined triplet loss value and the infoNCE loss value is used as the contrastive learning loss between fine-grained videos.

[0143] The representation method of the contrastive learning loss between fine-grained videos is roughly the same as the representation method of calculating the first contrastive learning loss above, the difference being that the coarse-grained similarity is replaced by the fine-grained similarity.

[0144] In addition, for each optimal segment of the fine-grained sample, the video segments at the left and right ends of this optimal segment of the fine-grained sample are used as the first type of negative sample of this optimal segment of the fine-grained sample, and the entire training video to which this optimal segment of the fine-grained sample belongs is used as the second type of hard negative sample. The triplet loss is calculated by combining the triplet loss function respectively. And the sum of all triplet losses of this optimal segment of the fine-grained sample is obtained to get the contrastive learning loss within the video.

[0145] The contrastive learning loss within a video can be expressed as:

[0146] L intra = L trip [S c (q,c), S n1 (q,v), Δ2)] + L trip [S c (q,v), S n2 (q,v), Δ2)]

[0147] + L trip [S c (q,v), S n3 (q,v), Δ3)]

[0148] where S n1 (q,v) and S n2 (q,v) respectively represent the cosine similarity between the segment features of a type of negative sample and the sentence-level text features, and S n3 (q,v) represents the similarity between the second type of hard negative samples and the sentence-level text features. Both Δ2 and Δ3 represent boundary hyperparameters.

[0149] The fine-grained inter-video contrastive learning loss and the contrastive learning loss within a video are taken as the second contrastive learning loss regarding fine-grainedness.

[0150] In step S16, the first contrastive learning loss and the second contrastive learning loss, that is, the first contrastive learning loss, the fine-grained inter-video contrastive learning loss, and the contrastive learning loss within a video, are weighted and summed to obtain the hybrid collaborative contrastive learning loss.

[0151] In the above manner, during the training process of the retrieval algorithm, a collaborative manner guided by the focus (i.e., the center point and scale) is adopted between the browsing branch and the gazing branch, which can enhance the semantic alignment ability of the retrieval algorithm model for the query text and the specific segments of the video. Additionally, a video feature vector generation mechanism based on multi-scale candidate segments is adopted in both the browsing branch and the gazing branch, considering contrastive learning in multiple granularities and multiple directions (i.e., multiple center points and multiple scales), and a hybrid collaborative contrastive learning loss is designed by combining the first contrastive learning loss and the second contrastive learning loss, enabling the retrieval algorithm (model) to simultaneously consider the discrimination ability for different videos and different segments of the same video. Thus, while the retrieval algorithm (model) realizes video-level retrieval and segment-level retrieval, it can improve the retrieval recall rate and the positioning accuracy of the target segments in the video.

[0152] In a possible implementation manner, referring to Figure 5 , the multi-granularity video retrieval method provided by the embodiment of the present application may further include the following steps. The retrieval algorithm trained in the above embodiment is deployed on Figure 1After the retrieval device, the multi-granularity video retrieval, i.e., video-level retrieval and segment-level retrieval, can be achieved through the following steps. The following steps are the video retrieval process for deploying the retrieval algorithm.

[0153] S21. Process the text to be queried to obtain the sentence-level text features corresponding to the text to be queried.

[0154] For the implementation manner of step S21, reference can be made to the manner of obtaining sentence-level text features in the above S10, and it will not be elaborated in this implementation manner.

[0155] S23. Obtain the feature information of each video data in the video library.

[0156] In this implementation manner, the feature information includes coarse-grained video features and fine-grained video features, and the coarse-grained video features are obtained by downsampling the fine-grained video features. The manner of obtaining the feature information of each video data in the video library is the same as the manner of obtaining the fine-grained video sample features and coarse-grained video sample features in the above step S12, and it will not be elaborated in this implementation manner.

[0157] It should be noted that for the feature information of each video data in the video library, it can be processed and obtained before performing video retrieval and stored in the database. In step S22, only loading and calling are required.

[0158] S25. Input the sentence-level text features into the pre-trained retrieval algorithm.

[0159] The pre-trained retrieval algorithm is the retrieval algorithm trained in the manner of the above steps S10 to S17.

[0160] S27. Through the retrieval algorithm, based on the sentence-level text features and the feature information, perform multi-center and multi-scale dual-branch collaborative feature processing to obtain the similarity data between the text to be queried and each video data.

[0161] Among them, the similarity data includes coarse-grained similarity and fine-grained similarity.

[0162] S29. Obtain the retrieval result according to the similarity data.

[0163] In this implementation manner, the retrieval result includes the overall-level video corresponding to the video-level retrieval and the segment-level video corresponding to the segment-level retrieval.

[0164] In the above multi-granularity video retrieval method, the query text is processed to obtain the corresponding sentence-level text features, and the coarse-grained video features and fine-grained video features of each video data in the video library are obtained. The sentence-level text features are input into a pre-trained retrieval algorithm. Thus, through this retrieval algorithm, based on the sentence-level text features and the coarse-grained video features and fine-grained video features of each video data in the video library, multi-center and multi-scale dual-branch collaborative feature processing is performed to obtain the similarity data between the query text and each video data, and based on this similarity data, the retrieval results including the overall-level videos corresponding to video-level retrieval and the segment-level videos corresponding to segment-level retrieval are obtained. Video-level retrieval and segment-level retrieval are performed through multi-granularity and multi-directional dual-branch hybrid collaboration, thereby being able to improve the retrieval recall rate and the positioning accuracy of the target segments in the video, and enhancing the retrieval accuracy.

[0165] The retrieval algorithm includes a browsing branch and a gazing branch. Refer to Figure 6 , Figure 6 which is the video retrieval logic diagram of the multi-granularity video retrieval method provided by this application in the deployment environment. In a possible implementation manner, refer to Figure 7 , step S27 can be further implemented as the following steps.

[0166] S271, through the browsing branch, based on the selected multiple center points and multiple scales, construct multiple coarse-grained candidate segments for each video data, and in combination with the coarse-grained video features and the sentence-level text features, obtain the coarse-grained optimal segment from the multiple coarse-grained candidate segments, and calculate the coarse-grained similarity between the query text and each video data.

[0167] S272, through the gazing branch, according to the center point and multiple scales of the coarse-grained optimal segment, construct multiple fine-grained candidate segments for each video data, and in combination with the fine-grained video features and the sentence-level text features, obtain the fine-grained optimal segment from the multiple fine-grained candidate segments, and calculate the fine-grained similarity between the query text and each video data.

[0168] It should be understood that for each video data, through the above S271 and step S272, the fine-grained similarity between the query text and each video data is calculated, and the corresponding fine-grained optimal segment is obtained.

[0169] For step S271, multiple center points and multiple scales can be selected. Taking each center point as the center point of a segment and each scale as the width of the segment, the video data is divided into multiple coarse-grained candidate segments. For each coarse-grained candidate segment, Gaussian weighted pooling aggregation is performed by combining the coarse-grained video features of the video data and the center point and width of this coarse-grained candidate segment, to obtain the segment feature of this coarse-grained candidate segment. Furthermore, based on the sentence-level text features of the text to be queried and the segment features, the cosine similarity between each coarse-grained candidate segment and the text to be queried is calculated, and the coarse-grained candidate segment with the maximum cosine similarity is taken as the optimal coarse-grained segment, and the cosine similarity of this optimal coarse-grained segment is taken as the coarse-grained similarity between this video data and the text to be queried.

[0170] Multiple (which can be n pc pieces) center points can be selected at equal intervals between 0 and 1, and multiple (which can be n pe ) scales are selected at equal intervals between a preset scale lower limit value and a scale upper limit value, and the training video is divided into multiple coarse-grained sample candidate segments according to taking each center point as the center point of a segment and each scale as the width of the segment, where the total number of coarse-grained sample candidate segments is n p = n pc ×n pw .

[0171] For example, assume there are 5 center points: A, B, C, D, E, and 5 scales: a, b, c, d, e. Then there are 25 coarse-grained sample candidate segments, which are: five coarse-grained sample candidate segments with A as the center point and widths of a, b, c, d, e respectively; five coarse-grained sample candidate segments with B as the center point and widths of a, b, c, d, e respectively; five coarse-grained sample candidate segments with C as the center point and widths of a, b, c, d, e respectively; five coarse-grained sample candidate segments with D as the center point and widths of a, b, c, d, e respectively; and five coarse-grained sample candidate segments with E as the center point and widths of a, b, c, d, e respectively.

[0172] The segment feature of the coarse-grained candidate segment can be expressed as:

[0173]

[0174] where, c j represents the segment feature of the j-th coarse-grained candidate segment, b c represents the number of coarse-grained video sample features, represents the i-th coarse-grained video feature, represents the center point of the j-th coarse-grained candidate segment, Characterize the width of the j-th coarse-grained candidate segment, and σ is the scaling factor.

[0175] In this embodiment, the coarse-grained similarity can be expressed as: Where S c (q, v) characterizes the coarse-grained similarity, and c np characterizes the segment feature of the np-th coarse-grained sample candidate segment, and q T characterizes the transpose of the sentence-level text feature vector.

[0176] In the above manner, the browsing branch adopts multiple center points, and each center point is set with multiple different scales to construct multiple candidate segments with multiple scales and multiple center points. When calculating the segment features of the candidate segments, considering the different semantic representativeness between different frames within the segment, the Gaussian weighted pooling method is used to aggregate the segment features, so as to be able to focus more on the video frames closer to the center of the segment, thereby greatly improving the video retrieval accuracy.

[0177] For step S272, multiple scales can be selected. Taking the center point of the coarse-grained optimal segment as the segment center point and each scale as the width of the segment, the video data is divided into multiple fine-grained candidate segments. And for each fine-grained candidate segment, combining the fine-grained video features and the center point and width of the fine-grained candidate segment, Gaussian weighted pooling aggregation is performed to obtain the segment feature of the fine-grained candidate segment. Furthermore, based on the sentence-level text feature and segment feature of the text to be queried, the cosine similarity between each fine-grained candidate segment and the text to be queried is calculated, and the fine-grained candidate segment with the largest cosine similarity is used as the fine-grained optimal segment, and the cosine similarity of the fine-grained optimal segment is used as the fine-grained similarity between the video data and the training query text.

[0178] It can be between a pre-set scale lower limit value and a scale upper limit value (i.e., w lb to w ub ), multiple (which can be ) scales are selected at equal intervals, and the video data is divided into multiple fine-grained candidate segments with the center point of the coarse-grained optimal segment as the segment center point and each scale as the width of the segment. The total number of coarse-grained candidate segments can be

[0179] For example, assume that the center point of the coarse-grained optimal segment is C, and there are 5 scales: a, b, c, d, e. Then the fine-grained candidate segments are 5: five fine-grained sample candidate segments with C as the center point and widths of a, b, c, d, e respectively.

[0180] The calculation method of the segment features of the fine-grained candidate segments can refer to the calculation method of the segment features of the coarse-grained candidate segments in S132 above. In this embodiment, it will not be elaborated. Similarly, the calculation formula of the fine-grained similarity can refer to the calculation formula of the coarse-grained similarity in step S133, and will not be elaborated here.

[0181] Through the above steps S271 and S272, based on the semantic alignment ability between the browsing branch and the gazing branch of the retrieval algorithm for the query text and the specific segments of the video, the video feature vector generation mechanism of the multi-scale candidate segments, and the discrimination ability for different videos and different segments of the same video, through a hybrid collaboration for multi-center and multi-scale dual-branch collaborative feature processing, the coarse-grained similarity and the fine-grained similarity between the query text and each video data are obtained.

[0182] For step S29, refer to Figure 8 , it can be further implemented as the following steps.

[0183] S291, for each video data in the video library, perform a weighted sum of the coarse-grained similarity of the fine-grained candidate segments corresponding to the video data and the coarse-grained similarity of the coarse-grained candidate segments to obtain the video-level similarity between the video data and the query text.

[0184] S292, sort all the video data according to the video-level similarity, and select a preset number of video data as the video-level retrieval results according to the sorting results.

[0185] S293, from each candidate video, select a preset number of fine-grained candidate segments with the highest cosine similarity as the candidate segments according to the fine-grained similarity.

[0186] It should be noted that the preset number k can be 10, 5, 12, etc., and is not specifically limited in this embodiment. The cosine similarity in step S293 is the cosine similarity between the fine-grained candidate segment and the query text, and its calculation method can refer to the calculation method in step S27 above. In this embodiment, it will not be elaborated.

[0187] S294, sort all the candidate segments according to the cosine similarity, and select a preset number of candidate segments as the segment-level retrieval results according to the sorting results.

[0188] Through the above steps S291 to S292, the video-level retrieval results are obtained, and through the above steps S293 to S294, the segment-level retrieval results are obtained. Moreover, for different retrieval requirements (video-level retrieval and segment-level retrieval), multiple retrieval sorting strategies are proposed, which have flexibility in use and can be applied to various usage scenarios.

[0189] To test the retrieval results of the retrieval algorithm (model) provided in this application and verify the performance of the multi-granularity video detection method provided in this application, test cases are also provided in this application.

[0190] During testing, two evaluation metrics were used for video-level retrieval and video segment-level retrieval respectively. For video-level retrieval, the R@K metric can be used, including R@1, R@5, R@10, R@100, etc., which represents the proportion of the first K items in the retrieval ranking results of candidate videos among all query texts that contain the true target item. For segment-level retrieval, the R@K metric with an IoU (Intersection over Union) threshold can be used, where the IoU thresholds include 0.3, 0.5, and 0.7, which represents the proportion that at least one item in the first K items of the ranking results of video candidate segments among all query texts has an IoU greater than the threshold with the true target item.

[0191] Test scenario 1

[0192] Model training and testing were carried out on the Charades-STA dataset, and the effects were compared with previous methods. The Charades-STA dataset contains a total of 6,670 videos, which cover various indoor activities. The average length of the videos is 30.0 seconds. Each video corresponds to an average of 2.4 natural language text descriptions, and each text description corresponds to a specific segment in a certain video. The average length of these segments is 8.1 seconds. The effect comparison of video-level retrieval is shown in Table 1, and the effect comparison of video segment-level retrieval is shown in Table 2.

[0193] Table 1

[0194] Method R@1 R@5 R@10 R@100 Indicators and XML Model 1.6 6.0 10.1 46.9 64.6 DE+++ Model 1.7 5.6 9.6 37.1 54.1 ReLoCLNet Model 1.2 5.4 10.0 45.6 62.3 RIVRL Model 1.6 5.6 9.4 37.7 54.3 MS-SL Model 1.8 7.1 11.8 47.7 68.4 Retrieval Algorithm (Model) 2.4 7.7 12.8 49.8 72.7

[0195] Table 2

[0196]

[0197]

[0198] As can be seen from Table 1 and Table 2 above, the retrieval algorithm provided in this application is superior to existing models (including XML model, DE+++ model, ReLoCLNet model, RIVRL model, and MS-SL model, etc.) in terms of performance in video-level retrieval and segment-level retrieval on the Charades-STA dataset.

[0199] Test scenario 2

[0200] The model is trained and tested on the Activitynet-Captions dataset, and the effects are compared with previous methods. The Activitynet-Captions dataset contains more than 20,000 videos, which involve 200 different types of indoor and outdoor activities. The average length of the videos is 117.6 seconds. There are more than 100,000 natural language text descriptions corresponding to the videos. These text descriptions all correspond to a specific segment in a certain video, and the average length of these segments is 36.2 seconds. The effect comparison of video-level retrieval is shown in Table 3, and the effect comparison of video segment-level retrieval is shown in Table 4.

[0201] Table 1

[0202] Method R@1 R@5 R@10 R@100 Indicators and XML Model 5.3 19.4 30.6 73.1 128.4 DE+++ Model 5.3 18.4 29.2 68.0 121.0 ReLoCLNet Model 5.7 18.9 30.0 72.0 126.6 RIVRL Model 5.2 18.0 28.2 66.4 117.8 MS-SL Model 7.1 22.5 34.7 75.8 140.1 Retrieval Algorithm (Model) 6.8 22.7 34.8 76.1 140.5

[0203] Table 2

[0204]

[0205]

[0206] As can be seen from Table 3 and Table 4 above, on the Activitynet-Captions dataset, the performance of both video-level retrieval and segment-level retrieval provided by the retrieval algorithm of the present application is also superior to existing models (including XML model, DE+++ model, ReLoCLNet model, RIVRL model, MS-SL model, etc.).

[0207] Test scenario three

[0208] On the Charades-STA dataset, the actual results of multi-granularity video content retrieval for a given query text using the multi-granularity video retrieval method proposed in the present application are as follows. Figure 9 As shown, for the two given query texts, the multi-granularity video retrieval method proposed in the present application has successfully ranked the correct target video first, and for each candidate video, segments with relatively high relevance to the query text are given.

[0209] Preferably, since the multi-granularity video retrieval method proposed in the present application adopts a retrieval method based on similarity ranking, the candidate videos ranked in the front positions all show semantic relevance to the query text.

[0210] Based on the same concept as the above multi-granularity video retrieval method, in a possible implementation manner, a multi-granularity video retrieval device 40 is provided. Referring to Figure 10 , it may include a preprocessing module 401, a feature acquisition module 402, an input module 403, a retrieval processing module 404, and a result acquisition module 405.

[0211] A preprocessing module 401 for processing the text to be queried to obtain sentence-level text features corresponding to the text to be queried.

[0212] A feature acquisition module 402 for acquiring feature information of each video data in the video library. Among them, the feature information includes coarse-grained video features and fine-grained video features, and the coarse-grained video features are obtained by downsampling the fine-grained video features.

[0213] An input module 403 for inputting the sentence-level text features into a pre-trained retrieval algorithm.

[0214] A retrieval processing module 404 for performing multi-center and multi-scale dual-branch collaborative feature processing based on the sentence-level text features and the feature information through the retrieval algorithm to obtain similarity data between the text to be queried and each video data. Among them, the similarity data includes coarse-grained similarity and fine-grained similarity.

[0215] A result acquisition module 405 for obtaining a retrieval result according to the similarity data. Among them, the retrieval result includes an overall-level video corresponding to video-level retrieval and a segment-level video corresponding to segment-level retrieval.

[0216] In a possible implementation manner, it further includes a training module, and the training module is used to execute the training steps of the above steps S10 to S17 to obtain a mature retrieval algorithm.

[0217] In the above multi-granularity video retrieval device 40, through the collaborative action of the preprocessing module 401, the feature acquisition module 402, the input module 403, the retrieval processing module 404, and the result acquisition module 405, the text to be queried is processed to obtain corresponding sentence-level text features, and the coarse-grained video features and fine-grained video features of each video data in the video library are acquired, and the sentence-level text features are input into a pre-trained retrieval algorithm. Thus, through this retrieval algorithm, based on the sentence-level text features and the coarse-grained video features and fine-grained video features of each video data in the video library, multi-center and multi-scale dual-branch collaborative feature processing is performed to obtain similarity data between the text to be queried and each video data, and a retrieval result including an overall-level video corresponding to video-level retrieval and a segment-level video corresponding to segment-level retrieval is obtained according to the similarity data. Video-level retrieval and segment-level retrieval are performed through multi-granularity and multi-directional dual-branch hybrid collaboration, so as to improve the retrieval recall rate and the positioning accuracy of the target segment in the video, and enhance the retrieval accuracy.

[0218] For the specific limitations of the multi-granularity video retrieval device 40, reference may be made to the limitations on the multi-granularity video retrieval method in the foregoing text, which will not be elaborated herein. Each module in the above multi-granularity video retrieval device 40 can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the electronic device 50 in hardware form or be independent of it, or can be stored in the memory of the electronic device 50 in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0219] In one embodiment, an electronic device 50 is provided, and its internal structural diagram can be as Figure 11 shown. The electronic device 50 includes a processor, a memory, a communication interface, and an input device connected through a system bus. Among them, the processor of the electronic device 50 is used to provide computing and control capabilities. The memory of the electronic device 50 includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device 50 is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements the multi-granularity video retrieval method provided in the above embodiment.

[0220] Figure 11 The structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the electronic device 50 to which the solution of the present invention is applied. The specific electronic device 50 may include more or fewer components than Figure 11 shown in, or combine some components, or have different component arrangements.

[0221] In one embodiment, the multi-granularity video retrieval device 40 provided by the present invention applied to a deployed device can be implemented in the form of a computer program, and the computer program can run on an electronic device 50 as Figure 11 shown. The memory of the electronic device 50 can store each program module that makes up the multi-granularity video retrieval device 40. For example, Figure 10 the preprocessing module 401, the feature acquisition module 402, the input module 403, the retrieval processing module 404, and the result acquisition module 405 shown in. The computer program composed of each program module enables the processor to execute the steps in the multi-granularity video retrieval method described in this specification.

[0222] For example, Figure 11 the electronic device 50 shown in can be through as Figure 10The preprocessing module 401 in the multi-granularity video retrieval device 40 shown executes step S21. The electronic device 50 may execute step S23 through the feature acquisition module 402. The electronic device 50 may execute step S25 through the input module 403. The electronic device 50 may execute step S27 through the retrieval processing module 404. The electronic device 50 may execute step S29 through the result acquisition module 405.

[0223] In one embodiment, an electronic device 50 is provided, including: a processor and a memory for storing one or more programs; when the one or more programs are executed by the processor, the following steps are implemented: processing the text to be queried to obtain sentence-level text features corresponding to the text to be queried; obtaining feature information of each video data in the video library; inputting the sentence-level text features into a pre-trained retrieval algorithm; through the retrieval algorithm, based on the sentence-level text features and the feature information, performing multi-center and multi-scale dual-branch collaborative feature processing to obtain similarity data between the text to be queried and each video data; and obtaining a retrieval result according to the similarity data.

[0224] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: processing the text to be queried to obtain sentence-level text features corresponding to the text to be queried; obtaining feature information of each video data in the video library; inputting the sentence-level text features into a pre-trained retrieval algorithm; through the retrieval algorithm, based on the sentence-level text features and the feature information, performing multi-center and multi-scale dual-branch collaborative feature processing to obtain similarity data between the text to be queried and each video data; and obtaining a retrieval result according to the similarity data.

[0225] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0226] In addition, each functional module in various embodiments of this application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0227] If the described functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0228] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A multi-granularity video retrieval method, characterized in that, The method includes: processing the text to be queried to obtain the sentence-level text feature corresponding to the text to be queried; acquiring the feature information of each video data in the video library; wherein, the feature information includes a coarse-grained video feature and a fine-grained video feature, and the coarse-grained video feature is obtained by downsampling the fine-grained video feature; inputting the sentence-level text feature into a pre-trained retrieval algorithm; through the retrieval algorithm, based on the sentence-level text feature and the feature information, performing multi-center and multi-scale dual-branch collaborative feature processing to obtain the similarity data between the text to be queried and each video data; wherein, the similarity data includes a coarse-grained similarity and a fine-grained similarity; obtaining a retrieval result according to the similarity data; wherein, the retrieval result includes an overall-level video corresponding to video-level retrieval and a segment-level video corresponding to segment-level retrieval; the retrieval algorithm includes a browsing branch and a gazing branch; the step of performing multi-center and multi-scale dual-branch collaborative feature processing through the retrieval algorithm, based on the sentence-level text feature and the feature information, to obtain the similarity data between the text to be queried and each video data includes: through the browsing branch, based on a plurality of selected center points and a plurality of scales, constructing a plurality of coarse-grained candidate segments of each video data, and combining the coarse-grained video feature and the sentence-level text feature, obtaining a coarse-grained optimal segment from the plurality of coarse-grained candidate segments, and calculating the coarse-grained similarity between the text to be queried and each video data; through the gazing branch, according to the center point and a plurality of scales of the coarse-grained optimal segment, constructing a plurality of fine-grained candidate segments of each video data, and combining the fine-grained video feature and the sentence-level text feature, obtaining a fine-grained optimal segment from the plurality of fine-grained candidate segments, and calculating the fine-grained similarity between the text to be queried and each video data.

2. The multi-granularity video retrieval method according to claim 1, wherein The method further includes the step of training to obtain a retrieval algorithm, including: processing each sample pair in the training dataset to obtain the sentence-level text feature corresponding to the query text sample in each sample pair, and the frame-level feature of the video sample in each sample pair; wherein, the query text sample in the sample pair is a natural language description of the segment of the video sample in the sample pair; selecting a preset number of sample pairs from the training dataset as training samples, and inputting each training sample pair into an initial retrieval algorithm; wherein, the initial retrieval algorithm includes an initial Transformer model, an initial browsing branch and an initial gazing branch, and the training sample pair includes a training video and a training query text; processing the frame-level feature based on the initial Transformer model to obtain a fine-grained video sample feature and a coarse-grained video sample feature; Through the initial browsing branch, based on multiple center points and multiple scales, construct multiple coarse-grained sample candidate segments of the training video, and combine the coarse-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal coarse-grained sample segment from the multiple coarse-grained sample candidate segments, and calculate the coarse-grained similarity between the training video and the training query text; Through the initial gaze branch, according to the center point and multiple scales of the optimal coarse-grained sample segment, construct multiple fine-grained sample candidate segments of the training video, and combine the fine-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal fine-grained sample segment from the multiple fine-grained sample candidate segments, and calculate the fine-grained similarity between the training video and the training query text; Based on the coarse-grained similarity and the fine-grained similarity, combine the training videos in all the training samples to calculate the first contrastive learning loss for the coarse-grained and the second contrastive learning loss for the fine-grained; Combine the first contrastive learning loss and the second contrastive learning loss to obtain the hybrid collaborative contrastive learning loss. Based on the hybrid collaborative contrastive learning loss, use an optimization algorithm to update the parameters of the initial retrieval algorithm to obtain a mature retrieval algorithm.

3. The multi-granularity video retrieval method according to claim 2, characterized in that The step of combining the coarse-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal coarse-grained sample segment from the multiple coarse-grained sample candidate segments and calculating the coarse-grained similarity between the training video and the training query text includes: For each of the coarse-grained sample candidate segments, perform Gaussian weighted pooling aggregation by combining the coarse-grained video sample features and the center point and width of the coarse-grained sample candidate segment to obtain the segment feature of the coarse-grained sample candidate segment; Based on the sentence-level text feature and the segment feature, calculate the cosine similarity between each of the coarse-grained sample candidate segments and the training query text, take the coarse-grained sample candidate segment with the maximum cosine similarity as the optimal coarse-grained sample segment, and take the cosine similarity of the optimal coarse-grained sample segment as the coarse-grained similarity between the training video and the training query text.

4. The multi-granularity video retrieval method according to claim 2, wherein The step of combining the fine-grained video sample features and sentence-level text features of the training sample pair to obtain the optimal fine-grained sample segment from the multiple fine-grained sample candidate segments and calculating the fine-grained similarity between the training video and the training query text includes: For each of the fine-grained sample candidate segments, perform Gaussian weighted pooling aggregation by combining the fine-grained video sample features and the center point and width of the fine-grained sample candidate segment to obtain the segment feature of the fine-grained sample candidate segment; Based on the sentence-level text features and the segment features, calculate the cosine similarity between each fine-grained sample candidate segment and the training query text, and take the fine-grained sample candidate segment with the maximum cosine similarity as the optimal fine-grained sample segment, and take the cosine similarity of the optimal fine-grained sample segment as the fine-grained similarity between the training video and the training query text.

5. The multi-granularity video retrieval method according to claim 2, wherein, The step of calculating the first contrastive learning loss for the coarse-grained and the second contrastive learning loss for the fine-grained, in combination with the training videos in all the training samples, based on the coarse-grained similarity and the fine-grained similarity, includes: According to the coarse-grained similarity, determine the first positive sample and the first negative sample from the training videos in all the training samples, and jointly use the triplet loss function and the infoNCE loss function to calculate the first contrastive learning loss for the coarse-grained. Based on the fine-grained similarity, determine the second positive sample, the second negative sample, the first type of negative sample, and the second type of hard negative sample from the training videos in all the training samples, and jointly use the triplet loss function and the infoNCE loss function to calculate the second contrastive learning loss for the fine-grained.

6. The multi-granularity video retrieval method according to claim 2, wherein The step of processing the frame-level features based on the initial Transformer model to obtain the fine-grained video sample features and the coarse-grained video sample features includes: Use the initial Transformer model to perform semantic modeling on the video samples in the training samples to obtain the fine-grained video sample features. Downsample the fine-grained video sample features to obtain the coarse-grained video sample features.

7. The multi-granularity video retrieval method according to claim 1, wherein The step of obtaining the retrieval result according to the similarity data includes: For each video data in the video library, perform weighted summation on the corresponding coarse-grained similarity and fine-grained similarity of the video data to obtain the video-level similarity between the video data and the text to be queried. Sort all the video data according to the video-level similarity, and select a preset number of video data as the video-level retrieval result according to the sorting result. From each candidate video, select a preset number of fine-grained candidate segments with the highest cosine similarity as the candidate segments to be selected; where the cosine similarity is the cosine similarity between the fine-grained candidate segment and the text to be queried. Sort all the candidate segments to be selected according to the cosine similarity, and select a preset number of candidate segments to be selected as the segment-level retrieval result according to the sorting result.

8. The multi-granularity video retrieval method according to claim 1, wherein The step of processing the text to be queried to obtain the sentence-level text features corresponding to the text to be queried includes: Use the pre-trained RoberTa model to extract the token-level features of the text to be queried, and perform context modeling and feature aggregation on the token-level features to obtain the sentence-level text features.

9. A multi-granularity video retrieval device, characterized in that, A multi-granularity video retrieval method for implementing any one of claims 1-8 includes a preprocessing module, a feature acquisition module, an input module, a retrieval processing module, and a result acquisition module. The preprocessing module is used to process the text to be queried to obtain the sentence-level text features corresponding to the text to be queried; The feature acquisition module is used to acquire the feature information of each video data in the video library; wherein, the feature information includes coarse-grained video features and fine-grained video features, and the coarse-grained video features are obtained by downsampling the fine-grained video features; The input module is used to input the sentence-level text features into a pre-trained retrieval algorithm; The retrieval processing module is used to perform multi-center and multi-scale dual-branch collaborative feature processing based on the sentence-level text features and the feature information through the retrieval algorithm to obtain the similarity data between the text to be queried and each video data; wherein, the similarity data includes coarse-grained similarity and fine-grained similarity; The result acquisition module is used to obtain a retrieval result according to the similarity data; wherein, the retrieval result includes the overall-level video corresponding to video-level retrieval and the segment-level video corresponding to segment-level retrieval.

Citation Information

Patent Citations

  • Long video retrieval method and device based on multi-scale multi-example similarity learning

    CN115408558A

  • End-to-end multi-granularity contrast learning method for video text retrieval

    CN115757713A