Video retrieval positioning method, training method, electronic equipment, storage medium and program product

By introducing video characterization module, coarse-grained and fine-grained positioning module and multi-loss function optimization, the problem of poor multi-modal alignment effect of video retrieval model under weak supervision training is solved, and more efficient video clip positioning and retrieval is achieved.

CN120234447APending Publication Date: 2025-07-01INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510300139.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing video retrieval and clip positioning models are difficult to overcome the modal gap under weak supervision training, resulting in poor multimodal alignment and the search positioning performance needs to be improved.

Method used

The video retrieval positioning model is adopted, including the video characterization module, the coarse-grained positioning module, the fine-grained positioning module and the query text characterization module. By calculating significant alignment loss, fragment boundary alignment loss and inter-sample comparison learning loss, the model parameters are optimized to improve the multimodal alignment degree and search performance.

Benefits of technology

It significantly improves the multimodal alignment and retrieval performance of the video retrieval model, and can significantly lead the existing methods on ActivityNet and Charades datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234447A_ABST
    Figure CN120234447A_ABST
Patent Text Reader

Abstract

The invention provides a video retrieval positioning method, a training method, electronic equipment, a storage medium and a program product. The training method comprises the following steps: inputting a coarse-grained visual feature of each video sample and a text feature of a matched query text into a coarse-grained positioning module to obtain a fragment positioning focusing position of the video sample for the matched query text; inputting each video sample into a fine-grained positioning module according to the fragment positioning focusing position of the matched query text, the fine-grained visual feature and the text feature of the matched query text to obtain the boundary and fragment matching degree of the video sample according to the positioning fragment of the matched query text; and determining the total loss of the video retrieval positioning model based on the fine-grained visual features of each video sample in the current batch of training samples, the text features of each query text, and the boundary and fragment matching degree of each video sample for the positioning fragment of the matched query text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of data processing technologies, and more specifically, to a video retrieval and positioning method, a training method, an electronic device, a storage medium, and a program product. Background Art

[0002] Video retrieval and segment positioning is a technology for retrieving relevant videos from a video set according to a query text and further positioning relevant segments. For the weakly supervised training method of the model used for video retrieval and segment positioning, the model training can be completed only by using video-level annotation data. However, due to reasons such as the difficulty for the model to overcome the modality gap and achieve better multimodal alignment effects with separate video and text encoding modules, the retrieval and positioning performance of the model still needs to be further improved. Summary of the Invention

[0003] Exemplary embodiments of the present disclosure are directed to providing a video retrieval and positioning method, a training method, an electronic device, a storage medium, and a program product, which can solve at least one of the above problems existing in the prior art.

[0004] According to the first aspect of the embodiments of the present disclosure, a training method for a video retrieval and localization model is provided. The video retrieval and localization model includes a video representation module, a coarse-grained localization module, a fine-grained localization module, and a query text representation module. Wherein, the training method includes: obtaining the current batch of training samples, where the current batch of training samples includes: a plurality of training samples, and each training sample includes a video sample and a query text that matches the video sample; inputting each video sample in the current batch of training samples into the video representation module to obtain the coarse-grained visual features and fine-grained visual features of each video sample. Wherein, the fine-grained visual features of the video sample include: the visual features of multiple sampled frames obtained by performing a first sampling process on the video sample, and the coarse-grained visual features of the video sample are obtained by performing a second sampling process on the multiple sampled frames; inputting each query text in the current batch of training samples into the query text representation module to obtain the text features of each query text; inputting the coarse-grained visual features of each video sample and the text features of the matched query text into the coarse-grained localization module to obtain the segment localization focus position of the video sample for the matched query text; inputting the segment localization focus position of each video sample for the matched query text, the fine-grained visual features, and the text features of the matched query text into the fine-grained localization module to obtain the boundary and segment matching degree of the localization segment of the video sample for the matched query text. Wherein, the localization segment of the video sample for the matched query text is: the video segment with the highest matching degree with the query text in the video sample; determining the total loss of the video retrieval and localization model based on the fine-grained visual features of each video sample, the text features of each query text, the boundary and segment matching degree of the localization segment of each video sample for the matched query text in the current batch of training samples; updating the model parameters of the video retrieval and localization model based on the total loss of the video retrieval and localization model.

[0005] Optionally, the step of determining the total loss of the video retrieval and localization model includes: calculating the saliency alignment loss of the video retrieval and localization model based on the fine-grained visual features of each video sample, the text features of each query text, and the boundaries of the localization segments of each video sample for the matched query text, where the saliency alignment loss is used to enable the video retrieval and localization model to learn that: the matching degree between the sampled frames within the localization segment of the video sample for the matched query text should be significantly higher than the matching degree between other sampled frames of the video sample and the query text; calculating the segment boundary alignment loss of the video retrieval and localization model based on the fine-grained visual features of each video sample and the boundaries of the localization segments of each video sample for the matched query text, where the segment boundary alignment loss is used to enable the video retrieval and localization model to learn that: the semantic consistency segment corresponding to each sampled frame within the localization segment of the video sample for the matched query text should be aligned with the boundaries of the localization segment, where the semantic consistency segment corresponding to the sampled frame is a video segment composed of sampled frames that are adjacent to and semantically consistent with the sampled frame; calculating the inter-sample contrastive learning loss of the video retrieval and localization model based on the segment matching degrees of the localization segments of each video sample for different query texts and the visual similarities between the localization segments of different video samples for the same query text, where the inter-sample contrastive learning loss is used to enable the video retrieval and localization model to learn to distinguish visually similar video samples; calculating the total loss of the video retrieval and localization model based on the saliency alignment loss, the segment boundary alignment loss, and the inter-sample contrastive learning loss of the video retrieval and localization model.

[0006] Optionally, the step of calculating the saliency alignment loss of the video retrieval and localization model includes: for each video sample, using a first neural network to map the visual features of each sampled frame of the video sample and the text features of the query text matched by the video sample to the same shared space, and calculating the similarity between the mapped result of the visual feature of each sampled frame and the mapped result of the text feature of the query text as the matching score between the sampled frame and the query text; calculating the saliency alignment loss of the video retrieval and localization model based on the matching score between each sampled frame and the query text and the target value of the matching score corresponding to the sampled frame; where, for each sampled frame of the video sample, if the sampled frame is within the localization segment of the video sample for the matched query text, the target value of the matching score corresponding to the sampled frame is 1, otherwise it is 0.

[0007] Optionally, the steps of calculating the segment boundary alignment loss of the video retrieval and localization model include: taking each sampled frame within the localization segment of each video sample for the matched query text as the target sampled frame respectively, predicting the distance from the target sampled frame to the left boundary of its corresponding semantic consistency segment based on the visual features of the target sampled frame using a second neural network, and predicting the distance from the target sampled frame to the right boundary of its corresponding semantic consistency segment using a third neural network; determining the boundaries of the semantic consistency segment corresponding to the target sampled frame based on the distances from the target sampled frame to the left and right boundaries of its corresponding semantic consistency segment; calculating the segment boundary alignment loss of the video retrieval and localization model based on the boundaries of the semantic consistency segments corresponding to the sampled frames within the localization segments of each video sample for the matched query text and the boundaries of the localization segments of each video sample for the matched query text.

[0008] Optionally, the steps of calculating the inter-sample contrastive learning loss of the video retrieval and localization model include: for each training sample in the current batch of training samples, taking the segment matching degree of the video sample in the training sample for the matched query text as the first visual-text matching degree corresponding to the training sample; for each training sample in the current batch of training samples, taking the visual similarity between the localization segment of the video sample in the training sample for the matched query text and the localization segment of the video sample not in the training sample for the same query text as the segment visual similarity corresponding to the training sample; for each training sample in the current batch of training samples, taking the segment matching degree of the video sample in the training sample for the query text not in the training sample as the second visual-text matching degree corresponding to the training sample; for each training sample in the current batch of training samples, taking the segment matching degree of the video sample not in the training sample for the query text in the training sample as the third visual-text matching degree corresponding to the training sample; calculating the inter-sample contrastive learning loss of the video retrieval and localization model based on the first visual-text matching degree, segment visual similarity, second visual-text matching degree, and third visual-text matching degree corresponding to each training sample in the current batch of training samples.

[0009] Optionally, it further includes: updating the network parameters in the first neural network, the second neural network, and the third neural network based on the total loss of the video retrieval and localization model.

[0010] According to a second aspect of the embodiments of the present disclosure, there is provided a video retrieval and positioning method, including: receiving a query text; inputting the query text and each video in a video set into a video retrieval and positioning model to obtain the boundaries and segment matching degrees of the positioning segments of each video for the query text, where the positioning segment of a video for the query text is: the video segment with the highest matching degree with the query text in the video; screening out the positioning segments that meet a preset condition according to the segment matching degrees of the positioning segments of each video for the query text; outputting the boundaries of the positioning segments that meet the preset condition and the identification information of the videos to which they belong; where the video retrieval and positioning model is trained by executing the training method described above.

[0011] According to a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing instructions, which, when executed by a processor of an electronic device, enable the electronic device to execute the training method and / or the video retrieval and positioning method of the video retrieval and positioning model described above.

[0012] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, where the electronic device includes: at least one processor; at least one memory storing computer-executable instructions, where the computer-executable instructions, when run by the at least one processor, cause the at least one processor to execute the training method and / or the video retrieval and positioning method of the video retrieval and positioning model described above.

[0013] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer-executable instructions, which, when executed by at least one processor, implement the training method and / or the video retrieval and positioning method of the video retrieval and positioning model described above.

[0014] The video retrieval and positioning method, training method, electronic device, storage medium, and program product according to the exemplary embodiments of the present disclosure can improve the multi-modal alignment degree and retrieval performance of the model.

[0015] In the following description, some aspects and / or advantages of the general concept of the present disclosure will be set forth, and some aspects and / or advantages will be learned from the following description or the implementation of the general concept of the present disclosure. Description of the Drawings

[0016] From the following detailed description of the embodiments of the present application in conjunction with the drawings, these and / or other aspects and advantages of the present application will become clearer and easier to understand, where:

[0017] Figure 1 A flowchart showing a training method of a video retrieval and positioning model according to an exemplary embodiment of the present disclosure;

[0018] Figure 2 A flowchart showing a method for determining the total loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure;

[0019] Figure 3 A flowchart showing a method for calculating the saliency alignment loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure;

[0020] Figure 4 A flowchart showing a method for calculating the segment boundary alignment loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure;

[0021] Figure 5 A flowchart showing a method for calculating the inter-sample contrastive learning loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure;

[0022] Figure 6 An example of a training method for a video retrieval localization model according to an exemplary embodiment of the present disclosure;

[0023] Figure 7 A flowchart showing a video retrieval localization method according to an exemplary embodiment of the present disclosure;

[0024] Figure 8 A block diagram showing the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0025] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein the same reference numerals always refer to the same components. The following embodiments will be described with reference to the accompanying drawings to explain the present disclosure.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0027] It should be noted here that "at least one of several items" as used in this disclosure all represents three parallel cases, namely "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including both A and B. Another example, "performing at least one of step one and step two" means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.

[0028] Figure 1 The flowchart showing the training method of the video retrieval and localization model according to an exemplary embodiment of the present disclosure.

[0029] The video retrieval and localization model is used to locate the video segment with the highest matching degree with the query text (i.e., the localization segment) from each video, and provide the segment matching degree between each localization segment and the query text, so as to subsequently select the localization segments whose segment matching degree meets the preset conditions (for example, the top M with the highest segment matching degree or the segment matching degree exceeding a predetermined threshold) as the final retrieval and localization results.

[0030] The video retrieval and localization model includes a video representation module, a coarse-grained localization module, a fine-grained localization module, and a query text representation module. As an exemplary embodiment, the video retrieval and localization model can be a Joint Searching and Grounding (JSG) model.

[0031] Referring to Figure 1 , in step S101, obtain the current batch of training samples.

[0032] The current batch of training samples includes: a plurality of training samples, and each training sample (q, v) includes a video sample v and a query text q that matches the video sample (i.e., the text annotated for the video sample).

[0033] In step S102, input each video sample in the current batch of training samples into the video representation module to obtain the coarse-grained visual features and fine-grained visual features of each video sample.

[0034] The fine-grained visual features of the video sample include: the visual features of multiple sampled frames obtained by performing a first sampling process on the video sample, and the coarse-grained visual features of the video sample are obtained by performing a second sampling process on the multiple sampled frames.

[0035] In step S103, input each query text in the current batch of training samples into the query text representation module to obtain the text features of each query text.

[0036] In step S104, the coarse-grained visual features of each video sample and the text features of the matched query text are input into the coarse-grained localization module to obtain the segment localization focus position of the video sample for the matched query text.

[0037] In step S105, the segment localization focus position of each video sample for the matched query text, the fine-grained visual features, and the text features of the matched query text are input into the fine-grained localization module to obtain the boundaries and segment matching degrees of the localization segments of the video sample for the matched query text (i.e., the query text that belongs to the same training sample as the video sample).

[0038] The localization segment of the video sample for the matched query text is: the video segment with the highest matching degree to the query text in the video sample. The segment matching degree of the localization segment of the video sample for the matched query text is: the matching degree of the localization segment to the query text.

[0039] The segment localization focus position FP (Focus Point) of the video sample for the matched query text is used to specify: the central position of the localization segment of the video sample for the query text. That is, the coarse-grained localization module is used to give the central position of the localization segment, and the fine-grained localization module is used to give the boundary positions of the localization segment. As an example, as Figure 6 shown, the coarse-grained localization module can determine the central position of the localization segment through a multi-instance learning paradigm, and the fine-grained localization module can determine the localization segment through a multi-instance learning paradigm. By using the coarse-grained localization module and the fine-grained localization module, the localization of video segments can be carried out from coarse to fine.

[0040] In step S106, based on the fine-grained visual features of each video sample in the current batch of training samples, the text features of each query text, and the boundaries and segment matching degrees of the localization segments of each video sample for the matched query text, the total loss of the video retrieval localization model is determined.

[0041] The following will be combined with Figure 2 to describe an exemplary embodiment of step S106, which will not be elaborated here for the time being.

[0042] In step S107, based on the total loss of the video retrieval localization model, the model parameters of the video retrieval localization model are updated.

[0043] Figure 2 The flowchart showing the method for determining the total loss of the video retrieval localization model according to an exemplary embodiment of the present disclosure.

[0044] Refer to Figure 2, in step S201, based on the fine-grained visual features of each video sample, the text features of each query text, and the boundaries of the localization segments of each video sample for the matched query text, calculate the saliency alignment loss of the video retrieval localization model.

[0045] The saliency alignment loss is used to enable the video retrieval localization model to learn that: the degree of match between the sampled frames within the localization segment of the video sample for the matched query text and the query text should be significantly higher than the degree of match between other sampled frames (i.e., the sampled frames not within the above-mentioned localization segment) of the video sample and the query text.

[0046] Next, an exemplary embodiment of step S201 will be described in conjunction with Figure 3 and will not be elaborated here for the time being.

[0047] In step S202, based on the fine-grained visual features of each video sample and the boundaries of the localization segments of each video sample for the matched query text, calculate the segment boundary alignment loss of the video retrieval localization model.

[0048] The segment boundary alignment loss is used to enable the video retrieval localization model to learn that: for each sampled frame within the localization segment of the video sample for the matched query text, the corresponding semantic consistency segment should be aligned with the boundaries of the localization segment. The semantic consistency segment corresponding to a sampled frame is a video segment composed of sampled frames that are adjacent to and semantically consistent with the sampled frame.

[0049] Next, an exemplary embodiment of step S202 will be described in conjunction with Figure 4 and will not be elaborated here for the time being.

[0050] In step S203, based on the segment matching degrees of each video sample for different query texts and the visual similarities between the localization segments of different video samples for the same query text, calculate the inter-sample contrastive learning loss of the video retrieval localization model.

[0051] The inter-sample contrastive learning loss is used to enable the video retrieval localization model to learn to distinguish visually similar video samples.

[0052] Next, an exemplary embodiment of step S203 will be described in conjunction with Figure 5 and will not be elaborated here for the time being.

[0053] In step S204, based on the saliency alignment loss, segment boundary alignment loss, and inter-sample contrastive learning loss of the video retrieval localization model, calculate the total loss of the video retrieval localization model.

[0054] It should be understood that in addition to the above three losses, other types of losses can also be calculated and jointly used to calculate the total loss, and the present disclosure places no restrictions on this.

[0055] As an exemplary embodiment, the total loss can be obtained by weighted summation of the saliency alignment loss, the segment boundary alignment loss, the contrastive learning loss between samples, and other suitable types of losses.

[0056] Figure 3 A flowchart showing a method for calculating the saliency alignment loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure.

[0057] Referring to Figure 3 , in step S301, for each video sample, the visual features of each sampled frame of the video sample and the text features of the query text matched by the video sample are mapped to the same shared space using a first neural network, and the similarity between the mapping result of the visual features of each sampled frame and the mapping result of the text features of the query text is calculated as the matching score between the sampled frame and the query text.

[0058] As an example, the fine-grained visual features of video sample v can be expressed as: where L f represents the number of sampled frames obtained by performing the first sampling process on video sample v, and d represents the feature dimension. The text features of query text q matched by video sample v can be expressed as: The boundaries (i.e., start and end coordinates) of the localization segment of video sample v for query text q can be expressed as In addition, the coarse-grained visual features of video sample v can be expressed as:

[0059] As an example, first, the shared three-layer neural network FC shared (i.e., the first neural network) is used to map and F q to the same shared space to obtain the mapping result F v ′ of and the mapping result F q of q ′:

[0060]

[0061] Then, the cosine similarity can be calculated frame by frame as the matching score between the sampled frame and the query text

[0062]

[0063] In step S302, based on the matching score between each sampled frame and the query text and the target value of the matching score corresponding to the sampled frame, the saliency alignment loss of the video retrieval localization model is calculated.

[0064] For each sampled frame of a video sample, if the sampled frame is within the localization segment of the video sample for the matched query text, the target value of the matching score corresponding to the sampled frame is 1; otherwise, it is 0. For example, the target value of the matching score corresponding to the sampled frame i of the video sample v is:

[0065]

[0066] As an example, for the current batch of training samples B of size |B|, the binary loss function BCELoss can be used to calculate the saliency alignment loss L qsa :

[0067]

[0068] Figure 4 A flowchart showing a method for calculating the segment boundary alignment loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure.

[0069] Referring to Figure 4 , in step S401, for each sampled frame within the localization segment of each video sample for the matched query text, each sampled frame is used as a target sampled frame. Based on the visual features of the target sampled frame, a second neural network is used to predict the distance from the target sampled frame to the left boundary of its corresponding semantic consistency segment, and a third neural network is used to predict the distance from the target sampled frame to the right boundary of its corresponding semantic consistency segment.

[0070] In step S402, based on the distances from the target sampled frame to the left and right boundaries of its corresponding semantic consistency segment, the boundaries of the semantic consistency segment corresponding to the target sampled frame are determined.

[0071] As an example, for the target sampled frame i (with visual features denoted as ), a fully connected layer prediction network FC l (i.e., the second neural network) and a fully connected layer prediction network FC r (i.e., the third neural network) are used to predict: the relative distances from the frame to the left and right boundaries of its corresponding semantic consistency segment, and then convert them into the relative start and end coordinates of the semantic consistency segment corresponding to the frame and

[0072]

[0073] where σ represents the sigmoid activation function.

[0074] In step S403, based on the boundaries of the semantic consistency segments corresponding to each sampling frame within the localization segment of the query text matched by each video sample, and the boundaries of the localization segments of the query text matched by each video sample, the segment boundary alignment loss of the video retrieval localization model is calculated.

[0075] As an example, the boundaries of the localization segment The actual left and right start and end coordinates can be represented as l (q,v) and r (q,v) , and the calculation method of the segment boundary alignment loss EBA is as follows:

[0076]

[0077] where represents the number of frames in the localization segment.

[0078] Figure 5 FIG. shows a flowchart of a method for calculating the inter-sample contrastive learning loss of a video retrieval localization model according to an exemplary embodiment of the present disclosure.

[0079] Referring to Figure 5 , in step S501, for each training sample (q, v) in the current batch of training samples, the segment matching degree s(q, v) of the localization segment of the query text matched by the video sample in the training sample is used as the first visual-text matching degree corresponding to the training sample.

[0080] As an example, assume that the current batch of training samples includes three training samples (q1, v1), (q2, v2), (q3, v3). Taking the training sample (q1, v1) as an example, s(q, v) is specifically: s(q1, v1) (i.e., Figure 6 in ), which represents the matching degree of the localization segment of v1 for the matched q1 with q1.

[0081] In step S502, for each training sample (q, v) in the current batch of training samples, the visual similarity between the localization segment of the query text matched by the video sample in the training sample and the localization segment of the query text matched by the video samples in the non-training samples of the training sample is used as the segment visual similarity corresponding to the training sample.

[0082] Specifically, for each training sample (q, v) in the current batch of training samples, the video sample v in the training sample is regarded as a positive sample, and the video samples in the other training samples in the current batch of training samples (i.e., the video samples not in the training sample) are regarded as negative samples, marked as v - .

[0083] As an example, assume that the current batch of training samples includes three training samples (q1, v1), (q2, v2), and (q3, v3). Taking the training sample (q1, v1) as an example, v2 and v3 are the corresponding negative samples v - , Specifically, it includes: (i.e., Figure 6 in ) and (i.e., Figure 6 in ). Taking as an example, it represents the visual similarity between the localization segment of v1 for q1 and the localization segment of v2 for q1. Regarding the localization segment of v2 for q1, relevant information can be obtained through the following method: input the coarse-grained visual features of v2 and the text features of q1 into the coarse-grained localization module to obtain the segment localization focus position of v2 for q1; input the segment localization focus position of v2 for q1, the fine-grained visual features of v2, and the text features of q1 into the fine-grained localization module to obtain the boundary and segment matching degree of the localization segment of v2 for q1 (i.e., the matching degree s(q1, v2) between the localization segment of v2 for q1 and q1).

[0084] In step S503, for each training sample in the current batch of training samples, the segment matching degree s(q - , v) of the video sample in this training sample for the query text in non-this training sample is used as the second visual-text matching degree corresponding to this training sample.

[0085] Specifically, for each training sample (q, v) in the current batch of training samples, the query text q in this training sample is regarded as the positive label, and the query texts in other training samples in the current batch of training samples (i.e., the query texts in non-this training sample) are regarded as negative labels, marked as q - .

[0086] As an example, assume that the current batch of training samples includes three training samples (q1, v1), (q2, v2), and (q3, v3). Taking the training sample (q1, v1) as an example, q2 and q3 are the corresponding negative labels q - , s(q - , v1) specifically includes: s(q2, v1) (i.e., Figure 6 in ) and s(q3, v1) (i.e., Figure 6 in ). Taking s(q2, v1) as an example, it represents the matching degree between the localization segment of v1 for q2 and q2.

[0087] In step S504, for each training sample in the current batch of training samples, the segment matching degree s(q, v - ) of the video samples that are not the training sample for the query text in the training sample is used as the third visual-text matching degree corresponding to the training sample.

[0088] As an example, assume that the current batch of training samples includes three training samples (q1, v1), (q2, v2), and (q3, v3). Taking the training sample (q1, v1) as an example, v2 and v3 are the corresponding negative samples v - , and s(q1, v - ) specifically includes: s(q1, v2) (i.e., Figure 6 in ) and s(q1, v3) (i.e., Figure 6 in ). Taking s(q1, v2) as an example, it represents the matching degree between the localization segment of v2 for q1 and q1.

[0089] In step S505, based on the first visual-text matching degree, segment visual similarity, second visual-text matching degree, and third visual-text matching degree corresponding to each training sample in the current batch of training samples, calculate the inter-sample contrastive learning loss of the video retrieval and localization model.

[0090] As an example, the calculation method of the inter-sample contrastive learning loss L w-nce is as follows:

[0091]

[0092] In addition, as an exemplary embodiment, the training method of the video retrieval and localization model according to the exemplary embodiment of the present disclosure may further include: updating the network parameters in the first neural network, the second neural network, and the third neural network based on the total loss of the video retrieval and localization model.

[0093] The present disclosure designs two frame-level fine-grained auxiliary alignment tasks for the weak supervision video retrieval and segment localization problems: the query-guided saliency alignment task (QSA), the event-aware boundary alignment task (EBA), and a weighted contrastive loss (WCL) for inter-sample contrastive learning.

[0094] QSA uses shared weights to map video and text features into the same space, calculates the matching degree between frames and text, and uses regression scores for fine-grained alignment of frames and text to improve the multi-modal alignment degree and retrieval performance of the model.

[0095] EBA, on the other hand, guides each frame within the localization segment to predict a semantically consistent segment, promotes semantic consistency within the localization segment, and improves the multi-modal alignment degree and retrieval performance of the model by promoting semantic consistency within the localization segment.

[0096] As Figure 6 shown, WCL uses a visually similar weighted contrastive learning matrix to give more loss weights to hard samples, so as to learn to distinguish hard samples with similar visual features and improve the retrieval performance of the model.

[0097] Figure 7 shows a flowchart of a video retrieval and localization method according to an exemplary embodiment of the present disclosure.

[0098] Referring Figure 7 , in step S701, a query text is received.

[0099] In step S702, the query text and each video in the video set are input into the video retrieval and localization model to obtain the boundaries and segment matching degrees of the localization segments of each video for the query text.

[0100] The localization segment of a video for a query text is: the video segment in the video with the highest matching degree with the query text.

[0101] The video retrieval and localization model is trained by executing the training method described in the above exemplary embodiment.

[0102] In step S703, according to the segment matching degrees of the localization segments of each video for the query text, the localization segments that meet the preset conditions are filtered out.

[0103] As an example, the M localization segments with the highest segment matching degrees can be filtered out as the video segments matching the query text, or the localization segments with segment matching degrees exceeding a predetermined threshold can be filtered out as the video segments matching the query text. M is an integer greater than 0.

[0104] In step S704, the boundaries of the localization segments that meet the preset conditions and the identification information (e.g., name) of the videos to which they belong are output, that is, the retrieval and localization results for the received query text are output.

[0105] The present disclosure introduces QSA and EBA frame-level alignment tasks, improving the alignment degree of modalities and the semantic consistency of segments; and introduces a contrastive learning loss weighted by visual similarity, enabling the model to better distinguish difficult samples. This method significantly outperforms existing methods on the ActivityNet and Charades datasets.

[0106] Figure 8 A structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0107] Referring to Figure 8 , the electronic device includes: at least one memory 600 and at least one processor 700. A set of computer-executable instructions is stored in the at least one memory 600. When the set of computer-executable instructions is executed by the at least one processor 700, at least one of the following items is performed: the training method of the video retrieval and localization model as described in the above exemplary embodiment, and the video retrieval and localization method as described in the above exemplary embodiment.

[0108] As an example, the electronic device may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device does not have to be a single electronic device, and may also be any aggregate of devices or circuits that can individually or jointly execute the above instructions (or instruction sets). The electronic device may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0109] In the electronic device, the processor 700 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 700 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0110] The processor 700 may run the instructions or code stored in the memory 600, where the memory 600 may also store data. The instructions and data may also be sent and received via the network interface device over the network, where the network interface device may employ any known transmission protocol.

[0111] The memory 600 may be integrated with the processor 700. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. In addition, the memory 600 may include separate devices such as external disk drives, storage arrays, or other storage devices that may be used by any database system. The memory 600 and the processor 700 may be operatively coupled or may communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 700 can read files stored in the memory.

[0112] In addition, the electronic device may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device may be connected to each other via a bus and / or a network.

[0113] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are run by at least one processor, the at least one processor is caused to perform at least one of the following: the method for training a video retrieval and positioning model as described in the above exemplary embodiment, and the video retrieval and positioning method as described in the above exemplary embodiment. Examples of the computer-readable storage medium herein include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-RLTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as a client, host, proxy device, server, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0114] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, and the instructions in the computer program product may be executed by at least one processor to complete at least one of the following: the method for training a video retrieval and positioning model as described in the above exemplary embodiment, and the video retrieval and positioning method as described in the above exemplary embodiment.

[0115] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0116] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for a video retrieval and positioning model, characterized in that: The video retrieval positioning model includes a video representation module, a coarse-grained positioning module, a fine-grained positioning module, and a query text representation module, wherein the training method includes: Obtaining a current batch of training samples, wherein the current batch of training samples includes: a plurality of training samples, each training sample includes a video sample and a query text matching the video sample; Inputting each video sample in the current batch of training samples into the video representation module to obtain coarse-grained visual features and fine-grained visual features of each video sample, wherein the fine-grained visual features of the video sample include: visual features of multiple sampling frames obtained by performing a first sampling process on the video sample, and the coarse-grained visual features of the video sample are obtained by performing a second sampling process on the multiple sampling frames; Input each query text in the current batch of training samples into the query text representation module to obtain the text features of each query text; Inputting the coarse-grained visual features of each video sample and the text features of the matched query text into a coarse-grained positioning module to obtain a segment positioning focus position of the video sample for the matched query text; Inputting the segment positioning focus position, fine-grained visual features, and text features of the matched query text of each video sample into a fine-grained positioning module to obtain the boundary and segment matching degree of the positioning segment of the video sample for the matched query text, wherein the positioning segment of the video sample for the matched query text is: the video segment in the video sample with the highest matching degree with the query text; Determine the total loss of the video retrieval and localization model based on the fine-grained visual features of each video sample in the current batch of training samples, the text features of each query text, the boundaries of the localization segments of each video sample for the matched query text, and the segment matching degree; Based on the total loss of the video retrieval and positioning model, the model parameters of the video retrieval and positioning model are updated.

2. The training method according to claim 1, characterized in that: The steps to determine the total loss of the video retrieval localization model include: Based on the fine-grained visual features of each video sample, the text features of each query text, and the boundaries of the positioning segments of each video sample for the matched query text, the saliency alignment loss of the video retrieval positioning model is calculated, wherein the saliency alignment loss is used to enable the video retrieval positioning model to learn that the matching degree between the sampled frames in the positioning segments of the video sample for the matched query text and the query text should be significantly higher than the matching degree between other sampled frames of the video sample and the query text; Based on the fine-grained visual features of each video sample and the boundary of the positioning segment of each video sample for the matched query text, the segment boundary alignment loss of the video retrieval positioning model is calculated, wherein the segment boundary alignment loss is used to enable the video retrieval positioning model to learn that the semantically consistent segment corresponding to each sampling frame in the positioning segment of the video sample for the matched query text should be aligned with the boundary of the positioning segment, wherein the semantically consistent segment corresponding to the sampling frame is a video segment composed of sampling frames that are adjacent to the sampling frame and semantically consistent with the sampling frame; Based on the segment matching degree of each video sample for the localization segments of different query texts and the visual similarity between the localization segments of different video samples for the same query text, the inter-sample contrastive learning loss of the video retrieval localization model is calculated, wherein the inter-sample contrastive learning loss is used to enable the video retrieval localization model to learn: distinguish visually similar video samples; Based on the saliency alignment loss, segment boundary alignment loss, and inter-sample contrastive learning loss of the video retrieval and positioning model, the total loss of the video retrieval and positioning model is calculated.

3. The training method according to claim 2, characterized in that: The steps to calculate the saliency alignment loss for the video retrieval localization model include: For each video sample, use the first neural network to map the visual features of each sampling frame of the video sample and the text features of the query text matched by the video sample to the same shared space, and calculate the similarity between the mapping result of the visual features of each sampling frame and the mapping result of the text features of the query text as the matching score between the sampling frame and the query text; Based on the matching score between each sample frame and the query text and the matching score target value corresponding to the sample frame, the saliency alignment loss of the video retrieval and positioning model is calculated; For each sampling frame of the video sample, if the sampling frame is in the positioning segment of the video sample for the matched query text, the matching score target value corresponding to the sampling frame is 1, otherwise it is 0.

4. The training method according to claim 3, characterized in that: The steps to calculate the segment boundary alignment loss for the video retrieval localization model include: Each sampling frame in the positioning segment of each video sample corresponding to the matched query text is respectively taken as a target sampling frame, and based on the visual features of the target sampling frame, a second neural network is used to predict the distance from the target sampling frame to the left boundary of its corresponding semantically consistent segment, and a third neural network is used to predict the distance from the target sampling frame to the right boundary of its corresponding semantically consistent segment; Determine the boundary of the semantically consistent segment corresponding to the target sampling frame based on the distance from the target sampling frame to the left boundary and the right boundary of the semantically consistent segment corresponding to the target sampling frame; Based on the boundaries of semantically consistent segments corresponding to each sampling frame within the positioning segment of each video sample for the matched query text and the boundaries of the positioning segment of each video sample for the matched query text, the segment boundary alignment loss of the video retrieval positioning model is calculated.

5. The training method according to claim 2, characterized in that: The steps to calculate the inter-sample contrastive learning loss for the video retrieval and localization model include: For each training sample in the current batch of training samples, the segment matching degree of the video sample in the training sample to the positioning segment of the matched query text is used as the first visual text matching degree corresponding to the training sample; For each training sample in the current batch of training samples, the visual similarity between the positioning segment of the video sample in the training sample for the matched query text and the positioning segment of the video sample not in the training sample for the query text is used as the segment visual similarity corresponding to the training sample; For each training sample in the current batch of training samples, the segment matching degree of the video sample in the training sample to the positioning segment of the query text that is not in the training sample is used as the second visual text matching degree corresponding to the training sample; For each training sample in the current batch of training samples, the segment matching degree of the video sample not in the training sample to the positioning segment of the query text in the training sample is used as the third visual text matching degree corresponding to the training sample; Based on the first visual text matching degree, fragment visual similarity, second visual text matching degree, and third visual text matching degree corresponding to each training sample in the current batch of training samples, the inter-sample contrast learning loss of the video retrieval and positioning model is calculated.

6. The training method according to claim 4, characterized in that: Also includes: Based on the total loss of the video retrieval and positioning model, the network parameters in the first neural network, the second neural network and the third neural network are updated.

7. A video retrieval and positioning method, characterized in that: include: Receive query text; Input the query text and each video in the video set into the video retrieval positioning model to obtain the boundary and segment matching degree of each video positioning segment for the query text, wherein the positioning segment of the video for the query text is: the video segment with the highest matching degree with the query text in the video; According to the segment matching degree of each video to the positioning segment of the query text, the positioning segment that meets the preset conditions is screened out; Output the boundary of the positioning segment that meets the preset conditions and the identification information of the video to which it belongs; Wherein, the video retrieval and positioning model is trained by executing the training method described in any one of claims 1 to 6.

8. A computer-readable storage medium storing instructions, characterized in that: When the instruction is executed by a processor of an electronic device, the electronic device is enabled to execute the training method of the video retrieval and positioning model as described in any one of claims 1 to 6 and / or the video retrieval and positioning method as described in claim 7.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the training method of the video retrieval and positioning model as described in any one of claims 1 to 6 and / or the video retrieval and positioning method as described in claim 7.

10. A computer program product comprising computer executable instructions, characterized in that: When the computer executable instructions are executed by at least one processor, the training method of the video retrieval and positioning model according to any one of claims 1 to 6 and / or the video retrieval and positioning method according to claim 7 are implemented.