Video library clip retrieval method and device for subject-text joint query

By adopting the joint subject-text query method in the video library, combining video characterization sequence and subject information, efficient retrieval and precise positioning of video library fragments is achieved, and the problem of inability to distinguish video subjects and lack of generalization capabilities in the existing technology is solved, and the search accuracy and flexibility are improved.

CN120086410APending Publication Date: 2025-06-03TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510263770.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art cannot effectively distinguish the specific subjects in the video library segment retrieval, which makes it difficult to ensure that the search results contain designated subjects, and lack generalization capabilities, so that they cannot process unseen subjects.

Method used

The subject-text joint query method is adopted to obtain the representation sequence of videos in the video library, combine subject information and text query, and perform interactive calculations to fuse context semantics and calculate the similarity between video and query, thereby accurately positioning related video clips.

Benefits of technology

It realizes efficient search of video clips containing designated subjects in the video library, has better generalization ability, can handle unseen subjects, and significantly improves the accuracy and flexibility of video content retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086410A_ABST
    Figure CN120086410A_ABST
Patent Text Reader

Abstract

The invention provides a video library fragment retrieval method and device for subject-text joint query, and relates to the technical field of computer vision. The method comprises the following steps: acquiring a first video representation sequence of each video in a video library; performing representation extraction on the subject query and the text query in the subject-text joint query to obtain a first query representation sequence; performing interactive calculation on the first video representation sequence to obtain a second video representation sequence fused with context semantics, and performing interactive calculation on the first query representation sequence to obtain a second query representation sequence; according to the semantic representation in the second query representation sequence and the second video representation sequence, calculating the similarity between the subject-text joint query and each video, and taking the video corresponding to the maximum similarity as a retrieval video; and according to the second query representation sequence and the starting and ending timestamps of a second video representation sequence prediction fragment corresponding to the retrieval video, obtaining a retrieval video fragment corresponding to the subject-text joint query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a method and apparatus for retrieving video library segments by subject-text joint query. Background Art

[0002] The Video Corpus Grounding (VCG) task has received extensive attention in the field of multimedia research. This task aims to retrieve the video segments most relevant to a text description from a large-scale video library. This task has important potential in multiple downstream applications such as video editing, recommendation systems, and content creation.

[0003] However, related technologies retrieve video segments based on text descriptions and cannot effectively distinguish specific subjects in the video. For example, one may hope to find video segments of a certain movie star dancing, but related methods can only locate the dancing segments and cannot ensure that the retrieved videos contain this movie star. Therefore, related technologies have the following problems: it is difficult to ensure that the retrieval results contain the specified subject, and they lack generalization ability and cannot handle unseen subjects. Summary of the Invention

[0004] In view of the above problems, embodiments of this application provide a method and apparatus for retrieving video library segments by subject-text joint query to overcome or at least partially solve the above problems.

[0005] In a first aspect of the embodiments of this application, a method for retrieving video library segments by subject-text joint query is disclosed. The method includes: Obtaining a first video representation sequence of each video in the video library, where each video representation in the first video representation sequence includes subject information and other image information except the subject information; Respectively performing representation extraction on the subject query and the text query in the subject-text joint query to obtain a first query representation sequence; Performing interaction calculation on the first video representation sequence to obtain a second video representation sequence that integrates context semantics, and performing interaction calculation on the first query representation sequence to obtain a second query representation sequence, where the second query representation sequence includes semantic representations with subject-text joint query semantics; Calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and taking the video corresponding to the maximum similarity as the retrieved video; Predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain the retrieved video segment corresponding to the subject-text joint query.

[0006] Optionally, obtaining the first video representation sequence of each video in the video library includes: Uniformly sampling the video to obtain multiple frames of images; Performing main body information extraction on each frame of image, and performing other image information extraction on each frame of image to obtain the video representation corresponding to the image; Concatenating the video representations corresponding to the multiple frames of images according to the frame numbers to obtain the first video representation sequence.

[0007] Optionally, the main body information includes face information, and the other image information includes action information and environmental information; performing main body information extraction on each frame of image, and performing other image information extraction on each frame of image to obtain the video representation corresponding to the image, including: Performing face detection and face representation extraction on the image, and converting the dimension of the extracted face representation to the feature dimension required by the large model to obtain the final face representation; Performing other image representation extraction on the image, and converting the dimension of the extracted other image representation to the feature dimension required by the large model to obtain the final other image representation; Obtaining the video representation corresponding to the image according to the final face representation and the final other image representation.

[0008] Optionally, the main body query is an image containing the main body; performing representation extraction on the main body query and the text query in the main body-text joint query respectively to obtain the first query representation sequence, including: Performing face detection and face representation extraction on the main body query, and converting the dimension of the extracted face representation to the feature dimension required by the large model to obtain the main body query representation; Converting the text query into a word vector to obtain the text query representation; Obtaining the first query representation sequence according to the main body query representation and the text query representation.

[0009] Optionally, performing interactive calculation on the first query representation sequence to obtain the second query representation sequence, including: Adding a semantic representation to the first query representation sequence, where the semantic representation is used to learn the semantics of each query representation in the first query representation sequence; Performing interactive calculation on the first query representation sequence with the added semantic representation to obtain the second query representation sequence.

[0010] Optionally, calculating the similarity between the main body-text joint query and each video according to the semantic representation and the second video representation sequence, including: Calculate the similarity between the semantic representation and each video representation in the second video representation sequence in turn to obtain a plurality of similarity values; Take the maximum value among the plurality of similarity values as the similarity between the subject-text joint query and the video.

[0011] Optionally, predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain the retrieved video segment corresponding to the subject-text joint query, including: Concatenate the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain a first concatenated representation sequence; Add a fusion representation to the first concatenated representation sequence and perform interactive calculation on the first concatenated representation sequence with the fusion representation added, where the fusion representation is used to learn the semantics of each representation in the first concatenated representation sequence, to obtain a second concatenated representation sequence; Predict the start and end timestamps of the segment according to the fusion representation in the second concatenated representation sequence to obtain the target video segment for answering the subject-text joint query.

[0012] Optionally, the video library segment retrieval method for the subject-text joint query is implemented based on a retrieval model, and the retrieval model includes at least a large model; Among them, performing interactive calculation on the first video representation sequence, performing interactive calculation on the first video representation sequence, and calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and taking the video corresponding to the maximum similarity as the retrieved video is implemented using the layers at the bottom of the large model; Predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video is implemented using the layers at the top of the large model, where L represents the total number of layers of the large model.

[0013] Optionally, the loss function for training the retrieval model includes a first loss, a second loss, a third loss, and a fourth loss; Among them, the first loss represents the contrastive learning loss of the similarity between the subject-text joint query sample and the video sample; The second loss represents the L1 loss between the start and end timestamps of the predicted video segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; The third loss represents the generalized intersection over union loss between the start and end timestamps of the predicted segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; The fourth loss characterizes whether each frame image is a frame classification loss of the video segment to be retrieved.

[0014] In a second aspect of the embodiments of the present application, a video library segment retrieval device for subject-text joint query is disclosed. The device includes: A feature acquisition module, configured to acquire a first video feature sequence of each video in the video library, where each video feature in the first video feature sequence includes subject information and other image information except the subject information; A feature extraction module, configured to perform feature extraction on the subject query and the text query in the subject-text joint query respectively to obtain a first query feature sequence; An interaction calculation module, configured to perform interaction calculation on the first video feature sequence to obtain a second video feature sequence integrating context semantics, and perform interaction calculation on the first query feature sequence to obtain a second query feature sequence, where the second query feature sequence includes semantic features with subject-text joint query semantics; A similarity calculation module, configured to calculate the similarity between the subject-text joint query and each video according to the semantic features and the second video feature sequence, and use the video corresponding to the maximum similarity as the retrieved video; A segment prediction module, configured to predict the start and end timestamps of the segment according to the second query feature sequence and the second video feature sequence corresponding to the retrieved video, and obtain the retrieved video segment corresponding to the subject-text joint query.

[0015] In a third aspect of the embodiments of the present application, an electronic device is disclosed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the video library segment retrieval method for subject-text joint query described in the first aspect of the embodiments of the present application are implemented.

[0016] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video library segment retrieval method for subject-text joint query described in the first aspect of the embodiments of the present application are implemented.

[0017] In a fifth aspect of the embodiments of the present application, a computer program product is disclosed, including a computer program. When the computer program is executed by a processor, the steps of the video library segment retrieval method for subject-text joint query described in the first aspect of the embodiments of the present application are implemented.

[0018] The embodiments of the present application have the following advantages: In the embodiments of the present application, video library segment retrieval is performed based on subject-text joint query. By processing each video in the video library into a first video representation sequence, feature extraction is performed on the subject query and text query in the subject-text joint query to obtain a first query representation sequence; interactive calculation is performed on the first video representation sequence to obtain a second video representation sequence that integrates context semantics, and interactive calculation is performed on the first query representation sequence to obtain a second query representation sequence. Since the second query representation sequence includes semantic representations with subject-text joint query semantics, the similarity between the subject-text joint query and each video can be calculated based on this semantic representation and the second video representation sequence, and then the video corresponding to the maximum similarity is used as the retrieved video; finally, the retrieved video segment corresponding to the subject-text joint query is accurately located from the retrieved video.

[0019] This method performs video library segment retrieval based on subject-text joint query. It can not only achieve efficient retrieval through text description, but also combine subject information (subject query) to achieve more accurate video positioning. Moreover, this method has good generalization ability and can perform accurate retrieval for unseen subjects. Thus, this method can significantly improve the accuracy and flexibility of video content retrieval and meet the more personalized needs of users. Brief Description of the Drawings

[0020] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of the steps of a method for retrieving video library segments by subject-text joint query provided by the embodiments of the present application; Figure 2 It is a schematic diagram of a task for retrieving video library segments by subject-text joint query provided by the embodiments of the present application; Figure 3 It is a schematic diagram of the structure of a retrieval model provided by the embodiments of the present application; Figure 4 It is a schematic diagram of the structure of a device for retrieving video library segments by subject-text joint query provided by the embodiments of the present application; Figure 5 It is a schematic diagram of the structure of an electronic device provided by the embodiments of the present application. Detailed Embodiments

[0022] To make the above objects, features, and advantages of the present application more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0023] In related technologies, fragment localization can only be based on text descriptions and fails to consider the visual information of specific subjects in the video. To retrieve relevant fragments in a video library, current methods usually adopt a two-stage approach of first retrieving relevant videos and then performing fragment localization. For example, a video fragment retrieval method uses a two-tower model to jointly learn video retrieval and fragment localization objectives; a video fragment retrieval method introduces contrastive learning between query-video and query-frame to encourage alignment at different granularities; a video fragment retrieval method uses causal inference to prevent the model from learning incorrect retrieval biases.

[0024] Although these methods have achieved good performance in the video library fragment retrieval task, there are still the following problems: it is difficult to ensure that the retrieval results contain the specified subject, and they lack generalization ability. They can only represent the subject by a person's name, and if the model has not seen this person's name (subject name) during training, it is impossible to correspond this person's name to the subject in the video during inference. Moreover, because they cannot accept the image input of the subject (i.e., cannot handle the joint subject and text query), these methods cannot be directly applied to the Identity-Text Video Corpus Grounding (ITVCG) task of video library fragment retrieval with subject-text joint queries.

[0025] To overcome the limitations of related technologies, the embodiments of the present application provide a video library fragment retrieval method with subject-text joint queries, aiming to locate the most relevant fragments in the video library by combining text descriptions and subject images. Compared with traditional video library fragment retrieval, the method implemented in the present application has a wider application scenario. It can not only achieve efficient retrieval through text descriptions but also combine subject information to achieve more accurate video localization, significantly improving the accuracy and flexibility of video content retrieval and meeting the more personalized needs of users.

[0026] The following will describe the video library fragment retrieval method with subject-text joint queries in the embodiments of the present application with reference to the accompanying drawings.

[0027] Refer to Figure 1 as shown, Figure 1 is the flowchart of the steps of a video library fragment retrieval method with subject-text joint queries provided by the embodiments of the present application. As Figure 1As shown in the figure, the method for retrieving video library segments by subject-text joint query may include steps S110 to S150: Step S110: Obtain the first video representation sequence of each video in the video library, where each video representation in the first video representation sequence contains subject information and other image information except the subject information.

[0028] There are multiple videos in the video library. Feature extraction is performed on each video in the video library respectively to obtain the corresponding first video representation sequence for the video. The video and the first video representation sequence are in one-to-one correspondence. For example, if there are 5 videos in the video library, there are 5 first video representation sequences corresponding to each of the 5 videos.

[0029] The first video representation sequence includes multiple video representations, and each video representation corresponds to a frame image in the video. Each video representation contains subject information and other image information except the subject information. Among them, the subject information may be a face, and the other image information may be action information, environmental information, etc.

[0030] Step S120: Perform feature extraction on the subject query and the text query in the subject-text joint query respectively to obtain the first query representation sequence.

[0031] Among them, the subject-text joint query includes a subject query and a text query. The subject query may be an image of the subject, and the text query refers to a descriptive text about actions and / or scenes. The subject-text joint query can be expressed as <subject 1> and <subject 2> + "text description", where <subject 1> and <subject 2> are the subject queries, and "text description" is the text query. In some embodiments, the subject query may be an image containing the subject. For example, a picture of a movie star + "text" can be used as the subject-text joint query.

[0032] Perform feature extraction on the subject query and the text query respectively. For example, perform feature extraction on the subject query to obtain the subject query representation, perform feature extraction on the text query to obtain the text query representation, and then obtain the first query representation sequence according to the subject query representation and the text query representation. That is to say, the first query representation sequence includes multiple query representations, and each query representation is a subject query representation or a text query representation.

[0033] Step S130: Perform interactive calculation on the first video representation sequence to obtain a second video representation sequence that integrates context semantics, and perform interactive calculation on the first query representation sequence to obtain a second query representation sequence. The second query representation sequence includes semantic representations with subject-text joint query semantics.

[0034] After obtaining the first video representation sequence and the first query representation sequence, it is necessary to align the representations to achieve fast video retrieval. In some embodiments, the layers (i.e., the layers close to the input end) can be used to perform interactive calculations on the first video representation sequence and the first query representation sequence respectively.

[0035] Among them, each video representation in the first video representation sequence has the semantics corresponding to a frame of image. By performing interactive calculations on the first video representation sequence, each video representation can fuse the context semantics (or the semantics before and after), so as to obtain the second video representation sequence that fuses the context semantics.

[0036] The semantics of the specific word or subject corresponding to each query representation in the first query representation sequence need to use the total semantics of the entire subject-text joint query for video segment retrieval. Interactive calculations can be performed based on the first query representation sequence to obtain the second query representation sequence, and the second query representation sequence includes a semantic representation that learns the subject-text joint query semantics.

[0037] Step S140: Calculate the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and use the video corresponding to the maximum similarity as the retrieved video.

[0038] Since the second query representation sequence includes a semantic representation with the subject-text joint query semantics, the similarity between the subject-text joint query and each video can be calculated based on this semantic representation and the second video representation sequence. The greater the similarity, the more the subject-text joint query matches the video. Furthermore, the video corresponding to the maximum similarity is used as the retrieved video.

[0039] For example, there are second video representation sequences corresponding to 5 videos. If the similarity between the subject-text joint query and the 5 videos is calculated as 0.1, 0.5, 0.6, 0.8, 0.2 respectively according to the semantic representation and each second video representation sequence, then the video corresponding to the similarity of 0.8 is used as the retrieved video.

[0040] Step S150: Predict the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video, and obtain the retrieved video segment corresponding to the subject-text joint query.

[0041] After completing the video retrieval, it is also necessary to locate the relevant segments from the retrieved videos. Specifically, an interaction calculation is performed based on the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain a fusion representation that contains all semantics (i.e., the semantics of the subject-text joint query and the semantics of the retrieved video). Then, the segment locator (a three-layer fully connected neural network) predicts the start and end timestamps of the segment based on this fusion representation, obtaining the retrieved video segment corresponding to the subject-text joint query, that is, the retrieved video segment for the subject-text joint query.

[0042] Adopting the technical solution of the embodiment of the present application, the most relevant segments in the video library are located by combining text queries and subject queries. As Figure 2 shown, for traditional text-based queries, if the input query is "Cars and Beckett are dancing together", the model may not be able to answer; while the method of the embodiment of the present application can quickly retrieve relevant segments by uploading a picture of a movie star and inputting a text representation. For example, inputting a subject-text joint query of "Images of Cars and Images of Beckett + dancing together", the retrieval model outputs relevant video segments. Therefore, the method implemented in the present application has a wider application scenario. This method performs video library segment retrieval based on subject-text joint queries, which can not only achieve efficient retrieval through text descriptions, but also combine subject information (subject queries) to achieve more accurate video localization. Moreover, this method has good generalization ability and can perform accurate retrieval for unseen subjects. In this way, this method can significantly improve the accuracy and flexibility of video content retrieval and meet the more personalized needs of users.

[0043] Combining the above embodiments, in one implementation manner, the embodiment of the present application also provides a method for retrieving video library segments for subject-text joint queries. In this method, the step of "obtaining the first video representation sequence of each video in the video library" in the above step S110 specifically includes sub-steps S110-1 to step S110-3: Step S110-1: Uniformly sample the video to obtain multiple frames of images.

[0044] In some embodiments, the time lengths of all videos in the video library may be the same. Therefore, for all videos in the video library, uniform sampling can be performed according to the same number of frames. For example, all videos are uniformly sampled 100 frames. In some embodiments, the time lengths of all videos in the video library may not be the same. Therefore, for each video in the video library, different numbers of multiple frames of images can be collected according to different time lengths.

[0045] Step S110-2: Extract subject information from each frame of image, and extract other image information from each frame of image to obtain the video representation corresponding to the image.

[0046] For each frame of image, the subject information and other image information therein need to be considered. Since the subject information and other image information are obtained in different ways, the subject information extraction can be performed on each frame of image respectively, and the other image information extraction can be performed on each frame of image, so as to obtain the subject representation and other image representations, and further obtain the video representation corresponding to the image.

[0047] In some embodiments, the subject information includes face information, and the other image information includes action information and environmental information. In the above step S110-2, "performing subject information extraction on each frame of image, and performing other image information extraction on each frame of the first video representation sequence of images to obtain the video representation corresponding to the image" specifically includes sub-steps S110-2-1 to step S110-2-3: Step S110-2-1: Perform face detection and face representation extraction on the image, and convert the dimension of the extracted face representation to the feature dimension required by the large model to obtain the final face representation.

[0048] For the subject information, perform face detection on the image, and perform human representation extraction on the detected face through a pre-trained face representation extractor. For the finally obtained video representation, it can be processed by the large model. Since the feature dimension of the face representation is different from the feature dimension of the large model, an adapter (i.e., a two-layer fully connected neural network) is used to perform dimension conversion, so that the dimension of the face representation is converted into the feature dimension required by the large model, and then the final face representation is obtained.

[0049] Step S110-2-2: Perform other image representation extraction on the image, and convert the dimension of the extracted other image representation to the feature dimension required by the large model to obtain the final other image representation.

[0050] For the other image information, directly use the pre-trained visual encoder to perform representation extraction. Similarly, since the feature dimension of the extracted other image representation is different from the feature dimension of the large model, an adapter is used for dimension conversion to convert the dimension of the other image representation into the feature dimension required by the large model, and then the final other image representation is obtained.

[0051] Step S110-2-3: Obtain the video representation corresponding to the image according to the final face representation and the final other image representation.

[0052] Fuse (add) the two representations obtained through the above steps to obtain the video representation corresponding to the image.

[0053] Step S110-3: Concatenate the video representations corresponding to the multiple frames of images according to the frame numbers to obtain the first video representation sequence.

[0054] Each video corresponds to a first video representation sequence, which corresponds to multiple frames of images of the corresponding video. After obtaining the video representations corresponding to each frame of image, they are concatenated according to the frame numbers to obtain the first video representation sequence corresponding to the video. For example, there are 100 frames of images, and the video representation corresponding to the i-th frame of image is , where i represents the image frame number, and 0 represents the input representation for the large model (i.e., the subsequent processing model). Then the first video representation sequence can be expressed as , and represent the 1st, 2nd, and 100th video representations respectively.

[0055] Adopting the technical solution of the embodiment of the present application, by uniformly sampling the video and respectively extracting the subject information and other image information from the sampled multiple frames of images, the video representations corresponding to the images are obtained, and then the first video representation sequence corresponding to the video is obtained based on the video representations corresponding to the multiple frames of images. Since each video representation contains subject information and other image information, when performing video segment retrieval, the subject with features can be identified based on the subject information in the video representation, and the most relevant retrieval video segment for the subject-text joint query can be determined by combining the other image information in the video representation.

[0056] Combined with the above embodiments, in one implementation manner, the embodiment of the present application further provides a method for retrieving video library segments for subject-text joint query. In this method, the subject query is an image containing a subject. The step of "respectively performing representation extraction on the subject query and the text query in the subject-text joint query to obtain a first query representation sequence" in the above step S120 specifically includes sub-steps S120-1 to step S120-3: Step S120-1: Perform face detection and face representation extraction on the subject query, and convert the dimension of the extracted face representation to the feature dimension required by the large model to obtain the subject query representation.

[0057] Step S120-2: Perform word vector conversion on the text query to obtain the text query representation.

[0058] Step S120-3: Obtain the first query representation sequence according to the subject query representation and the text query representation.

[0059] Exemplarily, for a subject-text joint query, such as "<Subject 1> and <Subject 2> are dancing", where <Subject 1> and <Subject 2> are both in the form of images.

[0060] For the subject query, the same processing method as that for the video can be adopted, that is, using a pre-trained face feature extractor to extract human features, and adding an adapter for dimension conversion to convert the dimension of the extracted face features into the feature dimension required by the large model, so as to obtain the subject query representation.

[0061] For the text query, since the large model itself can accept text input, the processing method of the large model for the text is adopted, that is, converting the text into a representation through word vectors to obtain the text query representation.

[0062] Finally, according to the subject query representation and the text query representation, the first query representation sequence is obtained , where n is the length of the query, and for different subject-text joint queries, the lengths of the corresponding queries may be different. respectively represent the first, second, and nth query representations, and each query representation is a subject query representation or a text query representation.

[0063] By adopting the technical solution of the embodiment of the present application, the subject query and the text query in the subject-text joint query are converted into corresponding representations, and then the subject in the first query representation sequence can be used to identify each video in the video library based on the subject query representation, and efficient retrieval can be realized based on the text query representation, so as to determine the retrieval video segment most relevant to the subject-text joint query.

[0064] Combining the above embodiments, in one implementation manner, the embodiment of the present application further provides a method for retrieving video library segments for subject-text joint queries. In this method, in the above step S130, "performing interaction calculation on the first query representation sequence to obtain a second query representation sequence" specifically includes sub-steps S130-1 to step S130-2: Step S130-1: Adding a semantic representation to the first query representation sequence, where the semantic representation is used to learn the semantics of each query representation in the first query representation sequence.

[0065] Step S130-2: Performing interaction calculation on the first query representation sequence with the added semantic representation to obtain a second query representation sequence.

[0066] For the query representation sequence , each query representation only represents its corresponding word or subject, and the overall semantics of the entire subject-text joint query is required for video retrieval. Therefore, a semantic representation is added at the end of the first query representation sequence , and interaction calculation is performed through the first query representation sequence with the added semantic representation to obtain a second query representation sequence.

[0067] Specifically, the interactive calculation of the first query representation sequence with added semantic representation can be implemented through the layer at the bottom of the large model: (1) Among them, represents the second query representation sequence. For the large model, the input and output sequence lengths are the same, and the representations within the sequence can perform interactive calculations. Therefore, actually contains the semantics of the entire subject-text joint query.

[0068] Adopting the technical solution of the embodiment of the present application, by adding semantic representation to the first query representation sequence for interactive calculation, the semantic representation can learn the semantics of the entire subject-text joint query, so that accurate video segment retrieval can be achieved based on this semantic representation.

[0069] Furthermore, for the first video representation sequence , each video representation only corresponds to the semantics of its own frame of image, while video retrieval hopes it can fuse the semantics before and after. Therefore, in the above step S130, "performing interactive calculation on the first video representation sequence to obtain a second video representation sequence that fuses context semantics" can also be implemented through the layer at the bottom of the large model: (2) Among them, represents the second video representation sequence.

[0070] In this way, a semantic representation containing the subject-text joint query, and a second video representation sequence that fuses context semantics are obtained, so that the retrieval video segment corresponding to the subject-text joint query can be accurately retrieved based on the semantic representation and the second video representation sequence subsequently.

[0071] In practical applications, for a large number of videos in the video library, the interactive calculation of the video representation sequence needs to be performed using the above formula (2), resulting in a large amount of computational overhead. However, it is noted that formula (2) has nothing to do with the subject-text joint query. Therefore, the video platform can calculate and store the results in advance using formula (2) (i.e., store the second video representation sequences of each video), so as to achieve fast video segment retrieval.

[0072] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a method for retrieving video library segments for subject-text joint query. In this method, in step S140 above, "calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence" specifically includes sub-steps S140-1 to step S140-2: Step S140-1: Calculate the similarity between the semantic representation and each video representation in the second video representation sequence in turn to obtain a plurality of similarity values.

[0073] Step S140-2: Use the maximum value among the plurality of similarity values as the similarity between the subject-text joint query and the video.

[0074] In the embodiments of the present application, the similarity between the subject-text joint query and the video is defined as: the maximum value of the similarity between the subject-text joint query and each frame image. And the similarity between the subject-text joint query and each frame image is the similarity between the semantic representation and each video representation in the second video representation sequence.

[0075] Exemplarily, taking the cosine similarity as an example, the similarity between the subject-text joint query and the video can be expressed as: (3) Where Q represents the subject-text joint query, represents the video, represents the semantic representation in the second query representation sequence, represents the i-th video representation in the second video representation sequence.

[0076] Adopting the technical solution of the embodiments of the present application, based on the maximum value of the similarity between the semantic representation and each video representation in the second video representation sequence as the similarity between the subject-text joint query and the video, and then calculating the similarity between the subject-text joint query and each video, and finally using the video corresponding to the maximum similarity as the retrieved video.

[0077] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a method for retrieving video library segments for subject-text joint query. In this method, in step S150 above, "predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain the retrieved video segment corresponding to the subject-text joint query" specifically includes sub-steps S150-1 to step S150-3: Step S150-1: Concatenate the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain a first concatenated representation sequence.

[0078] Step S150-2: Add a fusion representation to the first spliced representation sequence, and perform interactive calculation on the first spliced representation sequence with the fusion representation added to obtain a second spliced representation sequence. The fusion representation is used to learn the semantics of each representation in the first spliced representation sequence.

[0079] Step S150-3: Predict the start and end timestamps of the segment based on the fusion representation in the second spliced representation sequence to obtain the retrieved video segment corresponding to the subject-text joint query.

[0080] In the embodiments of the present application, after the complete video retrieval, it is necessary to locate the relevant segments in the retrieved video. For this purpose, the second query representation sequence and the second video representation sequence corresponding to the retrieved video are spliced together. In order to accurately predict the start and end timestamps of the segment based on all the semantics (i.e., the semantics of the subject-text joint query and the semantics of the retrieved video), a fusion representation is added to the first spliced representation sequence, and then in the subsequent interactive calculation process, the semantics of each representation in the first spliced representation sequence (i.e., all the semantics) are learned through the fusion representation.

[0081] Specifically, the interactive calculation of the first spliced representation sequence with the fusion representation added can be implemented through the layer at the top of the large model:

[0082] Among them, represents the fusion representation, represents the first spliced representation sequence with the fusion representation added, represents the second spliced representation sequence.

[0083] The fusion representation in the second spliced representation sequence has learned the semantics of the subject-text joint query and the semantics of the retrieved video. Therefore, the start and end timestamps of the segment are predicted based on this fusion representation. Specifically, the fusion representation in the second spliced representation sequence can be input into a segment locator (a three-layer fully connected neural network) to output the start and end timestamps of the predicted segment based on the segment locator, and the retrieved video segment corresponding to the subject-text joint query is obtained according to the start and end timestamps of the predicted segment.

[0084] Adopting the technical solution of the embodiments of the present application, the semantics of the subject-text joint query and the semantics of the retrieved video are learned through the fusion representation, and then the start and end timestamps of the segment are accurately predicted based on this fusion representation, and the video segment most relevant to the subject-text joint query is accurately located in the retrieved video.

[0085] Combined with the above embodiments, in one implementation, the embodiments of the present application further provide a method for retrieving video library segments for subject-text joint query. In this method, the method for retrieving video library segments for subject-text joint query is implemented based on a retrieval model, and the retrieval model at least includes a large model; Among them, performing interactive calculation on the first video representation sequence, performing interactive calculation on the first video representation sequence, and calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and taking the video corresponding to the maximum similarity as the retrieved video is implemented using the layer at the bottom of the large model; Predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video is implemented using the layer at the top of the large model, where L represents the total number of layers of the large model.

[0086] As Figure 3 shown, Figure 3 is a schematic structural diagram of a retrieval model provided by the embodiments of the present application. The retrieval model includes a video processing module, a query processing module, and a large model. Among them, the video processing module implements the method of the above step S110. The video processing module includes a visual encoder, a visual adapter, a face representation extractor, and a face representation adapter 1. Specifically, multiple frames of images are uniformly sampled from the video, and the face representation extractor is used to extract face representations from the images. The face representation adapter 1 is used to convert the dimension of the extracted face representations into the feature dimension required by the large model to obtain the final face representations. At the same time, the visual encoder is used to extract other image representations from the images, and the visual adapter is used to convert the dimension of the extracted other image representations into the feature dimension required by the large model to obtain the final other image representations. According to the final face representations and the final other image representations, the video representation corresponding to the image is obtained. Finally, the video representations corresponding to multiple frames of images are spliced according to the frame numbers to obtain the first video representation sequence.

[0087] The query processing module is used to implement the method of the above step S120. The query processing module includes a face representation extractor, a face representation adapter 2, and a text word vector representation layer. Specifically, the face representation extractor is used to extract face representations from the subject query, and the face representation adapter 2 is used to convert the dimension of the extracted face representations into the feature dimension required by the large model to obtain the subject query representation. At the same time, the text word vector representation layer is used to perform word vector conversion on the text query to obtain the text query representation. Finally, according to the subject query representation and the text query representation, the first query representation sequence is obtained.

[0088] The large model is used for video-subject-text multimodal alignment and multimodal fine-grained fusion, that is, the method for implementing the above steps S130 to S150 through the large model. The large model includes L layers and a segment locator. Specifically, through the layers at the bottom of the large model, the first video representation sequence and the first query representation sequence are respectively subjected to interactive calculation, the second video representation sequence with fused context semantics, and the second query representation sequence; and according to the semantic representation in the second query representation sequence and the second video representation sequence, the similarity between the subject-text joint query and each video is calculated, and the video corresponding to the maximum similarity is used as the retrieved video.

[0089] Next, the second query representation sequence and the second video representation sequence corresponding to the retrieved video are spliced to obtain a first spliced representation sequence, a fusion representation is added to the first spliced representation sequence, and the layers at the top of the large model are used to perform interactive calculation on the first spliced representation sequence with the added fusion representation to obtain a second spliced representation sequence. Finally, the fusion representation in the second spliced representation sequence is input into the segment locator to obtain the retrieved video segment corresponding to the subject-text joint query.

[0090] In this way, the video library segment retrieval based on the subject-text joint query is realized through the retrieval model. This retrieval model can not only achieve efficient retrieval through text description, but also combine subject information (subject query) to achieve more accurate video positioning. Moreover, it has good generalization ability and can perform accurate retrieval for unseen subjects.

[0091] Furthermore, the loss function for training the retrieval model includes a first loss, a second loss, a third loss, and a fourth loss; wherein, the first loss represents the contrastive learning loss of the similarity between the subject-text joint query sample and the video sample; the second loss represents the L1 loss between the start and end timestamps of the predicted video segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; the third loss represents the generalized intersection over union loss between the start and end timestamps of the predicted segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; the fourth loss represents the frame classification loss of whether each frame of image is a frame of the video segment to be retrieved.

[0092] In the embodiments of the present application, in order to ensure that the retrieval model can realize the video library segment retrieval based on the subject-text joint query, a loss function including a first loss, a second loss, a third loss, and a fourth loss is used to train the retrieval model.

[0093] To accurately retrieve videos from a video library that match a subject-text joint query, when the subject-text joint query Q and the video V match, the similarity between the two needs to be as large as possible, and vice versa. Therefore, a contrastive learning loss (i.e., the first loss) of the similarity between the subject-text joint query samples and the video samples is used to train the retrieval model. Specifically, given a training batch of subject-text joint query samples and the corresponding video samples , the first loss is expressed as:

[0094] where is the temperature parameter of contrastive learning, and B is the batch size of training.

[0095] The retrieval model is required to accurately locate the video segment corresponding to the query through three losses (i.e., the second loss, the third loss, and the fourth loss). Among them, the second loss and the third loss are segment localization losses, and the fused representation passes through a segment locator (a three-layer fully connected neural network) to output the start and end timestamps of the predicted segment, and calculates the L1 loss (i.e., the second loss ) and the generalized intersection over union loss (the third loss ) with the start and end timestamps of the segment to be retrieved; the fourth loss is the frame classification loss , and respectively pass through a binary classifier to determine whether a certain frame is a frame in the segment to be retrieved.

[0096] Exemplarily, the loss function can be expressed as:

[0097] It should be noted that the video processing module and the query processing module in the retrieval model are pre-trained. During the actual training process, for a given training batch, the total loss value is calculated based on the loss function , and the parameters of the large model are updated based on the total loss value. Thus, when the training end condition is met, a trained retrieval model is obtained, and based on this trained retrieval model, video library segment retrieval based on subject-text joint queries can be realized.

[0098] Based on the same inventive concept, an embodiment of the present application also provides a video library segment retrieval device for subject-text joint queries. Referring to Figure 4 as shown Figure 4 is a schematic structural diagram of a video library segment retrieval device for subject-text joint queries provided by an embodiment of the present application. The device includes: A characterization acquisition module 410, configured to acquire a first video characterization sequence of each video in a video library, where each video characterization in the first video characterization sequence includes subject information and other image information except the subject information; A characterization extraction module 420, configured to perform characterization extraction on a subject query and a text query in a subject-text joint query respectively to obtain a first query characterization sequence; An interaction calculation module 430, configured to perform interaction calculation on the first video characterization sequence to obtain a second video characterization sequence integrating context semantics, and perform interaction calculation on the first query characterization sequence to obtain a second query characterization sequence, where the second query characterization sequence includes a semantic characterization with subject-text joint query semantics; A similarity calculation module 440, configured to calculate the similarity between the subject-text joint query and each video according to the semantic characterization and the second video characterization sequence, and use the video corresponding to the maximum similarity as the retrieved video; A segment prediction module 450, configured to predict the start and end timestamps of a segment according to the second query characterization sequence and the second video characterization sequence corresponding to the retrieved video to obtain a retrieved video segment corresponding to the subject-text joint query.

[0099] In an optional embodiment, the characterization acquisition module includes: An adoption module, configured to uniformly sample a video to obtain multiple frames of images; An information extraction module, configured to extract subject information from each frame of image and extract other image information from each frame of image to obtain a video characterization corresponding to the image; A splicing module, configured to splice the video characterizations corresponding to the multiple frames of images according to the frame numbers to obtain a first video characterization sequence.

[0100] In an optional embodiment, the subject information includes face information, and the other image information includes action information and environment information; The information extraction module is specifically configured to: perform face detection and face characterization extraction on the image, convert the dimension of the extracted face characterization into a feature dimension required by a large model to obtain a final face characterization; perform other image characterization extraction on the image, convert the dimension of the extracted other image characterization into a feature dimension required by a large model to obtain a final other image characterization; and obtain a video characterization corresponding to the image according to the final face characterization and the final other image characterization.

[0101] In an alternative embodiment, the subject query is an image containing the subject. The feature extraction module is specifically configured to: perform face detection and face feature extraction on the subject query, and convert the dimension of the extracted face feature into the feature dimension required by the large model to obtain a subject query feature; perform word vector conversion on the text query to obtain a text query feature; and obtain a first query feature sequence according to the subject query feature and the text query feature.

[0102] In an alternative embodiment, the interaction calculation module includes: A first interaction calculation sub-module, configured to add a semantic feature to the first query feature sequence, where the semantic feature is used to learn the semantics of each query feature in the first query feature sequence; and perform interaction calculation on the first query feature sequence with the added semantic feature to obtain a second query feature sequence.

[0103] In an alternative embodiment, the similarity calculation module is specifically configured to: sequentially calculate the similarity between the semantic feature and each video feature in the second video feature sequence to obtain a plurality of similarity values; and use the maximum value among the plurality of similarity values as the similarity between the subject-text joint query and the video.

[0104] In an alternative embodiment, the segment prediction module is specifically configured to: splice the second query feature sequence and the second video feature sequence corresponding to the retrieved video to obtain a first spliced feature sequence; add a fusion feature to the first spliced feature sequence, and perform interaction calculation on the first spliced feature sequence with the added fusion feature to obtain a second spliced feature sequence, where the fusion feature is used to learn the semantics of each feature in the first spliced feature sequence; and predict the start and end timestamps of the segment according to the fusion feature in the second spliced feature sequence to obtain the retrieved video segment corresponding to the subject-text joint query.

[0105] The embodiment of the present application further provides an electronic device. Refer to Figure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by the embodiment of the present application. As Figure 5 shown, the electronic device 500 includes: a memory 510 and a processor 520. The memory 510 is communicatively connected to the processor 520 through a bus. A computer program is stored in the memory 510, and the computer program can run on the processor 520 to implement the steps of the video library segment retrieval method for subject-text joint query described in the embodiment of the present application.

[0106] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the video library segment retrieval method for entity-text joint query described in the embodiments of the present application are implemented.

[0107] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the video library segment retrieval method for entity-text joint query described in the embodiments of the present application are implemented.

[0108] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0109] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto the computer or other programmable data processing terminal devices, so that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0110] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0111] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0112] The above has introduced in detail a method and apparatus for retrieving video library segments through subject-text joint query provided by this application. Specific examples are used in this document to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A video library clip retrieval method for subject-text joint query, characterized in that: The method comprises: Acquire a first video representation sequence of each video in a video library, wherein each video representation in the first video representation sequence includes subject information and other image information except the subject information; Representation extraction is performed on the subject query and the text query in the subject-text joint query respectively to obtain a first query representation sequence; Interactively calculating the first video representation sequence to obtain a second video representation sequence that incorporates context semantics, and interactively calculating the first query representation sequence to obtain a second query representation sequence, wherein the second query representation sequence includes a semantic representation having subject-text joint query semantics; Calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and taking the video corresponding to the maximum similarity as the retrieval video; The retrieval video segment corresponding to the subject-text joint query is obtained by predicting the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieval video.

2. The method according to claim 1, characterized in that Obtaining the first video representation sequence of each video in the video library, including: Uniformly sample the video to obtain multiple frames of images; Extracting subject information from each frame of image, and extracting other image information from each frame of image, to obtain a video representation corresponding to the image; The video representations corresponding to the multiple frames of images are spliced ​​according to the frame sequence numbers to obtain a first video representation sequence.

3. The method according to claim 2, characterized in that The subject information includes face information, and the other image information includes action information and environment information; Extract the main information of each frame of image, and extract other image information of each frame of image to obtain the video representation corresponding to the image, including: Perform face detection and face representation extraction on the image, and convert the dimension of the extracted face representation into the feature dimension required by the large model to obtain the final face representation; Extracting other image representations from the image, and converting the dimensions of the extracted other image representations into feature dimensions required by the large model to obtain final other image representations; A video representation corresponding to the image is obtained according to the final face representation and the final other image representations.

4. The method according to claim 1, characterized in that: The subject query is an image containing a subject; the subject query and the text query in the subject-text joint query are respectively represented and extracted to obtain a first query representation sequence, including: Performing face detection and face representation extraction on the subject query, and converting the dimension of the extracted face representation into the feature dimension required by the large model to obtain the subject query representation; Convert the text query into a word vector to obtain a text query representation; A first query representation sequence is obtained according to the subject query representation and the text query representation.

5. The method according to claim 1, characterized in that Interactively calculating the first query representation sequence to obtain a second query representation sequence includes: Adding a semantic representation to the first query representation sequence, wherein the semantic representation is used to learn the semantics of each query representation in the first query representation sequence; Interactive calculation is performed on the first query representation sequence with added semantic representation to obtain a second query representation sequence.

6. The method according to claim 1, characterized in that Calculating the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence includes: sequentially calculating the similarity between the semantic representation and each video representation in the second video representation sequence to obtain a plurality of similarity values; The maximum value among the multiple similarity values ​​is used as the similarity between the subject-text joint query and the video.

7. The method according to claim 1, characterized in that The method of obtaining a retrieval video segment corresponding to the subject-text joint query according to the second query representation sequence and the start and end timestamps of the segment predicted by the second video representation sequence corresponding to the retrieval video comprises: splicing the second query representation sequence and the second video representation sequence corresponding to the retrieved video to obtain a first spliced ​​representation sequence; Adding a fused representation to the first concatenated representation sequence, and performing interactive calculation on the first concatenated representation sequence to which the fused representation is added, to obtain a second concatenated representation sequence, wherein the fused representation is used to learn the semantics of each representation in the first concatenated representation sequence; The retrieval video segment corresponding to the subject-text joint query is obtained according to the start and end timestamps of the segment predicted by the fused representation in the second spliced ​​representation sequence.

8. The method according to any one of claims 1 to 7, characterized in that: The video library segment retrieval method for subject-text joint query is implemented based on a retrieval model, and the retrieval model at least includes a large model; The interactive calculation of the first video representation sequence is performed, and the similarity between the subject-text joint query and each video is calculated based on the semantic representation and the second video representation sequence, and the video corresponding to the maximum similarity is used as the retrieval video, which is based on the bottom of the large model. Layer implementation; The start and end timestamps of the segment are predicted based on the second query representation sequence and the second video representation sequence corresponding to the retrieved video, using the top of the large model layers, where L represents the total number of layers in the large model.

9. The method according to claim 8, characterized in that The loss function used to train the retrieval model includes a first loss, a second loss, a third loss, and a fourth loss; Wherein, the first loss represents the contrastive learning loss of the similarity between the subject-text joint query sample and the video sample; The second loss represents the L1 loss between the start and end timestamps of the predicted video segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; The third loss represents the generalized intersection-over-union loss of the start and end timestamps of the predicted segment output by the retrieval model and the start and end timestamps of the video segment to be retrieved; The fourth loss represents the frame classification loss of whether each frame image is the video segment to be retrieved.

10. A video library clip retrieval device for subject-text joint query, characterized in that: The device comprises: A representation acquisition module, used to acquire a first video representation sequence of each video in the video library, wherein each video representation in the first video representation sequence includes subject information and other image information except the subject information; A representation extraction module, used to extract representations of the subject query and the text query in the subject-text joint query respectively, to obtain a first query representation sequence; An interactive computing module, configured to interactively compute the first video representation sequence to obtain a second video representation sequence incorporating context semantics, and interactively compute the first query representation sequence to obtain a second query representation sequence, wherein the second query representation sequence includes a semantic representation having subject-text joint query semantics; A similarity calculation module, used to calculate the similarity between the subject-text joint query and each video according to the semantic representation and the second video representation sequence, and use the video corresponding to the maximum similarity as the retrieval video; The segment prediction module is used to predict the start and end timestamps of the segment according to the second query representation sequence and the second video representation sequence corresponding to the retrieved video, so as to obtain the retrieved video segment corresponding to the subject-text joint query.