Video retrieval feature extraction and retrieval positioning method, electronic equipment, storage medium and program product
Through a two-stage method based on dense video text description, the problem of poor dependence on annotation and interpretation of video retrieval and clip positioning in the prior art is solved, and efficient and resource-saving video retrieval and clip positioning are achieved, which is compatible with downstream tasks.
Patent Information
- Application Number
- CN202510272632.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
Existing video retrieval and clip positioning techniques rely on fine-grained or coarse-grained video-text annotation, which consumes a lot of manpower and material resources, and the end-to-end training methods are poorly interpretable and difficult to compatible with downstream tasks.
A two-stage video retrieval and clip positioning method based on dense video text description is proposed. The video description text collection is generated through a machine learning model, and the description text of low-description quality video frames is updated to the description text of the adjacent high-description quality video frames. The video clips are divided and the search characteristics are extracted. The text similarity of sentences and keywords is searched and positioned.
No clip-level or video-level annotation is required, which saves manpower and material resources, is highly interpretable, and is easy to be compatible with downstream tasks, such as video Q&A and video summary, achieving efficient video retrieval and clip positioning.
Smart Images

Figure CN120196786A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data processing technologies, and more specifically, to a method for video retrieval feature extraction and retrieval positioning, an electronic device, a storage medium, and a program product. Background Art
[0002] Video retrieval and segment positioning is a technology for retrieving relevant videos from a video set according to a query text and further positioning relevant segments. Existing video retrieval and segment positioning technologies generally use fully supervised or semi-supervised training methods to train a video retrieval locator based on segment-level or video-level annotations, and then use the video retrieval locator to perform video retrieval and segment positioning.
[0003] Existing video retrieval and segment positioning technologies rely on fine-grained or coarse-grained video-text annotations, which require a large amount of manpower and material resources. In addition, the end-to-end training method has poor interpretability and is difficult to be further compatible with downstream tasks, such as video question answering, video summarization, etc. Summary of the Invention
[0004] Exemplary embodiments of the present disclosure are directed to providing a method for video retrieval feature extraction and retrieval positioning, an electronic device, a storage medium, and a program product, which can solve at least one of the above problems existing in the prior art.
[0005] According to the first aspect of the embodiments of the present disclosure, a method for extracting video retrieval features is provided, including: respectively taking each video in the video set as a target video, and using a machine learning model to generate a set of description texts for the target video, where the set of description texts includes: description texts of each video frame of the target video; for each video frame of the target video, taking the matching degree score between the image content of the video frame and the description text as the description quality score of the video frame; updating the description text of the low-description-quality video frame in the target video to the description text of the adjacent high-description-quality video frame to obtain an updated set of description texts, where the low-description-quality video frame is a video frame whose description quality score is lower than the first preset threshold, and the high-description-quality video frame is a video frame whose description quality score is higher than the first preset threshold or the description text has been updated; based on the updated set of description texts, dividing the target video into multiple video segments so that video frames with similar semantics in the description text are divided into the same video segment; based on the updated set of description texts, extracting the retrieval features of the target video, where the retrieval features of the target video include: retrieval features of each video segment of the target video, and the retrieval feature of each video segment includes: the representative description text of the video segment and the retrieval keywords of the video segment, the representative description text of each video segment is the description text of the video frame with the highest description quality score in the video segment, and the retrieval keywords of each video segment include: keywords in the description texts of each video frame of the video segment; where the retrieval features of each video in the video set are used to retrieve videos matching a given query text from the video set, and / or, the retrieval features of each video in the video set are used to locate video segments matching the given query text from the video.
[0006] Optionally, the step of updating the description text of the low-description-quality video frame in the target video to the description text of the adjacent high-description-quality video frame includes: sequentially determining whether each video frame of the target video is a low-description-quality video frame according to the display order of the video frames; in response to any video frame being determined to be a low-description-quality video frame, determining whether there is a video frame with a description quality score higher than or equal to the first preset threshold among a predetermined number of video frames that are closest to and before the video frame in the display order; if there is a video frame with a description quality score higher than or equal to the first preset threshold among the predetermined number of video frames, updating the description text of the video frame to: the description text of the video frame that is closest to the video frame among the video frames with a description quality score higher than or equal to the first preset threshold; if there is no video frame with a description quality score higher than or equal to the first preset threshold among the predetermined number of video frames, updating the description text of the video frame to: the updated description text of the video frame that is closest to the video frame among the predetermined number of video frames.
[0007] Optionally, the step of dividing the target video into multiple video segments based on the updated set of description texts includes: constructing a video frame text similarity matrix M of the target video based on the updated set of description texts s , where update the description quality scores of the low-description-quality video frames in the target video to the description quality scores of adjacent high-description-quality video frames, and construct a video frame description quality score matrix M of the target video q , where determine the video frame similarity matrix of the target video where slide a preset convolutional kernel along the diagonal of the video frame similarity matrix of the target video. Whenever the central element of the preset convolutional kernel overlaps with an element on the diagonal of the video frame similarity matrix, take the convolution of the preset convolutional kernel and the video frame similarity matrix as the boundary score of the video frame corresponding to this element; regard the video frames in the target video with boundary scores exceeding the second preset threshold as video segment boundary points; divide the target video into multiple video segments based on the respective video segment boundary points of the target video; where represents the matrix composed of the text encoding results obtained by inputting the description texts of the respective video frames of the target video into the text encoder represents the transpose matrix of represents the matrix composed of the description quality scores of the respective video frames of the target video represents the transpose matrix of, and ⊙ represents matrix multiplication operation
[0008] According to a second aspect of the embodiments of the present disclosure, there is provided a video retrieval method, including: receiving a query text; for each video in the video set, taking the maximum value among the similarities between the representative description texts of the respective video segments of the video and the query text as the sentence-level text similarity between the video and the query text; for each video in the video set, determining the keyword-level text similarity between the video and the query text based on the similarities between the retrieval keywords of the respective video segments of the video and the query keywords in the query text; for each video in the video set, determining the comprehensive text similarity between the video and the query text based on the sentence-level text similarity and the keyword-level text similarity between the video and the query text; determining, from the video set, the videos that match the query text according to the comprehensive text similarities between the respective videos in the video set and the query text; where the retrieval features of the respective videos in the video set are obtained by performing the video retrieval feature extraction method as described above, and the retrieval feature of each video includes: the retrieval features of the respective video segments of the video, and the retrieval feature of each video segment includes: the representative description text of the video segment and the retrieval keyword of the video segment
[0009] Optionally, the step of determining the keyword-level text similarity between the video and the query text based on the similarity between the retrieval keywords of each video segment of each video and the query keywords in the query text includes: for each query keyword in the query text, taking the maximum value among the similarities between each retrieval keyword in the retrieval keyword set of the video and the query keyword as the similarity between the query keyword and the video; taking the average value of the similarities between all query keywords in the query text and the video as the keyword-level text similarity between the video and the query text; wherein, the retrieval keyword set of each video includes: the retrieval keywords of each video segment of the video.
[0010] According to a third aspect of the embodiments of the present disclosure, there is provided a method for video segment localization, including: receiving a query text; taking the similarity between the representative description text of each video segment of the video to be localized in the video set and the query text as the sentence-level text similarity between each video segment and the query text; determining the keyword-level text similarity between each video segment and the query text based on the similarity between the retrieval keywords of each video segment of the video to be localized and the query keywords of the query text; determining the comprehensive text similarity between each video segment and the query text based on the sentence-level text similarity and the keyword-level text similarity between each video segment and the query text; localizing the video segment matching the query text from the video to be localized according to the comprehensive text similarity between each video segment in the video to be localized and the query text; wherein, the retrieval features of the video to be localized are obtained by performing the video retrieval feature extraction method as described above, and the retrieval features of the video to be localized include: the retrieval features of each video segment of the video, and the retrieval features of each video segment include: the representative description text of the video segment and the retrieval keywords of the video segment.
[0011] Optionally, the step of determining the keyword-level text similarity between each video segment and the query text based on the similarity between the retrieval keywords of each video segment of the video to be localized and the query keywords of the query text includes: for each query keyword in the query text, taking the maximum value among the similarities between each retrieval keyword of each video segment in the video to be localized and the query keyword as the similarity between the query keyword and the video segment; taking the average value of the similarities between all query keywords in the query text and the video segment as the keyword-level text similarity between the video segment and the query text.
[0012] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing instructions, which, when executed by a processor of an electronic device, enable the electronic device to execute the video retrieval feature extraction method and / or the video retrieval method and / or the video segment localization method as described above.
[0013] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: at least one processor; and at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the video retrieval feature extraction method and / or the video retrieval method and / or the video clip localization method as described above.
[0014] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer-executable instructions, which when executed by at least one processor, implement the video retrieval feature extraction method and / or the video retrieval method and / or the video clip localization method as described above.
[0015] For the video retrieval feature extraction and retrieval localization method, electronic device, storage medium and program product according to the exemplary embodiments of the present disclosure, in view of the video retrieval and clip localization problems, a two-stage zero-shot video retrieval and clip localization method based on dense video text description is designed. In the construction stage, dense video text descriptions (i.e., retrieval features) at the video clip level are constructed; in the retrieval stage, sentence- and keyword-level multi-granularity text similarity is used to locate relevant videos and clips. On the one hand, it does not require annotation at the clip level or video level, saving manpower and material resources; on the other hand, it has strong interpretability and is convenient for further compatibility with downstream tasks such as video question answering and video summarization.
[0016] In the following description, some aspects and / or advantages of the general concept of the present disclosure will be set forth, and some aspects and / or advantages will be learned from the following description or the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] These and / or other aspects and advantages of the present application will become clearer and easier to understand from the following detailed description of the embodiments of the present application in conjunction with the drawings, wherein:
[0018] Figure 1 A flowchart showing a video retrieval feature extraction method according to an exemplary embodiment of the present disclosure;
[0019] Figure 2 A flowchart showing a method for updating a description text set according to an exemplary embodiment of the present disclosure;
[0020] Figure 3 A flowchart showing a method for dividing a video into multiple video clips according to an exemplary embodiment of the present disclosure;
[0021] Figure 4 A flowchart showing a video retrieval method according to an exemplary embodiment of the present disclosure;
[0022] Figure 5 A flowchart showing a method for locating video segments according to an exemplary embodiment of the present disclosure;
[0023] Figure 6 An example showing a method for extracting video retrieval features and retrieval location according to an exemplary embodiment of the present disclosure;
[0024] Figure 7 A structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0025] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same reference numerals always refer to the same components. The following embodiments will be described with reference to the drawings to explain the present disclosure.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0027] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations including "any one of the several items", "any combination of several items among them", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0028] Figure 1 A flowchart showing a method for extracting video retrieval features according to an exemplary embodiment of the present disclosure.
[0029] Refer to Figure 1 , in step S101, each video in the video set is respectively used as a target video, and a description text set of the target video is generated using a machine learning model.
[0030] The description text set of the target video includes: description texts of each video frame of the target video.
[0031] The video set includes N videos, and the video set can be expressed as where V i represents the i-th video in the video set (i.e., the target video), and i is an integer greater than 0 and less than or equal to N.
[0032] As an exemplary embodiment, a multimodal large model VLLM can be used to generate frame-level text descriptions of the target video, that is, to generate a set of description texts R for the target video i , represents the description text of the l-th video frame of the i-th video, and l is an integer greater than 0 and less than or equal to L i and L i represents the number of video frames of the i-th video. As an example, the multimodal large model VLLM can be specifically LLaVA, MiniGPT, etc., and the present disclosure does not limit this.
[0033] In step S102, for each video frame of the target video, the matching degree score between the image content of the video frame and the description text is used as the description quality score of the video frame.
[0034] As an exemplary embodiment, for the l-th video frame f of the i-th video l i with the description text r l i , the matching degree score between the image content and the description text can be calculated using, for example, the BLIP multimodal model as the description quality score of the video frame
[0035]
[0036] where E v , E t respectively represent the video encoder and the text encoder of the BLIP model, and <·,·> represents the inner product operator.
[0037] In step S103, the description text of the low-description-quality video frames in the target video is updated to the description text of the adjacent high-description-quality video frames, obtaining an updated set of description texts (i.e., a denoised set of description texts).
[0038] Low-description-quality video frames are video frames with a description quality score lower than a first preset threshold, and high-description-quality video frames are video frames with a description quality score higher than the first preset threshold or description texts that have been updated.
[0039] According to the exemplary embodiment of the present disclosure, it is possible to effectively remove low-quality video frame descriptions caused by factors such as multimodal large model hallucinations and low-quality video frames.
[0040] The following will be combined with Figure 2 to describe the exemplary embodiments of step S103, which will not be elaborated here.
[0041] In step S104, based on the updated description text set, the target video is divided into multiple video segments so that video frames with similar semantics in the description text are divided into the same video segment.
[0042] The following will be combined with Figure 3 to describe the exemplary embodiments of step S104, which will not be elaborated here.
[0043] In step S105, based on the updated description text set, the retrieval features of the target video are extracted.
[0044] The retrieval features of the target video include: the retrieval features of each video segment of the target video; the retrieval features of each video segment include: the representative description text of the video segment and the retrieval keywords of the video segment; the representative description text of each video segment is the description text of the video frame with the highest description quality score in the video segment; the retrieval keywords of each video segment include: the keywords in the description texts of each video frame of the video segment.
[0045] The retrieval features of each video in the video set can be used to retrieve videos from the video set that match a given query text. The retrieval features of each video in the video set can be used to locate video segments in the video that match the given query text.
[0046] As an exemplary embodiment, for the h-th video segment of the i-th video of the description texts of each video frame For example, the Spacy natural language processing library can be used to extract all the verbs and nouns in the description texts of each video frame, and the union of all the extracted verbs and nouns is used as the retrieval keyword set of the segment
[0047]
[0048] where NounVerb represents the verb and noun extraction operation, and the video segment can be represented by its start frame end frame representative description text and retrieval keyword set as:
[0049] Figure 2A flowchart showing a method for updating a set of description texts according to an exemplary embodiment of the present disclosure.
[0050] Referring to Figure 2 , in step S201, in the display order of video frames, it is sequentially determined whether each video frame of the target video is a low-description-quality video frame.
[0051] As an exemplary embodiment, when the original description text r of the l-th video frame of the i-th video l i has a description quality score whose numerical value is less than the first preset threshold θ d , it is determined that the original description text r l i is of low description quality and subsequent operations need to be performed to update it to a high-description-quality description text; otherwise, it is determined that the original description text r l i is of high description quality and does not need to be updated.
[0052] In step S202, in response to any video frame (hereinafter, also referred to as the video frame to be updated) being determined to be a low-description-quality video frame, it is determined whether there is a video frame among a predetermined number of video frames that are before the video frame to be updated in the display order and closest to the video frame to be updated, and whose description quality score is higher than or equal to the first preset threshold.
[0053] In step S203, if there is a video frame among the predetermined number of video frames whose description quality score is higher than or equal to the first preset threshold, the description text of the video frame to be updated is updated to: the description text of the video frame closest to the video frame to be updated among the video frames whose description quality score is higher than or equal to the first preset threshold.
[0054] As an exemplary embodiment, based on the fact that adjacent video frames have similar semantics, a sliding window can be used to filter out low-quality frame descriptions and replace them with adjacent high-quality frame descriptions to achieve description text update. This process can be formally described as:
[0055]
[0056]
[0057] where W represents the predetermined number, that is, the description text of the video frame whose description quality score is higher than or equal to θ and closest to the low-description-quality video frame on the left side of the low-description-quality video frame is used as the description text of this frame. d as the description text of this frame.
[0058] In step S204, if there is no video frame in the predetermined number of video frames whose description quality score is higher than or equal to the first preset threshold, update the description text of the video frame to be updated to: the updated description text of the video frame closest to the video frame to be updated among the predetermined number of video frames.
[0059] Specifically, if there is no video frame in the predetermined number of video frames whose description quality score is higher than or equal to the first preset threshold, it means that the original description texts of the predetermined number of video frames are all of low quality and have been updated to high quality. In this case, the updated description text of the video frame closest to the video frame to be updated among the predetermined number of video frames can be used as the updated description text of the video frame to be updated.
[0060] The set of high-quality video frame descriptions obtained after filtering (i.e., the set of updated description texts) can be expressed as:
[0061] Figure 3 A flowchart showing a method of dividing a video into multiple video segments according to an exemplary embodiment of the present disclosure.
[0062] Refer to Figure 3 , in step S301, based on the set of updated description texts, construct a video frame text similarity matrix of the target video.
[0063] As an example, for the i-th video, first construct a video frame text similarity matrix
[0064]
[0065] where E t represents a text encoder such as the BLIP model.
[0066] In step S302, update the description quality score of the video frames with low description quality in the target video to the description quality score of the adjacent high description quality video frames, and construct a video frame description quality score matrix of the target video.
[0067] Specifically, for any video frame, if its original description text is updated to the description text of another video frame, its original description quality score also needs to be updated to the description quality score of the above-mentioned another video frame.
[0068] When performing video segmentation, the text description quality of the video frames still needs to be considered. For this purpose, construct a video frame description quality score matrix of the i-th video and use it as the mask matrix of the video frame text similarity matrix:
[0069]
[0070] Among them, represents the matrix composed of the description quality scores (i.e., the updated description quality scores) of each video frame of the i-th video.
[0071] In step S303, determine the video frame similarity matrix of the target video.
[0072] As an example, for the i-th video, obtain the video frame similarity matrix
[0073]
[0074] where ⊙ represents matrix multiplication operation, for example, Hadamard matrix multiplication operation.
[0075] In step S304, slide the preset convolution kernel along the diagonal of the video frame similarity matrix of the target video. Whenever the central element of the preset convolution kernel overlaps with an element on the diagonal of the video frame similarity matrix, take the convolution of the preset convolution kernel and the video frame similarity matrix as the boundary score of the video frame corresponding to this element.
[0076] As an example, the preset convolution kernel can be the Uboco contrast convolution kernel, such as Figure 6 the convolution kernel shown on the left side in (b), and the central element of this convolution kernel is the element located in the third row and the third column.
[0077] Figure 6 The last figure in (b) corresponds to the video frame similarity matrix of the video. The straight line from the upper left corner to the lower right corner represents the diagonal of the video frame similarity matrix. The brightness of each point in this figure can characterize the boundary score of the video frame corresponding to this point. The higher the boundary score, the higher the brightness. In addition, it should be understood that the convolution of the preset convolution kernel and the video frame similarity matrix is: the convolution of the part of the video frame similarity matrix that overlaps with the current sliding position of the preset convolution kernel and the preset convolution kernel.
[0078] In addition, Figure 6 The middle figure in (b) shows the recognition result of obtaining the video frame boundary score based on the video frame visual features of the video. It can be seen that the recognition effect of the video frame boundary recognition based on the video frame text features proposed by the present disclosure is better.
[0079] In step S305, use the video frames in the target video whose boundary scores exceed the second preset threshold as video segment boundary points.
[0080] In step S306, based on the respective video segment boundary points of the target video, divide the target video into multiple video segments.
[0081] For the i-th video, frames with boundary scores exceeding the second preset threshold θ b are regarded as video segment boundary points, and finally H i video segments are obtained.
[0082] Figure 4 FIG. shows a flowchart of a video retrieval method according to an exemplary embodiment of the present disclosure. The retrieval features of each video in the video set are obtained by executing the video retrieval feature extraction method described in the above exemplary embodiment.
[0083] Referring to Figure 4 , in step S401, a query text is received.
[0084] In step S402, for each video in the video set, the maximum value among the similarities between the representative description texts of the respective video segments of the video and the query text is used as the sentence-level text similarity between the video and the query text.
[0085] For a given query text T, the sentence-level text similarity can be calculated using Sentence Transformer:
[0086]
[0087] where represents the representative description text of the h-th video segment of the i-th video and the similarity with the query text T, and E s represents the text encoder of Sentence Transformer.
[0088] The sentence-level text similarity between the i-th video and the query text is: H i represents the number of video segments of the i-th video.
[0089] In step S403, for each video in the video set, based on the similarities between the retrieval keywords of the respective video segments of the video and the query keywords in the query text, the keyword-level text similarity between the video and the query text is determined.
[0090] As an exemplary embodiment, step S403 may include: for each query keyword in the query text, taking the maximum value among the similarities between each retrieval keyword in the retrieval keyword set of the current video and the query keyword as the similarity between the query keyword and the video; and taking the average value of the similarities between all query keywords in the query text and the video as the keyword-level text similarity between the video and the query text.
[0091] The set of retrieval keywords for each video includes the retrieval keywords of each video segment of the video.
[0092] As an exemplary embodiment, for the text similarity at the keyword level, it can be calculated based on the Glove similarity:
[0093]
[0094] Among them, S w (T, K i ) represents the text similarity at the keyword level between the i-th video and the query text; represents the set of retrieval keywords of video V i ; represents any query keyword in the query text T (for example, verbs or nouns extracted using Spacy); |K T | represents the number of query keywords in the query text T.
[0095] In step S404, for each video in the video set, based on the sentence-level text similarity and keyword-level text similarity between the video and the query text, the comprehensive text similarity between the video and the query text is determined.
[0096]
[0097] Among them, S(T, V i ) represents the comprehensive text similarity between the i-th video and the query text, and α is a balancing factor.
[0098] In step S405, according to the comprehensive text similarity between each video in the video set and the query text, videos that match the query text are determined from the video set.
[0099] As an exemplary embodiment, M videos with the highest comprehensive text similarity to the query text in the video set can be selected as the videos that match the query text, or videos in the video set with a comprehensive text similarity exceeding a predetermined threshold can be selected as the videos that match the query text. M is an integer greater than 0.
[0100] Furthermore, the video retrieval method according to the exemplary embodiment of the present disclosure may further include: using the videos determined from the video set that match the query text as the videos to be located, and locating the video segments that match the query text from the videos to be located. The specific location method can refer to Figure 5 the steps S502 - S505 shown, which will not be elaborated here.
[0101] Figure 5A flowchart showing a method for locating video segments according to an exemplary embodiment of the present disclosure. The retrieval features of the video to be located are obtained by performing the video retrieval feature extraction method as described in the above exemplary embodiment.
[0102] Referring Figure 5 , in step S501, a query text is received.
[0103] In step S502, the similarity between the representative description text of each video segment of the video to be located in the video set and the query text is used as the sentence-level text similarity between each video segment and the query text.
[0104] As an example, the video to be located can be a video determined from the video set that matches the query text, or can be a certain video specified from the video set (for example, a video specified by the user from the video set).
[0105] For a given query text T, the sentence-level text similarity can be calculated using Sentence Transformer:
[0106]
[0107] Where represents the representative description text of the h-th video segment of the i-th video and the similarity with the query text T, E s represents the text encoder of Sentence Transformer.
[0108] In step S503, based on the similarity between the retrieval keywords of each video segment of the video to be located and the query keywords of the query text, the keyword-level text similarity between each video segment and the query text is determined.
[0109] As an exemplary embodiment, step S503 may include: for each query keyword in the query text, taking the maximum value among the similarities between each retrieval keyword of each video segment in the video to be located and the query keyword as the similarity between the query keyword and the video segment; and taking the average value of the similarities between all query keywords in the query text and the video segment as the keyword-level text similarity between the video segment and the query text.
[0110] For the keyword-level text similarity, it can be calculated based on Glove similarity:
[0111]
[0112] Where represents the keyword-level text similarity between the h-th video segment of the i-th video and the query text; Any retrieval keyword for the h-th video segment of the i-th video; Any query keyword in the query text T (e.g., verbs or nouns extracted using Spacy); |K T | represents the number of query keywords in the query text T.
[0113] In step S504, based on the sentence-level text similarity and keyword-level text similarity between each video segment and the query text, determine the comprehensive text similarity between each video segment and the query text.
[0114]
[0115] Wherein, represents the comprehensive text similarity between the h-th video segment of the i-th video and the query text, and α is a balancing factor.
[0116] In step S505, based on the comprehensive text similarity between each video segment in the video to be located and the query text, locate the video segments in the video to be located that match the query text.
[0117] As an exemplary embodiment, Z video segments with the highest comprehensive text similarity between the video to be located and the query text can be selected as the video segments that match the query text, or video segments with a comprehensive text similarity exceeding a certain threshold between the video to be located and the query text can be selected as the video segments that match the query text. Z is an integer greater than 0.
[0118] As Figure 6 shown, according to the exemplary embodiment of the present disclosure, in the construction phase, use a multimodal large model to generate frame-level text descriptions, and then use a sliding window-based filtering method to remove low-quality frame descriptions; then, based on the text similarity and text description quality score matrix of the frame descriptions, construct dense video text descriptions at the segment level. In the retrieval phase, use sentence and keyword multi-granularity text similarity to locate relevant videos and segments.
[0119] The present disclosure proposes a filtering method for low-quality video frame descriptions based on a sliding window to update them to high-quality video frame descriptions; proposes a video chunking method based on frame description similarity to convert continuous video descriptions at the frame level into video descriptions at the segment level; and proposes a retrieval method based on sentence and keyword multi-granularity text similarity to accurately retrieve relevant videos and segments according to the query text.
[0120] The present disclosure uses a multimodal large model to generate explicit video clip descriptions, with the advantages of good explicit and intuitive compatibility. It has achieved accuracy comparable to that of fully supervised and semi-supervised methods on multiple datasets and significantly outperformed existing methods on the ActivityNet dataset.
[0121] Figure 7 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.
[0122] Referring to Figure 7 , the electronic device includes: at least one memory 600 and at least one processor 700. A set of computer-executable instructions is stored in the at least one memory 600. When the set of computer-executable instructions is executed by the at least one processor 700, at least one of the following items is performed: the video retrieval feature extraction method as described in the above exemplary embodiment, the video retrieval method as described in the above exemplary embodiment, and the video clip localization method as described in the above exemplary embodiment.
[0123] As an example, the electronic device may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above set of instructions. Here, the electronic device does not have to be a single electronic device, and may also be any assembly of devices or circuits capable of executing the above instructions (or instruction set) individually or jointly. The electronic device may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected with a local or remote (e.g., via wireless transmission) interface.
[0124] In the electronic device, the processor 700 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processor 700 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0125] The processor 700 may run the instructions or code stored in the memory 600, where the memory 600 may also store data. The instructions and data may also be sent and received via a network interface device through a network, where the network interface device may adopt any known transmission protocol.
[0126] The memory 600 may be integrated with the processor 700. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. Additionally, the memory 600 may include separate devices such as external disk drives, storage arrays, or other storage devices that may be used by any database system. The memory 600 and the processor 700 may be operatively coupled or may communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 700 is able to read files stored in the memory.
[0127] In addition, the electronic device may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device may be connected to each other via a bus and / or a network.
[0128] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are run by at least one processor, the at least one processor is caused to perform at least one of the following: the video retrieval feature extraction method as described in the above exemplary embodiment, the video retrieval method as described in the above exemplary embodiment, and the video segment localization method as described in the above exemplary embodiment. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), card memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0129] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, wherein the instructions in the computer program product may be executed by at least one processor to complete at least one of the following: the video retrieval feature extraction method as described in the above exemplary embodiment, the video retrieval method as described in the above exemplary embodiment, and the video segment localization method as described in the above exemplary embodiment.
[0130] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0131] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video retrieval feature extraction method, characterized in that: include: Taking each video in the video set as a target video, respectively, and using a machine learning model to generate a description text set of the target video, wherein the description text set includes: description text of each video frame of the target video; For each video frame of the target video, the matching score between the image content of the video frame and the description text is used as the description quality score of the video frame; Update the description text of the low description quality video frame in the target video to the description text of the adjacent high description quality video frame to obtain an updated description text set, wherein the low description quality video frame is a video frame whose description quality score is lower than a first preset threshold, and the high description quality video frame is a video frame whose description quality score is higher than the first preset threshold or whose description text has been updated; Based on the updated description text set, the target video is divided into a plurality of video segments, so that video frames with description texts having similar semantics are divided into the same video segment; Based on the updated description text set, extract the retrieval features of the target video, wherein the retrieval features of the target video include: retrieval features of each video segment of the target video, the retrieval features of each video segment include: representative description text of the video segment and retrieval keywords of the video segment, the representative description text of each video segment is the description text of the video frame with the highest description quality score in the video segment, and the retrieval keywords of each video segment include: keywords in the description text of each video frame of the video segment; The retrieval features of each video in the video set are used to retrieve videos matching a given query text from the video set, and / or the retrieval features of each video in the video set are used to locate a video segment matching a given query text from the video.
2. The video retrieval feature extraction method according to claim 1, characterized in that: The step of updating the description text of the low description quality video frame in the target video to the description text of the adjacent high description quality video frame includes: According to the display order of the video frames, determining in sequence whether each video frame of the target video is a low description quality video frame; In response to any video frame being determined as a low description quality video frame, determining whether there is a video frame with a description quality score higher than or equal to a first preset threshold among a predetermined number of video frames that are before the video frame and are closest to the video frame in a display order; If there is a video frame whose description quality score is higher than or equal to the first preset threshold value among the predetermined number of video frames, updating the description text of the video frame to: the description text of the video frame that is closest to the video frame among the video frames whose description quality scores are higher than or equal to the first preset threshold value; If there is no video frame with a description quality score higher than or equal to the first preset threshold in the predetermined number of video frames, the description text of the video frame is updated to: the updated description text of the video frame closest to the video frame in the predetermined number of video frames.
3. The video retrieval feature extraction method according to claim 1, characterized in that: Based on the updated description text set, the step of dividing the target video into multiple video segments includes: Based on the updated description text set, construct the video frame text similarity matrix M of the target video s ,in, Update the description quality score of the low description quality video frame in the target video to the description quality score of the adjacent high description quality video frame, and construct the video frame description quality score matrix M of the target video q ,in, Determine the video frame similarity matrix of the target video in, The preset convolution kernel is made to slide along the diagonal of the video frame similarity matrix of the target video. Whenever the central element of the preset convolution kernel overlaps with an element on the diagonal of the video frame similarity matrix, the convolution of the preset convolution kernel and the video frame similarity matrix is used as the boundary score of the video frame corresponding to the element; Taking a video frame in the target video whose boundary score exceeds a second preset threshold as a video segment boundary point; Based on the boundary points of each video segment of the target video, the target video is divided into a plurality of video segments; in, represents the matrix formed by inputting the description text of each video frame of the target video into the text encoder to obtain the text encoding result. express The transposed matrix of Represents the matrix composed of the description quality scores of each video frame of the target video, express is the transposed matrix of , and ⊙ represents the matrix product operation.
4. A video retrieval method, characterized in that: include: Receive query text; For each video in the video set, the maximum value of the similarities between the representative description texts of each video segment of the video and the query text is used as the sentence-level text similarity between the video and the query text; For each video in the video set, based on the similarity between the search keywords of each video segment of the video and the query keywords in the query text, determine the keyword-level text similarity between the video and the query text; For each video in the video set, determine the comprehensive text similarity between the video and the query text based on the sentence-level text similarity and keyword-level text similarity between the video and the query text; Determine a video matching the query text from the video set based on the comprehensive text similarity between each video in the video set and the query text; Among them, the retrieval features of each video in the video set are obtained by executing the video retrieval feature extraction method as described in any one of claims 1 to 3, and the retrieval features of each video include: the retrieval features of each video segment of the video, and the retrieval features of each video segment include: the representative description text of the video segment and the retrieval keywords of the video segment.
5. The video retrieval method according to claim 4, characterized in that: Based on the similarity between the search keywords of each video segment of each video and the query keywords in the query text, the step of determining the keyword-level text similarity between the video and the query text comprises: For each query keyword in the query text, the maximum value among the similarities between each search keyword in the search keyword set of the video and the query keyword is used as the similarity between the query keyword and the video; The average of the similarities between all query keywords in the query text and the video is taken as the keyword-level text similarity between the video and the query text; The search keyword set for each video includes: search keywords for each video segment of the video.
6. A video segment positioning method, characterized in that: include: Receive query text; The similarity between the representative description text of each video segment of the to-be-located video in the video set and the query text is used as the sentence-level text similarity between each video segment and the query text; Determine the keyword-level text similarity between each video segment and the query text based on the similarity between the search keyword of each video segment of the video to be located and the query keyword of the query text; Determine the comprehensive text similarity between each video clip and the query text based on the sentence-level text similarity and keyword-level text similarity between each video clip and the query text; Locating video clips matching the query text from the video to be located according to the comprehensive text similarity between each video clip in the video to be located and the query text; Among them, the retrieval features of the video to be located are obtained by executing the video retrieval feature extraction method as described in any one of claims 1 to 3, and the retrieval features of the video to be located include: the retrieval features of each video segment of the video, and the retrieval features of each video segment include: the representative description text of the video segment and the retrieval keywords of the video segment.
7. The video segment positioning method according to claim 6, characterized in that: Based on the similarity between the search keywords of each video segment of the video to be located and the query keywords of the query text, the step of determining the keyword-level text similarity between each video segment and the query text comprises: For each query keyword in the query text, the maximum value of the similarities between each search keyword of each video segment in the video to be located and the query keyword is used as the similarity between the query keyword and the video segment; The average value of the similarities between all query keywords in the query text and the video clip is taken as the keyword-level text similarity between the video clip and the query text.
8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the video retrieval feature extraction method according to any one of claims 1 to 3 and / or the method according to any one of claims 4 to 7.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the video retrieval feature extraction method as described in any one of claims 1 to 3 and / or the method as described in any one of claims 4 to 7.
10. A computer program product comprising computer executable instructions, characterized in that: When the computer executable instructions are executed by at least one processor, the video retrieval feature extraction method according to any one of claims 1 to 3 and / or the method according to any one of claims 4 to 7 are implemented.