Video retrieval method and device, product, electronic equipment and storage medium

By dividing the video into multiple clips and calculating the similarity index value using a multimodal model, the problem of being unable to accurately locate video clips in the prior art is solved, and the accuracy and effectiveness of video retrieval is achieved.

CN119938982APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411833774.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the video clips corresponding to the search text cannot be accurately positioned in video retrieval, and only relevant videos can be output but specific clips in the video cannot be accurately positioned.

Method used

By acquiring multiple initial videos and dividing them into multiple video clips, the video feature data of each video clip is extracted, and the text feature data of the search text is obtained in response to a video search request. Using the pre-trained target multimodal model, the similarity index values ​​of video clips and search text are calculated, and then precisely positioned search results are generated.

Benefits of technology

The precise positioning of video clips is achieved, making the search results more in line with expectations, and can understand the text and accurately locate the relevant video clips.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938982A_ABST
    Figure CN119938982A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video retrieval method and device, a product, electronic equipment and a storage medium, and the method comprises the steps: obtaining a plurality of initial videos, dividing each initial video into a plurality of video clips, obtaining and storing video feature data of each video clip, responding to a video retrieval request, obtaining a retrieval text corresponding to the video retrieval request, and storing the retrieval text corresponding to the video retrieval request; obtaining text feature data of the retrieval text, determining a first similarity index value of the text feature data and each piece of video feature data, determining a plurality of target videos in the plurality of initial videos according to the first similarity index value, and determining a first target video clip corresponding to each target video, and inputting the first target video clip and the retrieval text into a pre-trained target multi-modal model to obtain a second similarity index value of the first target video clip and the retrieval text, generating a retrieval result according to the first target video clip and the second similarity index value, and returning the retrieval result, thereby realizing accurate positioning of the text retrieval video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and device, a product, an electronic device, and a storage medium for video retrieval. Background Art

[0002] With the development of artificial intelligence, big data and other data, information retrieval has become more and more convenient and faster. Among them, text-video retrieval aims to find the most relevant videos for a given retrieval text query. With the continuous increase in the number of online videos, this demand has gradually emerged.

[0003] In the related art, the feature vector of the search text is encoded, the cosine similarity is calculated, and the video search result of the search text is determined according to the cosine similarity. However, this solution has certain limitations, as it can only output the video and cannot accurately locate the video clip corresponding to the search text. Summary of the invention

[0004] In view of the above problems, a method and apparatus, product, electronic device, and storage medium for video retrieval are proposed to overcome the above problems or at least partially solve the above problems, including:

[0005] A video retrieval method, the method comprising:

[0006] Acquire multiple initial videos, and divide each of the initial videos into multiple video segments;

[0007] Obtaining and saving video feature data of each of the video clips;

[0008] In response to a video search request, obtaining a search text corresponding to the video search request;

[0009] Acquire text feature data of the search text, and determine a first similarity index value between the text feature data and each of the video feature data;

[0010] According to the first similarity index value, determining multiple target videos from the multiple initial videos, and determining a first target video segment corresponding to each target video;

[0011] Inputting the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text;

[0012] According to the first target video segment and the second similarity index value, a retrieval result is generated and returned.

[0013] Optionally, generating and returning a search result according to the first target video segment and the second similarity index value includes:

[0014] Selecting, from the first target video segments, a plurality of second target video segments whose second similarity index value is greater than or equal to a preset threshold;

[0015] determining a minimum video start time and a maximum video end time among the plurality of second target video segments;

[0016] According to the minimum video start time and the maximum video end time, intercepting the target video corresponding to the first target video segment, obtaining a third target video segment;

[0017] A retrieval result is generated according to the third target video segment and returned.

[0018] Optionally, generating and returning a search result according to the intercepted target video includes:

[0019] Determine an average value of a plurality of second similarity index values ​​corresponding to each of the third target video segments to obtain a target similarity value;

[0020] The third target video is sorted in descending order according to the target similarity value to obtain the search result, and the search result is returned.

[0021] Optionally, before acquiring the multiple initial videos, the method further includes:

[0022] Constructing an initial multimodal model;

[0023] Acquire video sample data and text sample data; wherein the video sample data is an image frame of a video, and the text sample data is a description text of the image frame;

[0024] Training the initial multimodal model according to the video sample data and the text sample data;

[0025] Adding a neural network layer to the trained initial multimodal model to obtain a target multimodal model; wherein the neural network layer is used to extract fusion features of the video sample data and the text sample data;

[0026] The target multimodal model is trained according to the video sample data and the text sample data to obtain a pre-trained target multimodal model.

[0027] Optionally, the method further comprises:

[0028] In the process of training the initial multimodal model, determining the timing information of the video sample data;

[0029] Feature extraction is performed on the video sample data through the self-attention mechanism of the initial multimodal model and the timing information.

[0030] Optionally, the neural network layer consists of alternating layers of local multi-head self-attention modules and multi-layer perceptrons.

[0031] A video retrieval device, comprising:

[0032] A video segment processing module, used for acquiring a plurality of initial videos and dividing each of the initial videos into a plurality of video segments;

[0033] A video feature data acquisition module, used to acquire and save the video feature data of each video clip;

[0034] A search text acquisition module, used to respond to a video search request and acquire a search text corresponding to the video search request;

[0035] A first similarity index value determination module, used for obtaining text feature data of the search text, and determining a first similarity index value between the text feature data and each of the video feature data;

[0036] a first target video segment determining module, configured to determine a plurality of target videos from the plurality of initial videos according to the first similarity index value, and determine a first target video segment corresponding to each target video;

[0037] A second similarity index value determination module, configured to input the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text;

[0038] The retrieval result generating module is used to generate and return a retrieval result according to the first target video segment and the second similarity index value.

[0039] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the computer program implements the video retrieval method as described above.

[0040] An electronic device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the above-mentioned video retrieval method when executed by the processor.

[0041] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the video retrieval method described above is implemented.

[0042] The embodiment of the present invention has the following advantages: by obtaining multiple initial videos, and dividing each initial video into multiple video segments, obtaining and saving video feature data of each video segment, in response to a video retrieval request, obtaining a retrieval text corresponding to the video retrieval request, obtaining text feature data of the retrieval text, and determining a first similarity index value between the text feature data and each video feature data; according to the first similarity index value, determining multiple target videos in the multiple initial videos, and determining a first target video segment corresponding to each target video; inputting the first target video segment and the retrieval text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the retrieval text; generating and returning a retrieval result according to the first target video segment and the second similarity index value; firstly roughly locating the video segment to be retrieved by using the first similarity index value, and then accurately locating the video segment in combination with the second similarity index value between the first target video segment and the retrieval text; the text can be accurately understood and the relevant video segment can be accurately located, so that the retrieval result is more in line with expectations. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0044] Figure 1 is a flowchart of the steps of a video retrieval method provided by one embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of the architecture of a multimodal model provided by an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of the system architecture of a video retrieval system provided by an embodiment of the present invention;

[0047] Figure 4 It is a structural block diagram of a video retrieval device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] Video-text retrieval is a fundamental research task in the field of multimodal video and language understanding, which aims to find the most relevant videos for a given text query. This demand has gradually emerged with the increasing number of online videos.

[0050] Based on the development of language models, large multi-modal models (LMM) have made significant progress in video understanding, such as GPT-3 and LLaMA (Large Language Model Meta AI), demonstrating strong general capabilities and laying the foundation for the realization of general artificial intelligence.

[0051] However, although these related technologies use advanced large language models, they still rely on image or video encoders to process visual input, and each encoder has its own limitations. Image encoders are good at capturing rich spatial details from frame sequences, but lack clear temporal context, which is critical in videos with complex action sequences. On the other hand, video encoders can provide temporal context, but are often limited by computational conditions, resulting in only sparse frames at lower resolutions, resulting in a lack of context and spatial understanding.

[0052] Currently, there is a rapidly growing demand for developing multimodal conversational models that can adapt to various input modalities including images and videos. The latest achievements in multimodal conversational models, such as MiniGPT-4 and LLaVA (Large Language and Vision Assistant), are committed to integrating vision into LLM. Despite the impressive results, existing methods are usually designed for processing image or video input. For example, methods that prioritize image input often use a large number of visual tags to obtain a more refined spatial understanding. Although some methods, such as Flamingo, can use query transformers to extract a fixed number of tokens for each image and video, their main focus is still on image understanding, and they lack the ability to effectively build temporal understanding, which leads to limitations in understanding videos. Therefore, achieving understanding of images and videos in a unified framework is both critical and challenging for LLM. Chat-UniVi is a large multimodal model for unified image and video understanding proposed by researchers from institutions such as Peking University and Sun Yat-sen University. However, in video retrieval, video action localization is also important, which determines whether the returned retrieval results are accurate.

[0053] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0054] Reference Figure 1 , shows a flowchart of a video retrieval method provided by an embodiment of the present invention, which may specifically include the following steps:

[0055] Step 101, obtaining a plurality of initial videos, and dividing each of the initial videos into a plurality of video segments;

[0056] Exemplarily, if there are M initial videos, each initial video may be split into N video segments.

[0057] Step 102, obtaining and saving video feature data of each of the video clips;

[0058] Exemplarily, the video feature data of each video clip can be extracted through a pre-trained model and stored in a database.

[0059] Step 103, in response to the video search request, obtaining a search text corresponding to the video search request;

[0060] The search text may be a phrase, a combination of words, etc., which is used to describe multiple features of the video clip that is desired to be searched, such as "three puppies playing on the beach".

[0061] Step 104, obtaining text feature data of the search text, and determining a first similarity index value between the text feature data and each of the video feature data;

[0062] In a specific implementation, the text feature data of the retrieved text may be extracted through a pre-trained model, and then a first similarity index value of the text feature data and the video feature data may be calculated through a preset algorithm (eg, cosine similarity or other algorithms).

[0063] For example, let the initial video be V, and the video feature data of its N video clips are F1, F2, F3, ..., F N , the text feature data is T, then T is respectively combined with F1, F2, F3...F N The first similarity index value may be a specific similarity value, for example, it may be expressed as a percentage.

[0064] Step 105, determining a plurality of target videos from the plurality of initial videos according to the first similarity index value, and determining a first target video segment corresponding to each target video;

[0065] In a specific implementation, the target video may be determined according to a specific value of the first similarity index value in combination with a threshold.

[0066] For example, there are three initial videos V1, V2, and V3, and each initial video corresponds to three video clips; the similarity between each video clip of V1 and the text feature data is 88%, 86%, and 98%, respectively; the similarity between each video clip of V2 and the text feature data is 90%, 70%, and 81%, respectively; and the similarity between each video clip of V3 and the text feature data is 64%, 49%, and 51%, respectively. If the threshold is set to 80%, then 3 video clips of V1, 2 videos of V2, and 0 video clips of V3 meet the requirements, and V1 and V2 are determined to be target videos, and the matching degree between V1 and the retrieved text is greater than that of V2; the video clips corresponding to V1 and V2 are the first target video clips. This step can also be called coarse positioning of the video clips to be retrieved.

[0067] Step 106, inputting the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text;

[0068] In this embodiment, the target multimodal model has the multi-task capabilities of video understanding and video positioning, which can identify the matching degree between the retrieved text and the video, as well as the corresponding position of the content of the retrieved text in the video (such as the specific start time and end time).

[0069] By inputting the first target video clip and the retrieval text into the pre-trained target multimodal model, a second similarity index value of each first target video clip and the retrieval text is obtained. The second similarity index value can specifically be a similarity value, a score value based on the similarity, etc.

[0070] In practical applications, the results of the target multimodal model can be Figure 2 As shown, the model supports multi-modal input of video, picture, text and voice. In step 106, Figure 2 The portion a framed by the dotted line obtains the second similarity index value between the first target video segment and the search text.

[0071] In some embodiments of the present invention, before acquiring the multiple initial videos, the method further includes:

[0072] Step 1001, constructing an initial multimodal model;

[0073] Exemplarily, an initial multimodal model can be constructed based on the deep learning model architecture of Transformer.

[0074] Step 1002, obtaining video sample data and text sample data; wherein the video sample data is an image frame of a video, and the text sample data is a description text of the image frame;

[0075] In this embodiment, video understanding training is first performed on the initial multimodal model based on video sample data and text sample data.

[0076] In some examples, a feature pyramid (a network structure used to handle multi-scale object detection) can be constructed for the initial multimodal model, and then the initial multimodal model can be trained.

[0077] Specifically, the model structure of the picture tower can be used as a video tower and a text tower for describing each picture (image frame of the video). The picture tower and the text tower have been aligned, which can train the model to accurately understand the content of the picture based on the input text. In addition, training videos based on the picture tower can obtain fine-grained features of the video frames, and complete video understanding training better and faster.

[0078] In practical applications, the LLM model can also be used to describe video frames to achieve arbitrary action understanding. Specifically, the description sentence is input into the LLM for expansion to obtain a complete description of the video sample, which is used to construct a training data set so that the multimodal model can understand actions in the open domain.

[0079] Step 1003: training the initial multimodal model according to the video sample data and the text sample data;

[0080] Continuing with the above example, since the video tower and the text tower have been aligned, during the video understanding training stage, only the video tower is trained, and the text tower is not trained.

[0081] In practical applications, the loss function for training the initial multimodal model can be obtained by combining the i2t (Image-to-Text, from image to text alignment) loss function and the t2i (Text-to-Image, from text to image alignment) loss function in the CLIP (Contrastive Language–Image Pre-training) model.

[0082] In some embodiments of the present invention, the method further comprises:

[0083] In the process of training the initial multimodal model, determining the timing information of the video sample data;

[0084] Feature extraction is performed on the video sample data through the self-attention mechanism of the initial multimodal model and the timing information.

[0085] Timing information refers to the information of video sample data in time series. For example, for an image frame at a certain time t in the video stream, the image frame at time t-1 and the image frame at time t+1 are extracted respectively, and the features of the image frames at times t-1, t, and t+1 are extracted through the self-attention mechanism.

[0086] Specifically, based on the formula of the self-attention mechanism:

[0087]

[0088] Among them, softmax is the activation function, q is the query matrix, K(Key) is the key, the characteristic information of the input data, V is the actual information of the input data, and d is the dimension of the key.

[0089] Combined with the temporal information of the video sample data, the self-attention mechanism is:

[0090]

[0091] Among them, K (t-1) That is, the key of the video sample data at time t-1, V (t-1) is the actual value of the video sample data at time t-1, and so on.

[0092] By adding time series information to the self-attention mechanism for feature extraction, the model can obtain better time information aggregation capabilities. In addition, since no additional parameters are added, the inference time does not increase.

[0093] Step 1004, adding a neural network layer to the trained initial multimodal model to obtain a target multimodal model; wherein the neural network layer is used to extract fusion features of the video sample data and the text sample data;

[0094] In this embodiment, after the initial multimodal model is trained for video understanding, the video positioning training phase is entered. By adding a neural network layer, for example, adding L Tansformer layers, the fusion features of the video sample data and the text sample data are extracted.

[0095] In some embodiments of the present invention, the neural network layer consists of alternating layers of local multi-head self-attention modules and multi-layer perceptrons.

[0096] In the Transformer model, the Multi-Head Self-Attention module (MSA) and the Multi-Layer Perceptron (MLP) are two core components, which are stacked alternately in the Transformer layers and together constitute the Transformer's encoder and decoder.

[0097] In practical applications, the detection head used by the decoder can be connected to the features of each layer of the feature pyramid (such as Figure 2 The h1, h2, and h3 features in the video-text fusion feature (they share parameters), that is, the detection head is attached to the features of each layer. Specifically, the feature pyramid can be the feature pyramid in VitDet (an object detection model based on the Vision Transformer (ViT) architecture).

[0098] Among them, the classification prediction head and regression prediction head in the detection head can be implemented using a one-dimensional convolutional network. The two have similar architectures. The main difference is that the final activation function used in the classification prediction head is the sofxmax function, and the final activation function used in the regression prediction head is the ReLU (Rectified Linear Unit) function.

[0099] Step 1005: train the target multimodal model according to the video sample data and the text sample data to obtain a pre-trained target multimodal model.

[0100] In this embodiment, during the training phase of video positioning, the video tower and text tower trained in the video understanding training are used, but they are not trained. Only the newly added network structure of the video tower (i.e., the newly added neural network layer) is trained, and the same sample data as that in the video understanding training phase is used for training.

[0101] Specifically, the loss function used during training may be a cross entropy loss function, which is used to calculate category loss, and an IoU loss function (Intersection over Union Loss), which is used to calculate position loss.

[0102] The cross entropy loss function is as follows:

[0103] L=-y*log(p)-(1-y)*log(1-p)

[0104] In practical applications, in order to realize open domain video positioning in multimodal models, when organizing data, LLM can be used to describe the image frames in the video sample data in text, and the multimodal model can be used to extract text features from the text. The text features are then aggregated to obtain video positioning information as soft labels for the video sample data, and the soft labels are then used to train the model to achieve the purpose of open domain video positioning.

[0105] In practical applications, in step 102, the video feature data of the video clip can be extracted based on the trained target multimodal model, and in step 104, the text feature data of the search text can be extracted, that is, mainly as follows Figure 2 The dotted-line framed portion a in the multimodal model shown extracts video feature data and text feature data.

[0106] Step 107: Generate and return a search result based on the first target video segment and the second similarity index value.

[0107] In practical applications, the first target video segments may be sorted in descending order according to the second similarity index values ​​as the search results and returned, or the first target video segments may be further processed before being returned.

[0108] In some embodiments of the present invention, generating and returning a search result according to the first target video segment and the second similarity index value includes:

[0109] Selecting, from the first target video segments, a plurality of second target video segments whose second similarity index value is greater than or equal to a preset threshold;

[0110] Exemplarily, for a target video V, it is split into 5 video segments, namely the first target video segments v1, v2, v3, v4, and v5, and the second similarity index values ​​are s1, s2, s3, s4, and s5, respectively. Those greater than or equal to the preset threshold smax are s1, s2, and s4, respectively. Then v1, v2, and v4 are determined as the second target video segments.

[0111] determining a minimum video start time and a maximum video end time among the plurality of second target video segments;

[0112] Continuing with the above example, let v1 be the video segment of video V[00:00-05:00], v2 be the video segment of video V[05:00-10:00], and v4 be the video segment of [15:00-20:00]. Then the minimum video start time is the video start time of v1, 00:00, and the maximum video end time is the end time of the v4 video segment, 20:00.

[0113] According to the minimum video start time and the maximum video end time, intercepting the target video corresponding to the first target video segment, obtaining a third target video segment;

[0114] A retrieval result is generated according to the third target video segment and returned.

[0115] Continuing with the above example, the minimum video start time is used as the start time T-start of the video V, and the maximum video end time is used as the end time T-end of the video V. The video V is edited using T-start and T-end to obtain the third target video segment vt.

[0116] Similarly, the above operation is performed for other target videos.

[0117] In some embodiments of the present invention, generating and returning a search result based on the captured target video includes:

[0118] Determine an average value of a plurality of second similarity index values ​​corresponding to each of the third target video segments to obtain a target similarity value;

[0119] The third target video is sorted in descending order according to the target similarity value to obtain the search result, and the search result is returned.

[0120] For example, following the above example, vt corresponds to three second similarity index values ​​s1, s2, and s4, then s1, s2, and s4 are averaged, and the average value is used as the target similarity value between vt and the search text. Similarly, the same operation is performed on other third target video clips.

[0121] Finally, the third target video clips are sorted in descending order according to the target similarity value and returned as the search results. The higher the ranking, the more the third target video clip matches the search text.

[0122] In some embodiments of the present invention, a video retrieval system is also provided. Figure 3 As shown, the system includes online and offline modules.

[0123] The offline module is mainly used to: cut the video to obtain video clips; obtain video clip features through the multimodal large model, and store the features to obtain a video clip feature library;

[0124] The online module (online) is mainly used for: obtaining prompt word features for the text prompt words input by the user through a multimodal large model; matching the prompt word features from the video clip feature library to obtain the top N videos (TopN videos) with similar retrieval; inputting the retrieved videos and text prompt words into the multimodal large model to obtain M video clips (TopM videos) after the video is accurately located.

[0125] The embodiment of the present invention has the following advantages: by obtaining multiple initial videos, and dividing each initial video into multiple video segments, obtaining and saving video feature data of each video segment, in response to a video retrieval request, obtaining a retrieval text corresponding to the video retrieval request, obtaining text feature data of the retrieval text, and determining a first similarity index value between the text feature data and each video feature data; according to the first similarity index value, determining multiple target videos in the multiple initial videos, and determining a first target video segment corresponding to each target video; inputting the first target video segment and the retrieval text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the retrieval text; generating and returning a retrieval result according to the first target video segment and the second similarity index value; firstly roughly locating the video segment to be retrieved by using the first similarity index value, and then accurately locating the video segment in combination with the second similarity index value between the first target video segment and the retrieval text; the text can be accurately understood and the relevant video segment can be accurately located, so that the retrieval result is more in line with expectations.

[0126] In addition, in some embodiments of the present invention, a multimodal model is proposed, which can support video, picture, text and voice input, and multimodality is aligned through text. This model integrates video understanding and positioning into the multimodal model for the first time. It can not only extract video multimodal features, but also understand and automatically edit the input video clips, and output video clips related to the text.

[0127] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0128] Reference Figure 4 , shows a schematic diagram of the structure of a video retrieval device provided by an embodiment of the present invention, which may specifically include the following modules:

[0129] The video segment processing module 401 is used to obtain multiple initial videos and divide each of the initial videos into multiple video segments;

[0130] A video feature data acquisition module 402, used to acquire and save the video feature data of each video clip;

[0131] A search text acquisition module 403 is used to respond to a video search request and acquire a search text corresponding to the video search request;

[0132] A first similarity index value determination module 404 is used to obtain text feature data of the search text and determine a first similarity index value between the text feature data and each of the video feature data;

[0133] A first target video segment determination module 405, configured to determine a plurality of target videos from the plurality of initial videos according to the first similarity index value, and determine a first target video segment corresponding to each target video;

[0134] A second similarity index value determination module 406 is used to input the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text;

[0135] The search result generation module 407 is used to generate and return a search result according to the first target video segment and the second similarity index value.

[0136] In some embodiments of the present invention, the search result generating module 407 includes:

[0137] A second target video segment screening submodule, configured to screen, from the first target video segments, a plurality of second target video segments whose second similarity index value is greater than or equal to a preset threshold;

[0138] A start and end time determination submodule, configured to determine a minimum video start time and a maximum video end time among the plurality of second target video segments;

[0139] A third target video segment generating module, configured to intercept a target video corresponding to the first target video segment according to the minimum video start time and the maximum video end time, to obtain a third target video segment;

[0140] The retrieval result generating submodule is used to generate and return the retrieval result according to the third target video segment.

[0141] In some embodiments of the present invention, the search result generation submodule is further used to:

[0142] Determine an average value of a plurality of second similarity index values ​​corresponding to each of the third target video segments to obtain a target similarity value;

[0143] The third target video is sorted in descending order according to the target similarity value to obtain the search result, and the search result is returned.

[0144] In some embodiments of the present invention, before acquiring the multiple initial videos, the apparatus further includes:

[0145] An initial multimodal model building module, used to build an initial multimodal model;

[0146] A sample data acquisition module, used to acquire video sample data and text sample data; wherein the video sample data is an image frame of a video, and the text sample data is a description text of the image frame;

[0147] A first training module, used for training the initial multimodal model according to the video sample data and the text sample data;

[0148] A neural network layer adding module, used for adding a neural network layer to the trained initial multimodal model to obtain a target multimodal model; wherein the neural network layer is used to extract fusion features of the video sample data and the text sample data;

[0149] The second training module is used to train the target multimodal model according to the video sample data and the text sample data to obtain a pre-trained target multimodal model.

[0150] In some embodiments of the present invention, the device further comprises:

[0151] A timing information determination module, used to determine the timing information of the video sample data during the training of the initial multimodal model;

[0152] A self-attention mechanism feature extraction module is used to extract features from the video sample data through the self-attention mechanism of the initial multimodal model and the timing information.

[0153] In some embodiments of the present invention, the neural network layer consists of alternating layers of local multi-head self-attention modules and multi-layer perceptrons.

[0154] Some embodiments of the present invention further provide an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the following steps when executed by the processor:

[0155] Acquire multiple initial videos, and divide each of the initial videos into multiple video segments;

[0156] Obtaining and saving video feature data of each of the video clips;

[0157] In response to a video search request, obtaining a search text corresponding to the video search request;

[0158] Acquire text feature data of the search text, and determine a first similarity index value between the text feature data and each of the video feature data;

[0159] According to the first similarity index value, determining multiple target videos from the multiple initial videos, and determining a first target video segment corresponding to each target video;

[0160] Inputting the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text;

[0161] According to the first target video segment and the second similarity index value, a retrieval result is generated and returned.

[0162] In some embodiments of the present invention, generating and returning a search result according to the first target video segment and the second similarity index value includes:

[0163] Selecting, from the first target video segments, a plurality of second target video segments whose second similarity index value is greater than or equal to a preset threshold;

[0164] determining a minimum video start time and a maximum video end time among the plurality of second target video segments;

[0165] According to the minimum video start time and the maximum video end time, intercepting the target video corresponding to the first target video segment, obtaining a third target video segment;

[0166] A retrieval result is generated according to the third target video segment and returned.

[0167] In some embodiments of the present invention, generating and returning a search result based on the captured target video includes:

[0168] Determine an average value of a plurality of second similarity index values ​​corresponding to each of the third target video segments to obtain a target similarity value;

[0169] The third target video is sorted in descending order according to the target similarity value to obtain the search result, and the search result is returned.

[0170] In some embodiments of the present invention, before acquiring the multiple initial videos, the method further includes:

[0171] Constructing an initial multimodal model;

[0172] Acquire video sample data and text sample data; wherein the video sample data is an image frame of a video, and the text sample data is a description text of the image frame;

[0173] Training the initial multimodal model according to the video sample data and the text sample data;

[0174] Adding a neural network layer to the trained initial multimodal model to obtain a target multimodal model; wherein the neural network layer is used to extract fusion features of the video sample data and the text sample data;

[0175] The target multimodal model is trained according to the video sample data and the text sample data to obtain a pre-trained target multimodal model.

[0176] In some embodiments of the present invention, the method further comprises:

[0177] In the process of training the initial multimodal model, determining the timing information of the video sample data;

[0178] Feature extraction is performed on the video sample data through the self-attention mechanism of the initial multimodal model and the timing information.

[0179] In some embodiments of the present invention, the neural network layer consists of alternating layers of local multi-head self-attention modules and multi-layer perceptrons.

[0180] Some embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned video retrieval method is implemented.

[0181] Some embodiments of the present invention further provide a computer program product, including a computer program, which implements the above video retrieval method when executed by a processor.

[0182] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0183] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0184] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0185] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0186] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0187] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0189] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0190] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the above elements.

[0191] The above provides a detailed introduction to the video retrieval method and device, product, electronic device, and storage medium. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A video retrieval method, characterized in that: The method comprises: Acquire multiple initial videos, and divide each of the initial videos into multiple video segments; Obtaining and saving video feature data of each of the video clips; In response to a video search request, obtaining a search text corresponding to the video search request; Acquire text feature data of the search text, and determine a first similarity index value between the text feature data and each of the video feature data; According to the first similarity index value, determining multiple target videos from the multiple initial videos, and determining a first target video segment corresponding to each target video; Inputting the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text; According to the first target video segment and the second similarity index value, a search result is generated and returned.

2. The method according to claim 1, characterized in that The step of generating and returning a search result according to the first target video segment and the second similarity index value includes: Selecting, from the first target video segments, a plurality of second target video segments whose second similarity index value is greater than or equal to a preset threshold; determining a minimum video start time and a maximum video end time among the plurality of second target video segments; According to the minimum video start time and the maximum video end time, intercepting the target video corresponding to the first target video segment, obtaining a third target video segment; A retrieval result is generated according to the third target video segment and returned.

3. The method according to claim 2, characterized in that The step of generating and returning a search result based on the intercepted target video includes: Determine an average value of a plurality of second similarity index values ​​corresponding to each of the third target video segments to obtain a target similarity value; The third target video is sorted in descending order according to the target similarity value to obtain the search result, and the search result is returned.

4. The method according to any one of claims 1 to 3, characterized in that: Before acquiring the multiple initial videos, the method further includes: Constructing an initial multimodal model; Acquire video sample data and text sample data; wherein the video sample data is an image frame of a video, and the text sample data is a description text of the image frame; Training the initial multimodal model according to the video sample data and the text sample data; Adding a neural network layer to the trained initial multimodal model to obtain a target multimodal model; wherein the neural network layer is used to extract fusion features of the video sample data and the text sample data; The target multimodal model is trained according to the video sample data and the text sample data to obtain a pre-trained target multimodal model.

5. The method according to claim 4, characterized in that The method further comprises: In the process of training the initial multimodal model, determining the timing information of the video sample data; Feature extraction is performed on the video sample data through the self-attention mechanism of the initial multimodal model and the timing information.

6. The method according to claim 4, characterized in that The neural network layer consists of alternating layers of local multi-head self-attention modules and multi-layer perceptrons.

7. A video retrieval device, characterized in that: The device comprises: A video segment processing module, used for acquiring a plurality of initial videos and dividing each of the initial videos into a plurality of video segments; A video feature data acquisition module, used to acquire and save the video feature data of each video clip; A search text acquisition module, used to respond to a video search request and acquire a search text corresponding to the video search request; A first similarity index value determination module, used for obtaining text feature data of the search text, and determining a first similarity index value between the text feature data and each of the video feature data; a first target video segment determining module, configured to determine a plurality of target videos from the plurality of initial videos according to the first similarity index value, and determine a first target video segment corresponding to each target video; a second similarity index value determination module, configured to input the first target video segment and the search text into a pre-trained target multimodal model to obtain a second similarity index value between the first target video segment and the search text; The retrieval result generating module is used to generate and return the retrieval result according to the first target video segment and the second similarity index value.

8. A computer program product, characterized in that The invention comprises a computer program, which implements the video retrieval method according to any one of claims 1 to 6 when being executed by a processor.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the video retrieval method according to any one of claims 1 to 6 when executed by the processor.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the video retrieval method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Combined fragment retrieval method and device

    CN122432361A