Video retrieval method and apparatus, computer device, and storage medium
By combining video encoding and text encoding with a multimodal model to calculate saliency markers, the problem of insufficient accuracy and convenience in traditional video retrieval methods is solved, realizing efficient and convenient video clip retrieval and improving the utilization efficiency of video materials.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-07-23
AI Technical Summary
In existing technologies, relying on traditional video production models to create media data such as product descriptions, video explanations, and promotional videos requires searching a large amount of material. Moreover, the accuracy and convenience of searching by relying on human memory or file names are limited, resulting in slow video production, high creative thresholds, and low material utilization efficiency, especially the inability to effectively utilize clips in long videos.
By acquiring video data and related text data, encoding them using video encoders and text encoders, calculating saliency markers using a multimodal model, and using a multilayer perceptual layer and decoder module for video segment retrieval, efficient and convenient video retrieval is achieved.
It improves the accuracy and convenience of video retrieval, enabling users to quickly find the video clips they need through text descriptions, reducing the difficulty of creation and learning, and improving the utilization efficiency of video materials.
Smart Images

Figure CN2025145334_23072026_PF_FP_ABST
Abstract
Description
Video retrieval methods, devices, computer equipment and storage media
[0001] This application claims priority to Chinese Patent Application No. 202510072749.5, filed on January 16, 2025, entitled "Video Retrieval Method, Apparatus, Computer Equipment and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of machine learning technology, and in particular to a video retrieval method, apparatus, computer device, and storage medium. Background Technology
[0003] With the rapid rise of online media technologies such as live streaming, digital humans, and short videos, interest in using these new technologies for product promotion, customer marketing, and public relations has grown stronger among tech companies in fields like smart healthcare and smart cities, as well as financial institutions such as banks, insurance companies, and securities firms. However, relying on traditional video production methods to create media data such as product explanations, video narrations, and promotional videos often requires searching through a large amount of material for editing. The accuracy and convenience of searching by human memory or filenames are often limited, preventing creators from efficiently and conveniently utilizing their accumulated video material. In particular, clips from longer videos cannot be effectively used, resulting in slow video production, high creative barriers, and low material utilization efficiency. Summary of the Invention
[0004] This application discloses a video retrieval method, apparatus, computer equipment, and storage medium, which solves the problem of limited accuracy and convenience for users to retrieve video clip data.
[0005] Firstly, this application provides a video retrieval method, including:
[0006] Acquire video data and related text data;
[0007] Video data is encoded using a video encoder to obtain video encoded data, and text data is encoded using a first text encoder and a second text encoder to obtain first text encoding and second text encoding.
[0008] Based on the first text encoding, the second text encoding, and the video encoding data, the correlation value data of the video segment is calculated using a multimodal model;
[0009] Multiple salient vectors are obtained, and salient labels are calculated based on these salient vectors, video segment correlation value data, and video encoding data.
[0010] The saliency label and the video data are input into a retrieval model to obtain target video data.
[0011] In a second aspect, the present application provides a model training device, comprising:
[0012] a data information acquisition module, configured to acquire video data and text data related to the video data;
[0013] a data feature extraction module, configured to encode the video data by using a video encoder to obtain video encoded data, and encode the text data by using a first text encoder and a second text encoder to obtain first text encoding and second text encoding;
[0014] a data feature interaction module, configured to calculate video clip correlation value data by using a multi-modal model according to the first text encoding, the second text encoding, and the video encoded data;
[0015] a saliency label calculation module, configured to acquire a plurality of saliency vectors, and calculate saliency labels according to the plurality of saliency vectors, the video clip correlation value data, and the video encoded data;
[0016] a target data retrieval module, configured to input the saliency labels and the video data into a retrieval model to obtain target video data.
[0017] In a third aspect, the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the video retrieval method provided by any of the embodiments of the present application is implemented.
[0018] In a fourth aspect, the present application provides a non-volatile computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor implements the video retrieval method provided by any of the embodiments of the present application.
[0019] The video retrieval method, device, computer device and storage medium described above, by obtaining video data and text data related to the video data, encoding the text data to obtain first text encoding and second text encoding, and encoding the video data to obtain video encoding data, calculating the attention weight by using a multi-modal model according to the first text encoding, the second text encoding and the video encoding data, calculating the video segment correlation value data according to the attention weight and the video encoding data, obtaining a plurality of saliency vectors, calculating the saliency mark according to the video segment correlation value data and the video encoding data, and inputting the saliency mark and the video data into a retrieval model to obtain target video data. By using multi-modal technologies such as feature extraction and feature interaction, users can search for the required video through text description, so as to more efficiently and conveniently carry out creation and editing work.
[0020] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0022] FIG. 1 is a step schematic flow chart of a video retrieval method according to an embodiment of the present application;
[0023] FIG. 2 is a step schematic flow chart of a data feature extraction method according to an embodiment of the present application;
[0024] FIG. 3 is a schematic block diagram of a neural network model according to an embodiment of the present application;
[0025] FIG. 4 is a step schematic flow chart of a saliency mark calculation method according to an embodiment of the present application;
[0026] FIG. 5 is a schematic block diagram of a neural network model according to an embodiment of the present application;
[0027] FIG. 6 is a step schematic flow chart of a video retrieval method according to an embodiment of the present application;
[0028] FIG. 7 is a step schematic flow chart of a model training method according to an embodiment of the present application;
[0029] FIG. 8 is a schematic block diagram of a video retrieval device according to an embodiment of the present application;
[0030] Figure 9 is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0031] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0034] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0035] It should be understood that, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first" and "third" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. For example, first data and third data are only used to distinguish different data and do not limit their order. Those skilled in the art will understand that the terms "first" and "third" do not limit the quantity or execution order, and the terms "first" and "third" are not necessarily different.
[0036] It should also be understood that the term "and / or" as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0037] To facilitate understanding of the embodiments of this application, some terms involved in the embodiments of this application will be briefly explained below.
[0038] 1. BERT (Bidirectional Encoder Representations from Transformers): A pre-trained language representation model based on the Transformer architecture. Its core feature is its ability to capture bidirectional contextual information of text, thus performing well in various natural language processing (NLP) tasks. Its input representations often include Token Embeddings, which convert words in the text into word embedding vectors; Segment Embeddings, which distinguish the two sentences in a sentence pair and are often used for question answering tasks; and Position Embeddings, which provide the model with positional information of words in the sentence.
[0039] 2. Transformer: The Transformer model consists of an encoder and a decoder, both of which contain multiple identical layers. Each layer consists of a self-attention sub-layer and a feed-forward fully connected sub-layer, with residual connections and layer normalization between these sub-layers.
[0040] 3. Token Embedding Layer: In the Transformer model, the token embedding layer is the first step in processing the input data. This layer converts words in the text into fixed-dimensional vectors to capture the semantic information of the words.
[0041] 4. Positional Encoding Layer: A crucial component of the Transformer, it provides the model with information about the position of words within the sequence. This is because the Transformer's self-attention mechanism itself does not contain information about word order. Positional encoding is typically achieved by adding a vector with the same dimension as the word embeddings, which contains information about the word's position.
[0042] 5. Multi-Head Attention Layer: One of the components of Transformer, used to learn information in parallel in different representation subspaces, allowing the model to focus on different parts of the input sequence at the same time, thereby better capturing complex dependencies.
[0043] 6. Self-Attention Layer: One of the components of Transformer, it allows the model to capture the dependencies between any two positions within a sequence when processing sequential data. It is widely used in the fields of Natural Language Processing (NLP) and Computer Vision (CV).
[0044] 7. Residual Layer: A key component in the Transformer, it helps the model efficiently pass information in deep networks by adding a skip connection, which helps solve the gradient vanishing problem in deep networks.
[0045] 8. Batch Normalization: In machine learning, this refers to a technique used to improve the training speed and stability of a model. Typically placed after each sub-layer in the model, it normalizes the input to reduce internal covariate shift, preventing variations in the network layer input distribution from adversely affecting the training process.
[0046] 9. Softmax: A commonly used activation function. The Softmax function is used to transform the raw scores of the output layer into probability distributions that represent the model's confidence in predicting each possible output.
[0047] 10. Activation Layer: A crucial component of neural networks, introducing non-linear characteristics that enable the network to learn and perform more complex tasks. Activation functions such as ReLU, Sigmoid, Leaky ReLU, and Softmax can be used to introduce non-linear factors, thereby enhancing the expressive power of deep neural networks. Among these, ReLU is a non-saturating activation function that can alleviate the gradient vanishing problem caused by the large number of layers in deep neural networks and accelerate convergence.
[0048] 11. ViT (Vision Transformer) model: A deep learning model based on the Transformer architecture for processing visual tasks. Its core idea is to extend the Transformer model from Natural Language Processing (NLP) to the field of computer vision, using attention mechanisms to capture features and patterns in images.
[0049] 12. ViViT: A ViT model that introduces inductive biases of locality in a video Transformer, including a video encoder capable of converting video data into video encoded data.
[0050] 13. Video Swin Transformer: A ViT model that improves the quality of video representations by combining large-scale pre-trained visual prior knowledge with video-level temporal bias.
[0051] 14.MVit: A multi-scale visual ViT model for video and image recognition.
[0052] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0053] From tech startups in smart healthcare and smart cities to financial institutions like banks, insurance companies, and securities firms, there's a growing interest in using online media technologies like virtual live streaming, digital humans, and short videos for product promotion, customer marketing, and public relations. However, relying on traditional video production methods to create product descriptions, video explanations, and promotional videos often requires retrieval of vast amounts of material for editing. The accuracy and convenience of searching by memory or filenames are often limited, preventing creators from efficiently utilizing their accumulated video footage. This is particularly true for segments from longer videos, leading to slow video production, high creative barriers, and low material utilization efficiency. Furthermore, it makes it difficult for clients to independently search for and learn about product information, hindering their ability to directly locate key segments within product videos that meet their needs.
[0054] In practice, financial institutions such as insurance companies often produce introductory or marketing videos for their financial products. However, these products typically involve numerous terms and complex content, requiring lengthy filming or being split into multiple episodes. Even when split into multiple episodes, each episode's explanations and elaborations remain relatively lengthy to ensure professionalism, compliance, and completeness. Because creators, editors, and investors often lack sufficient understanding of technical terminology or have very specific questions, they may struggle to quickly identify the most interesting and relevant parts based on the video title and summary, leading to inconvenience and difficulties in creation, learning, and comprehension.
[0055] Machine learning technology is being widely applied in information retrieval and data search. Among them, multimodal model technology has been widely used in autonomous driving, medical image analysis, computer-aided design and creation, and has the potential to continuously empower video retrieval and artistic creation. It may enable users to search for the videos they need through text descriptions, thereby making it easier to carry out creation, editing and investor education work.
[0056] Furthermore, while existing video retrieval technologies based on multimodal models each have their advantages, they are currently largely confined to the academic and laboratory stages. Therefore, limited by the computing power and data acquisition capabilities of laboratories, existing multimodal model-based video retrieval technologies are often trained on standardized datasets specific to particular scenarios such as movies, housework, and driving. This results in limited generalization ability and the inappropriate use of techniques like model distillation. While adaptable to the limited computing environment in laboratories, this affects the model's accuracy, generalization ability, neural network flexibility, and robustness. Thus, although existing technologies have achieved good results on specific datasets, they cannot adequately meet the specific needs of financial scenarios such as product marketing and customer consultation in terms of data processing, model architecture, and training processes. These needs include the requirements for the professionalism, rigor, and completeness of retrieval results, as well as the accuracy, robustness, and security requirements of commercial artificial intelligence.
[0057] To address the aforementioned problems, this application proposes a video retrieval method. Please refer to Figure 1, which is a flowchart illustrating the steps of a video retrieval method provided in this application.
[0058] As shown in Figure 1, the video retrieval method specifically includes steps S101 to S105.
[0059] S101. Obtain video data and text data related to the video data.
[0060] For example, video data may include one or more types of video data such as financial video data, product video data, music video data, health video data, promotional video data, introductory video data, educational video data, documentary video data, film and television video data, and short video data.
[0061] Furthermore, the text data may include one or more of the following text data related to the video data: query information, descriptive information, and subtitle information. For example, questions about insurance laws and regulations, questions about insurance product terms, and consultation information about the current insurance asset management portfolio.
[0062] S102. Video data is encoded using a video encoder to obtain video encoded data, and text data is encoded using a first text encoder and a second text encoder to obtain first text encoding and second text encoding.
[0063] For example, video encoders with video encoding capabilities, such as ViViT, MViT, or Video Swin Transformer, can be used to encode video data.
[0064] It is understandable that video encoded data can include multiple video feature representation vectors, and each video feature representation vector can correspond one-to-one with each frame in the video data.
[0065] The video data or video-encoded data may also include labeled data obtained by annotating the corresponding text data, in order to facilitate subsequent model training and improvement. The video data, video-encoded data, or labeled data may also include time information, contextual information, and event information.
[0066] For example, video data can be frame-stripped to obtain multiple video frames; these multiple video frames can be input into a video encoder to obtain video feature representation vectors corresponding to the video frames; based on the correlation between the corresponding video frames and text data, the video feature representation vectors can be labeled to obtain video encoded data, wherein the video encoded data includes multiple video feature representation vectors and corresponding labeled data.
[0067] It should be noted that the first text encoding may include multiple first text feature expression vectors, and the second text encoding may include multiple second text feature expression vectors. The first text feature expression vectors and the second text feature expression vectors can correspond one-to-one, and the first text feature expression vectors can correspond one-to-one with each sentence in the text data.
[0068] In some embodiments, please refer to FIG2, which is a schematic flowchart illustrating the steps of a data feature extraction method provided in this application embodiment. This data feature extraction method can be used to implement the above-described step S102.
[0069] As shown in Figure 2, the data feature extraction method specifically includes steps S102a to S102c.
[0070] It is understood that steps S102a to S102c can be implemented using a multimodal model, which may include a first text encoder, a second text encoder, and a video encoder.
[0071] S102a. Input the text data into the first text encoder to obtain the first text encoding.
[0072] For example, the encoder of a model with text encoding capabilities, such as BERT, T5, CLIP, or GPT, can be used to encode the text data to obtain the first text encoding.
[0073] S102b: Input the first text encoding into the second text encoder to obtain the feature text encoding.
[0074] The second text encoder may include one or more self-attention modules.
[0075] Please refer to Figure 3, which is a block diagram of a neural network model provided in an embodiment of this application. This neural network model can be used as a second text encoder to implement the above step S102b.
[0076] As shown in Figure 3, the neural network model can include multiple self-attention modules, wherein each self-attention module includes a normalization layer, a self-attention layer, a residual processing layer, and a feedforward network layer.
[0077] Among them, the self-attention layer can be a multi-head self-attention layer.
[0078] For example, for a self-attention module, the input vector of the self-attention module can be normalized to obtain an intermediate vector A; the intermediate vector A can be copied three times to obtain the Q vector, K vector, and V vector; using a multi-head self-attention algorithm, the intermediate vector B can be calculated based on the Q vector, K vector, and V vector; the average value of the intermediate vector B and the intermediate vector A can be calculated to obtain the intermediate vector C; the intermediate vector C can be input into a feedforward neural network to obtain the intermediate vector D; the intermediate vector D can be added to the intermediate vector B and then normalized to obtain the intermediate vector E.
[0079] The normalization method can be ReLU and / or SoftMax, and the feedforward neural network can be fully connected.
[0080] It should be noted that if the neural network model includes multiple self-attention modules, the intermediate vector E is used as the input vector of the next self-attention module; if the neural network model includes only one self-attention module or the self-attention module is the last self-attention module of the neural network model, then the intermediate vector E is the feature text encoding.
[0081] It should be noted that before inputting the first text encoding into the second text encoder, the process may include obtaining an initial virtual vector; concatenating the initial virtual vector with the first text encoding; and then inputting it into the first self-attention module of the second text encoder to obtain the virtual vector.
[0082] The method for obtaining the initial virtual vector includes initializing the value of the initial virtual vector according to a Gaussian distribution. The dimensions of the initial virtual vector and the virtual vector can be consistent with the dimensions of the first text encoding.
[0083] Understandably, traditional methods often fail to consider the varying importance of different words in text data. For example, in the financial and insurance fields, modal particles and auxiliary words are not very helpful in determining semantics and video retrieval, while specific nouns, verbs, and adverbs are much more helpful in determining semantics and video retrieval. The second text encoder used in this application leverages the advantages of the attention mechanism, enabling feature text encoding to focus more on keywords in the text data, capture richer semantic and contextual information, and improve the generalization ability and accuracy of the video retrieval method in this application.
[0084] Furthermore, the neural network model can include three self-attention modules. An appropriate number of self-attention modules also avoids the waste of computational resources and overfitting problems caused by stacking too many self-attention modules.
[0085] S102c: Concatenate the first text code with the feature text code to obtain the second text code.
[0086] Specifically, methods for concatenating the first text encoding with the feature text encoding can include simple concatenation and weighted summation.
[0087] It should be noted that the concatenated vector can contain both the original text features and high-level features extracted through the self-attention mechanism, thus providing a more comprehensive text representation. This helps the model capture more subtle language features and contextual relationships, thereby improving the model's performance in video retrieval tasks.
[0088] It should be noted that in some embodiments, the first text code may not be concatenated with the feature text code, and the feature text code may be directly used as the second text code to obtain the second text code.
[0089] For example, video data can be input into a video encoder to obtain video encoded data; text data can be input into a first text encoder to obtain a first text encoding; the first text encoding can be input into a second text encoder to obtain a feature text encoding. The feature text encoding can then be used as the second text encoding to obtain a second text encoding.
[0090] S103. Based on the first text encoding, the second text encoding, and the video encoding data, the correlation value data of the video segment is calculated using a multimodal model.
[0091] For example, the multimodal model may include one or more of a first text encoder, a second text encoder, and a video encoder, which may be part of the multimodal model or used independently.
[0092] For example, the multimodal model may also include three linear mapping layers, pQ, pK, and pV, and a self-attention layer. The multimodal model may also include one or more self-attention layers, where each self-attention layer may include the three linear mapping layers pQ, pK, and pV.
[0093] It should be noted that the self-attention layer may include a multi-head self-attention layer and / or an adaptive attention layer.
[0094] For example, the first text encoding can be input into the linear mapping layer pV to obtain the V distribution; the second text encoding can be input into the linear mapping layer pK to obtain the K distribution; the video encoding data can be input into the linear mapping layer pQ to obtain the Q distribution; the Q distribution, K distribution and V distribution can be input into the adaptive attention layer to calculate the attention weight W; and the video segment correlation value data can be calculated based on the attention weight W.
[0095] The Q-distribution, K-distribution, and V-distribution are input into the adaptive attention layer to calculate the attention weight W. The corresponding calculation formula may include:
[0096] Among them, L q L can represent the number of codes or word tokens contained in the first text encoding. d L can represent the number of codes or virtual vectors included in the feature text encoding. q +L d Then it can represent the number of codes included in the second text encoding, ⊙ can represent the dot product, h can represent the number of latent vector dimensions of the linear mapping layer, and W i,j V can represent the attention weights corresponding to the i-th video frame and the j-th first text code. j Q can represent the vector in the V distribution corresponding to the j-th code in the first text encoding. i K can represent the vector in the Q-distribution corresponding to the i-th video frame in the video coded data. j It can represent the vector in the K-distribution corresponding to the j-th code in the first text encoding, where K... k It can represent the vector in the K-distribution corresponding to the k-th code of the feature text encoding.
[0097] Furthermore, calculation formulas can be used. Based on the attention weight W, the relevance data of the video segment is calculated.
[0098] For example, the video segment correlation value data corresponding to the i-th video frame is
[0099] S104. Obtain multiple saliency vectors and calculate saliency markers based on video segment correlation value data and video encoding data.
[0100] Understandably, saliency markers can be used to express the degree of relevance between each video frame and its corresponding text in video coded data.
[0101] For example, obtaining multiple salient vectors P and calculating saliency markers based on video segment correlation value data and video encoding data can include: obtaining multiple salient vectors P and calculating saliency markers based on the video feature representation vector v and mean vector V corresponding to each video frame in the video encoding data. ctx Multiple salient vectors P and correlation values of video segments The correlation matrix is calculated; the correlation matrix is processed using the Softmax or ReLU algorithm to calculate the candidate weight vector C; the candidate weight vector C is sorted according to the magnitude of each dimension value to determine the set of the top K largest values {C}. (1) C (2) ,...,C (K)}; Based on the mean vector V ctx For each value in this set and its corresponding saliency vector P, the saliency label T is calculated. The corresponding calculation formula may include:
[0102] Where T can represent a significance marker, L v L can represent the total number of video frames or the video feature representation vectors corresponding to each video frame in the video coded data. p P can represent the total number of salient vectors. j The above L can represent p The j-th salient vector among the salient vectors, C j It can be represented as the mean vector V. ctx The weight vectors can be calculated using the arithmetic mean, geometric mean, or weighted mean of the video feature representation vectors v corresponding to each video frame in the video encoded data. The dimension of the candidate weight vectors can be the same as the number of the multiple salient vectors P mentioned above. K is a positive integer, and K is less than L. p .
[0103] Furthermore, the mean vector V ctx The arithmetic mean, geometric mean, or weighted mean used in the calculation can be the average value of each video feature representation vector across various dimensions, thus yielding the mean vector V. ctx It can be consistent with the number of dimensions of the video feature representation vector.
[0104] It should be noted that when obtaining multiple significant vectors, the number of significant vectors L... p The number of salient vectors can be between 5 and 15, and the dimension of the salient vector can be consistent with the number of dimensions of the video feature representation vector v.
[0105] Please refer to Figure 4, which is a schematic flowchart illustrating the steps of a saliency marker calculation method provided in an embodiment of this application, and can be used to implement the above-described step S104. In some embodiments, as shown in Figure 4, the saliency marker calculation method specifically includes steps S104a to S104d.
[0106] S104a. Based on the video coding data and multiple salient vectors, the correlation matrix is calculated.
[0107] For example, the mean vector V is calculated based on the video feature representation vector v corresponding to each video frame in the video encoded data. ctx Based on multiple saliency vectors P and correlation values of video segments and the mean vector V ctx The correlation matrix is then calculated.
[0108] S104b: Calculate the candidate weight vector based on the correlation value data and correlation matrix of the video segments.
[0109] For example, the correlation matrix is processed using the Softmax function, and the processing result is compared with the correlation values of the video segments. The transpose of the matrix is multiplied, and the matrix multiplication result is summed column-wise to obtain the candidate weight vector.
[0110] S104c. Based on the candidate weight vector and multiple salient vectors, determine and calculate the instantaneous description vector.
[0111] For example, the dimension with the highest value in the candidate weight vector is obtained, and its value is taken as a one-dimensional vector. The inner product of this vector and the salient vector corresponding to the ordinal number of the corresponding dimension is taken to obtain the instantaneous description vector.
[0112] Alternatively, obtaining the dimension with the highest value in the candidate weight vector can be replaced by obtaining multiple dimensions with the highest values. If multiple dimensions with the highest values in the candidate weight vector are obtained, the values of these multiple dimensions are used as 1x1 matrices, and the inner product of each matrix and the salient vector corresponding to the ordinal number of the respective dimension is calculated to obtain multiple instantaneous description vectors.
[0113] S104d. Calculate one or more saliency markers based on the instantaneous description vector.
[0114] For example, the instantaneous description vector and the mean vector V are compared. ctxSumming yields the significance markers.
[0115] For example, multiple instantaneous description vectors are respectively compared with the mean vector V. ctx Adding them together yields multiple saliency markers.
[0116] S105. Input the saliency markers and video data into the retrieval model to obtain the target video data.
[0117] For example, a saliency marker is input into a multilayer perceptron to obtain a mapping vector; the mapping vector is concatenated with video data to obtain an input vector for inputting into a retrieval model; the input vector is input into the retrieval model to obtain one or more time-segment data; and the video data is processed based on the one or more time-segment data to obtain the target video data.
[0118] The process of processing video data based on one or more time periods to obtain target video data may include using a Hungarian matching algorithm to segment the video data based on the one or more time periods and then evaluating and filtering the data to obtain the target video data.
[0119] It is understandable that the output of the retrieval model can be either target video data or time period data. For computer devices such as edge computing devices, mobile communication devices, and server devices, the actual performance of neural network models will vary due to limitations in hardware computing power and system design. Therefore, depending on requirements and actual conditions, time period data can be obtained on one device, and then another device can process the video data based on that time period data to obtain the target video data. This makes any of the video retrieval methods provided in the embodiments of this application more efficient and accurate.
[0120] Please refer to Figure 5, which is a block diagram of another neural network model provided in this application embodiment. This neural network model can be used as a retrieval model to implement the above step S105. As shown in Figure 5, the retrieval model may include at least a multi-layer perceptron (MLP), a fully connected layer, an encoder module, and a decoder module.
[0121] The encoder module and decoder module can be encoder and decoder modules of Transformer architecture such as Transformer, DETR or BERT.
[0122] Furthermore, this multilayer perceptron can use activation functions including Sigmoid, Tanh, ReLU, etc.
[0123] Furthermore, the self-attention layer may include a multi-head self-attention layer, a cross-attention layer, and / or an adaptive self-attention layer.
[0124] Furthermore, this feedforward network layer can be fully connected.
[0125] Furthermore, the retrieval model may also include a position encoder and / or a residual processing layer.
[0126] It is understandable that by adding a position encoder, computer devices can learn deep grammatical and semantic features based on the positional information of words in sentences or video frames in video data, and capture long-distance dependencies, thereby improving the model's understanding and generation capabilities, and thus enhancing the accuracy and robustness of the video retrieval method in this application embodiment.
[0127] For example, a saliency marker is input into the perception layer to obtain a mapping vector; the mapping vector is concatenated with the video encoded data to obtain the encoder input vector; this input vector is input into the encoder module to obtain a position encoding vector and a feature encoding vector; multiple auxiliary vectors are obtained and added to the position encoding vector respectively to obtain the decoder input vector; the encoder input vector is copied to obtain vector K and vector V; the decoder input vector is used as vector Q, and together with vector K and vector V, it is input into the decoder module to obtain time period data or target video data.
[0128] The position encoding vector can be obtained by using the position encoder in the encoder module to perform position encoding on the input vector. The position encoding method can include absolute position encoding or relative position encoding.
[0129] The dimensions of the multiple auxiliary vectors are consistent with the positional encoding. The initial multiple auxiliary vectors can all be zero vectors, and the number of auxiliary vectors can be between 10 and 15.
[0130] Representing key content in a video using multiple auxiliary vectors prevents overlap and duplication of retrieved target video data, allowing for a more efficient allocation of the model's attention. Using 10-15 auxiliary vectors enables semantic extraction of multiple key elements from the video data without significantly increasing the model's computational power consumption. This improves the model's versatility and the comprehensiveness of the retrieval method, preventing missed or false detections.
[0131] In some embodiments, please refer to FIG6, which is a schematic flowchart of another video retrieval method provided in this application embodiment. This video retrieval method can be used to implement the above-described step S105. As shown in FIG6, the video retrieval method specifically includes steps S105a to S105e.
[0132] S105a. Input the saliency marker into the multilayer perceptron and concatenate the output of the multilayer perceptron with the video data to obtain the encoder input vector.
[0133] For example, the multilayer perceptron (MLP module or MLP layer) may include activation layers and fully connected layers that utilize activation functions such as Sigmoid, Tanh or ReLU.
[0134] It is understandable that this multi-layer perception layer can be part of the retrieval model or a standalone multi-layer perception layer. This multi-layer perception layer can be deployed on the computer device where the retrieval model resides, or it can be deployed independently on another computer device and connected to the computer device where the retrieval model resides via communication technology. Distributed deployment allows for better integration of the design advantages of different computer devices, thereby improving system robustness and smoothness, reducing computational overhead, and saving human and material resources.
[0135] S105b: Input the encoder input vector into the encoder module to obtain the feature encoding vector and the position encoding vector.
[0136] As shown in Figure 5, the encoder module may include a position encoder, a self-attention layer, a residual processing layer, and a feedforward network layer.
[0137] For example, the encoder input vector is input to the position encoder of the encoder module to obtain the position encoded vector. The formula for position encoding may include:
[0138] In this context, the dimension of the position encoding vector is the same as the total dimension of the encoder input vector. PE can represent the position encoding vector, pos can represent the position, and d... model It can represent the total number of dimensions of the position encoding vector or encoder input vector, where i can represent the dimension.
[0139] Furthermore, the encoder input vector is input to the encoder module to obtain the feature encoding vector.
[0140] It should be noted that this encoder module can be an encoder based on open-source architectures such as Transformer or BERT, and it can include a position encoder. This encoder module can be part of a retrieval model.
[0141] Understandably, a position encoder enables the model to better capture deeper semantic and contextual information, thereby improving the accuracy of video retrieval methods.
[0142] S105c: Obtain multiple initialization vectors, and add the position encoding vector to each of the multiple initialization vectors to obtain multiple auxiliary vectors.
[0143] For example, using the zero vector as the initialization vector, 12 initialization vectors are obtained, and 12 position codes are obtained. These are then added to the 12 initialization vectors to obtain 12 auxiliary vectors.
[0144] Understandably, the number of auxiliary or initialization vectors can be adjusted based on the specific type of video data. For example, in the financial industry, professional explanation video data typically contains 10 to 15 salient elements; therefore, the number of auxiliary or initialization vectors can be set to within the range of 10 to 15. If the number is too small, the granularity of video salient segmentation may be too large, resulting in insufficient precision. If the number is too large, the granularity will be too fine, and overlapping results may occur, reducing the clarity and completeness of the search results.
[0145] S105d: Input the encoder input data, feature encoding vector, and multiple auxiliary vectors into the decoder module to obtain multiple decoder output vectors.
[0146] The decoder module can include decoders based on open-source architectures such as Transformer and BERT.
[0147] Furthermore, as shown in Figure 5, the decoder module may include a multi-head self-attention layer, a residual processing layer, and a feedforward network layer.
[0148] Understandably, the decoder module may also include a fully connected layer.
[0149] S105e inputs multiple decoder output vectors into a fully connected layer and uses the Hungarian matching algorithm to match and obtain the target video data.
[0150] For example, multiple decoder output vectors can be input into a fully connected layer to obtain multiple time-segment data; the inner product of each of these multiple time-segment data and the feature encoding vector can be calculated to obtain multiple hit values; based on these multiple hit values, the X time-segment data with the largest corresponding hit values can be determined; based on these X time-segment data, the video data can be cut to obtain the target video data.
[0151] Where X is a positive integer.
[0152] It is understandable that the fully connected layer can be either a fully connected layer included in the retrieval model or an independent fully connected layer.
[0153] Understandably, by utilizing the Hungarian matching algorithm, the processed target video data can be made to avoid overlap with each other, provided that they are highly relevant to the text, thereby ensuring the accuracy of the video retrieval method and the usability of the retrieval results.
[0154] In some embodiments, please refer to FIG7, which is a schematic flowchart of the model training method provided in this application embodiment. As shown in FIG7, the model training method specifically includes steps S201 to S204.
[0155] S201. Obtain the labeled data of the video data.
[0156] It is understandable that the video feature representation vectors corresponding to the video data can be labeled according to the correlation between the video data and the text data, thereby obtaining the labeled video data.
[0157] For example, if the first to fifth video frames in the video data corresponding to the first to fifth video feature representation vectors v1 to v5 in the video encoded data are related to a sentence in the text data, then the first to fifth video feature representation vectors v1 to v5 are labeled as the first video feature representation vector. Up to the fifth video feature representation vector
[0158] For example, if the sixth to ninth video frames in the video data corresponding to the sixth to ninth video feature representation vectors v6 to v9 in the video encoded data are related to a sentence in the text data, then the sixth to ninth video feature representation vectors v6 to v9 are labeled as the sixth video feature representation vector. Up to the ninth video feature representation vector
[0159] It should be noted that by labeling video encoded data or video feature representation vectors, we can obtain labeled data, which can identify video frames related to text data. This makes subsequent loss calculation and model training more convenient and improves the accuracy of the model.
[0160] S202. Based on the labeled data and the correlation value data of the video segments, calculate the cross-entropy loss data.
[0161] For example, the formula used in the calculation may include:
[0162] in, It can represent cross-entropy loss data, L v It can represent the total number of video frames or the total number of video feature representation vectors corresponding to each video frame in the video coded data. It can represent the relevant value data of video segments, a i It can represent the labeled data of the i-th frame of video data or its corresponding video feature representation vector.
[0163] It should be noted that the labeled data here can be understood as numbers with values of 0 or 1.
[0164] For example, if the video data of the i-th frame is related to the corresponding text data, then a i The value of a is 1. If the video data of the i-th frame is unrelated to the corresponding text data, then a i The value is 0.
[0165] S203. Calculate the orthogonal loss data based on the feature text encoding.
[0166] For example, the formula used in the calculation may include:
[0167] in, It can represent orthogonal loss data, L d It can represent the number of codes or virtual vectors included in the feature text encoding. It can represent the m-th code in the feature text encoding or its corresponding m-th virtual vector. It can represent the nth code in the feature text encoding or the corresponding nth virtual vector.
[0168] S204. Based on the cross-entropy loss data and orthogonal loss data, train a multimodal model using the backpropagation algorithm.
[0169] For example, the total loss data can be obtained by summing the cross-entropy loss data and the orthogonal loss data; based on the total loss data, a multimodal model can be trained using the backpropagation algorithm.
[0170] The summation method can include weighted summation.
[0171] It is understandable that a multimodal model can include a first text encoder, a second text encoder, or a video encoder. During parameter updates, some parameters of the first text encoder, second text encoder, or video encoder can be frozen to improve training efficiency and model quality, and reduce the risks of gradient explosion or gradient vanishing.
[0172] For example, during the training of a multimodal model, the parameters of the video encoder, the first text encoder, and the second text encoder can be frozen. Then, based on the total loss data obtained by weighted summation of the cross-entropy loss data and the orthogonal loss data, the multimodal model can be trained using the backpropagation algorithm.
[0173] Furthermore, one or more of the following algorithms can be used to train the multimodal model: Adam (Adaptive Moment Estimation), RMSprop (Root Mean Square Propagation), and Momentum.
[0174] It is understood that the multimodal model may include a first text encoder and a video encoder, or it may be distinguished from the first text encoder and the video encoder as a separate model.
[0175] As shown in Figure 8, which is a schematic diagram of a video retrieval device provided in an embodiment of this application, the video retrieval device is used to execute the aforementioned video retrieval method. The video retrieval device can be configured on a terminal or a server.
[0176] As shown in Figure 8, the video retrieval device 100 includes a data information acquisition module 101, a data feature extraction module 102, a data feature interaction module 103, a saliency marker calculation module 104, and a target data retrieval module 105.
[0177] The data information acquisition module 101 is used to acquire video data and text data related to the video data;
[0178] The data feature extraction module 102 is used to encode the video data using a video encoder to obtain video encoded data, and to encode the text data using a first text encoder and a second text encoder to obtain first text encoding and second text encoding.
[0179] The data feature interaction module 103 is used to calculate video segment correlation value data based on the first text encoding, the second text encoding and the video encoding data using a multimodal model;
[0180] The saliency marker calculation module 104 is used to obtain multiple saliency vectors and calculate a saliency marker based on the multiple saliency vectors, the video segment correlation value data, and the video encoding data;
[0181] The target data retrieval module 105 is used to input the saliency marker and the video data into the retrieval model to obtain the target video data.
[0182] In some embodiments, the video retrieval device 100 may further include a model training module, which, after calculating video segment correlation value data using a multimodal model based on the first text encoding, the second text encoding, and the video encoding data, calculates cross-entropy loss data based on the labeled data and the video segment correlation value data; calculates orthogonal loss data based on the feature text encoding; and trains the multimodal model using a backpropagation algorithm based on the cross-entropy loss data and the orthogonal loss data.
[0183] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and modules can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0184] The aforementioned device can be implemented as a computer program that can run on the computer device shown in Figure 9.
[0185] Please refer to Figure 9, which is a schematic block diagram of a computer device provided in an embodiment of this application. This computer device may be a server. Referring to Figure 9, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0186] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any video retrieval method.
[0187] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0188] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When these computer programs are executed by a processor, the processor can perform any video retrieval method.
[0189] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that the structure shown in Figure 9 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0190] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0191] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps:
[0192] The process involves: acquiring video data and related text data; encoding the video data using a video encoder to obtain video encoded data; encoding the text data using a first text encoder and a second text encoder to obtain first text encoding and second text encoding; calculating video segment relevance data using a multimodal model based on the first text encoding, the second text encoding, and the video encoded data; acquiring multiple saliency vectors; calculating saliency markers based on the video segment relevance data and the video encoded data; and inputting the saliency markers and the video data into a retrieval model to obtain target video data.
[0193] This application also provides a non-volatile computer-readable storage medium in its embodiments. The non-volatile computer-readable storage medium can be either non-volatile or volatile. The non-volatile computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the video retrieval method provided in any embodiment of this application. For example, the computer program includes program instructions, and the processor executes the program instructions to implement any of the video retrieval methods provided in the embodiments of this application.
[0194] The non-volatile computer-readable storage medium can be an internal storage unit of the computer device described in the foregoing embodiments, such as a hard disk or memory of the computer device. Alternatively, the non-volatile computer-readable storage medium can be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0195] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video retrieval method, comprising: Acquire video data and text data related to the video data; The video data is encoded using a video encoder to obtain video encoded data, and the text data is encoded using a first text encoder and a second text encoder to obtain first text encoding and second text encoding. Based on the first text encoding, the second text encoding, and the video encoding data, the correlation value data of the video segment is calculated using a multimodal model; Multiple salient vectors are obtained, and a salient marker is calculated based on the multiple salient vectors, the video segment correlation value data, and the video encoding data; The saliency markers and the video data are input into the retrieval model to obtain the target video data.
2. The method according to claim 1, wherein, The process of encoding the text data using a first text encoder and a second text encoder to obtain the first text code and the second text code includes: The text data is input into the first text encoder to obtain the first text encoding; The first text encoding is input into the second text encoder to obtain the feature text encoding; The first text code is concatenated with the feature text code to obtain the second text code.
3. The method according to claim 2, wherein, The multimodal model includes the first text encoder, the second text encoder, or the video encoder.
4. The method according to claim 2, wherein, The second text encoder includes multiple self-attention modules, each of which includes a normalization layer, a self-attention layer, a residual processing layer, and a feedforward network layer.
5. The method of claim 4, wherein, Before inputting the first text encoding into the second text encoder, the process includes: Obtain the initial virtual vector; The initial virtual vector is concatenated with the first text encoding and then input into the self-attention module of the second text encoder to obtain a virtual vector, which is used for cross-attention calculation.
6. The method according to claim 5, wherein, The process of obtaining the initial virtual vector includes: The initial virtual vector is initialized according to a Gaussian distribution, so that the dimensions of the initial virtual vector and the virtual vector are consistent with the dimensions of the first text encoding.
7. The method according to claim 3, wherein, After calculating the video segment correlation value data using a multimodal model based on the first text encoding, the second text encoding, and the video encoding data, the method further includes: Obtain the annotation data of the video data; Based on the labeled data and the correlation value data of the video segment, the cross-entropy loss data is calculated; Based on the feature text encoding, orthogonal loss data is calculated; The multimodal model is trained using the backpropagation algorithm based on the cross-entropy loss data and the orthogonal loss data.
8. The method according to claim 7, wherein, The process of training the multimodal model using the backpropagation algorithm includes: Freeze the parameters of the video encoder, the first text encoder, and the second text encoder; The total loss data is obtained by weighted summation of the cross-entropy loss data and the orthogonal loss data; Based on the total loss data, a multimodal model is trained using the backpropagation algorithm.
9. The method according to claim 1, wherein, The step of calculating saliency markers based on multiple saliency vectors, the video segment correlation value data, and the video encoding data includes: Based on the video encoded data and multiple salient vectors, a correlation matrix is calculated; Based on the correlation value data of the video segment and the correlation matrix, a candidate weight vector is calculated; The instantaneous description vector is determined and calculated based on the candidate weight vector and multiple salient vectors; One or more saliency markers are calculated based on the instantaneous description vector.
10. The method according to claim 9, wherein, The step of determining and calculating the instantaneous description vector based on the candidate weight vector and multiple salient vectors includes: Obtain the dimension with the highest value among the candidate weight vectors, treat its value as a one-dimensional vector, and take the inner product with the salient vector corresponding to the ordinal number of the corresponding dimension to obtain the instantaneous description vector; or... The highest-valued dimensions among the candidate weight vectors are obtained. The value of each selected dimension is used as a 1x1 matrix, and the inner product is calculated with the salient vector corresponding to the ordinal number of the corresponding dimension to obtain multiple instantaneous description vectors.
11. The method according to claim 9, wherein, The step of calculating the correlation matrix based on the video encoded data and multiple salient vectors includes: The mean vector is calculated based on the video feature representation vector corresponding to each video frame in the video encoded data. A correlation matrix is calculated based on the multiple salient vectors, the correlation value data of the video segments, and the mean vector.
12. The method according to claim 11, wherein, The step of calculating one or more saliency markers based on the instantaneous description vector includes: The saliency marker is obtained by summing the instantaneous description vector with the mean vector; or, Multiple instantaneous description vectors are added to the mean vector to obtain multiple saliency markers.
13. The method according to claim 1, wherein, The retrieval model includes at least a multi-layer perceptual layer, a fully connected layer, an encoder module, and a decoder module.
14. The method according to claim 13, wherein, The step of inputting the saliency marker and the video data into the retrieval model to obtain the target video data includes: The saliency marker is input into the multilayer perceptron, and the output of the multilayer perceptron is concatenated with the video data to obtain the encoder input vector; The encoder input vector is input to the encoder module to obtain a feature encoding vector and a position encoding vector; Multiple initialization vectors are obtained, and the position encoding vector is added to the multiple initialization vectors respectively to obtain multiple auxiliary vectors; The encoder input data, the feature encoding vector, and the multiple auxiliary vectors are input into the decoder module to obtain multiple decoder output vectors; The output vectors of the multiple decoders are input into the fully connected layer, and the target video data is obtained by matching using the Hungarian matching algorithm.
15. The method according to claim 14, wherein, The step of obtaining multiple initialization vectors and adding the position encoding vector to the multiple initialization vectors respectively to obtain multiple auxiliary vectors includes: Using the zero vector as the initialization vector, obtain multiple initialization vectors; Multiple position codes are obtained and added to multiple initialization vectors respectively to obtain multiple auxiliary vectors.
16. The method according to claim 15, wherein the number of the initialization vector, the position code, and the auxiliary vector are the same; in, The number of initialization vectors ranges from 10 to 15; or, The number of location codes ranges from 10 to 15; or, The number of auxiliary vectors ranges from 10 to 15.
17. The method of claim 14, wherein, The step of inputting the output vectors of the multiple decoders into the fully connected layer and using the Hungarian matching algorithm to match and obtain the target video data includes: The output vectors of the multiple decoders are input into the fully connected layer to obtain data for multiple time periods; The inner product of the multiple time period data and the feature encoding vector is calculated respectively to obtain multiple hit values; Based on the multiple hit values, X time periods with the highest hit values are determined, where X is a positive integer; Based on the X time period data, the video data is cut to obtain the target video data.
18. A video retrieval device, comprising: The data information acquisition module is used to acquire video data and text data related to the video data; The data feature extraction module is used to encode the video data using a video encoder to obtain video encoded data, and to encode the text data using a first text encoder and a second text encoder to obtain first text encoding and second text encoding. The data feature interaction module is used to calculate video segment correlation value data based on the first text encoding, the second text encoding, and the video encoding data using a multimodal model; A saliency marker calculation module is used to obtain multiple saliency vectors and calculate a saliency marker based on the multiple saliency vectors, the video segment correlation value data, and the video encoding data; The target data retrieval module is used to input the saliency marker and the video data into the retrieval model to obtain the target video data.
19. A computer device comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the video retrieval method as described in any one of claims 1-17.
20. A non-volatile computer-readable storage medium storing a computer program thereon, the computer program being executed by a processor to cause the processor to implement the video retrieval method as described in any one of claims 1-17.