Efficient Video Retrieval Model Based on Multimodal Feature Fusion

By using NetVLAD network and two-step training method in the video retrieval framework, the overfitting problem of video encoder is solved, and the accuracy and efficiency of video retrieval is improved. Text encoder representation suitable for video retrieval tasks is achieved more efficient video and text matching.

CN114564616BActive Publication Date: 2025-07-18NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210210095.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-07-18
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

In the existing video retrieval framework, especially the design of video encoder, there is an overfitting phenomenon, which cannot effectively process multiple modes of video. In the prior art, the randomization of parameter initialization of video encoder affects the performance of text encoder, resulting in poor retrieval effect.

Method used

The NetVLAD network is used instead of the Transformer network to fuse video features, and the text encoder parameters are frozen through a two-step training method and then the video encoder parameters are fine-tuned. Combined with the CLIP text encoder, the model training parameters are reduced and the video retrieval accuracy is improved.

Benefits of technology

It effectively solves the overfitting problem of video encoder, improves the accuracy and efficiency of video retrieval, and provides text encoder representation suitable for video retrieval tasks, which improves the similarity calculation accuracy of video and text matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114564616B_ABST
    Figure CN114564616B_ABST
Patent Text Reader

Abstract

This paper presents a video retrieval framework, which includes: a video encoder that obtains a video feature representation of an input video, including: a plurality of NetVLAD networks, each NetVLAD network including a convolutional neural network (CNN) and a NetVLAD layer, a connector that receives the outputs of the plurality of NetVLAD networks, and a fully connected network that receives the output of the connector; a text encoder that obtains a text feature representation of an input text; and a similarity calculation unit that calculates the similarity between the video feature representation and the text feature representation for determining the matching of the video and the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video processing technology and the application of neural network in the field of video processing, and more specifically, to a method, device and medium for video retrieval based on neural network. The present invention is particularly suitable for online video retrieval. Background Art

[0002] Video has become one of the most commonly used media forms due to its ability to capture dynamic events and provide direct visual and sound experience. Currently, online videos occupy an increasingly large proportion of video applications. There are hundreds of millions of hours of videos (or short videos) on various online video platforms. If we cannot access these videos efficiently, these videos cannot be effectively used. Therefore, how to retrieve relevant videos through retrieval becomes the key.

[0003] It is obviously impossible to manually add reasonable titles and descriptions to millions of videos. Even if the creator adds titles and descriptions to each video when it is produced, such titles and descriptions may not fully summarize the video content for subsequent video retrieval. Therefore, a lot of research is currently focused on how to use neural networks for efficient video retrieval.

[0004] For video retrieval, there are currently two tasks: "title to video" and "video to title". "Title to video" refers to retrieval given a title form (for example, "how to build a house"), and the retrieval target is the video that the title can best describe (for example, a video explaining how to build a house). The "title" here should represent various texts associated with the video content, such as the video title, video description text, etc. The "video" here includes a collection of pictures collected over time (i.e., visual video) in a narrow sense, and includes visual video, audio, voice, subtitles (embedded or separate subtitle files), various audio tracks (embedded or separate audio track files), related covers (such as movie covers used in DVD discs), time tags, location tags, video clips (for example, video clips used in DVDs and Blu-ray discs), various information related to video clips (for example, covers, time tags, subtitles, content descriptions, etc. for video clips), etc., which can form the components of various existing video contents. Examples of online videos can be various short videos on YouTube, Douyin, Tiktok, and Bilibili.

[0005] For the "title-to-video" task, for each specific retrieval, it is achieved by given a set of "title-video" pairs and sorting all video candidates so that the videos most relevant to the title are ranked highest. On the other hand, the purpose of the "video-to-title" task is to find the title (retrieval target) that can best describe the retrieved video among a set of title candidates.

[0006] A common approach for the above two video retrievals is similarity learning, that is, how we learn a function that can best describe the similarity between two elements (i.e., the query and the candidate). Then, we can rank the candidates (videos or titles) according to the similarity (similarity estimation) between each candidate and the query.

[0007] Therefore, the current mainstream framework for video retrieval includes three parts: a video encoder, a text encoder, and similarity estimation. The video encoder obtains the video feature representation of the input video; the text encoder obtains the text feature representation of the input text (i.e., titles, video description texts, etc., texts associated with the video content); similarity calculation finds the matching videos and texts by calculating the similarity between the video feature representation and the text feature representation. In this way, similarity learning is split into the learning of the video encoder and the text encoder and the similarity estimation function.

[0008] For example, in similarity learning (i.e., the training phase), assume X represents the set of videos used for training, and Y represents the relevant titles of all videos (also called "texts" in this article). Given a learning database of B pairs of data {(v1,c1),…,(v i ,c i ),…,(v B ,c B )}, where v i ∈X, c i ∈Y, similarity learning is to find the video feature representation F v and the text feature representation F c , and find the matching videos and texts by comparing the similarity scores. The formula is as follows:

[0009] s = d(F v (v i ), F c (c j )) (1)

[0010] where d represents the learned similarity function (or distance function); s is the similarity estimation for matching and ranking. In a specific embodiment, d can adopt the cosine similarity function, that is, given two attribute vectors, A and B, their cosine similarity θ is given by the dot product and the vector length, as follows:

[0011]

[0012] A here i and B i represent the components of vectors A and B respectively.

[0013] Since text analysis is relatively simple and research on text retrieval has been carried out for decades, text editors for generating text feature representations have become relatively mature and efficient. For example, Radford et al. used a CLIP text encoder and an image encoder based on the Transformer architecture in "Learning transferable visual models from natural language supervision" arXiv preprint arXiv:2103.00020. The authors collected 400 million text-image data pairs where the image and text pairs are semantically matched. After obtaining features through the text encoder and image encoder, the similarity is calculated, and then trained using a contrastive loss function. Finally, the similarity between matching text and pictures is high, and that between non-matching ones is low. The CLIP text encoder trained as above is available online.

[0014] Current research mainly focuses on the design of video encoders. Due to the diversity of video content (as described above), how to enable video encoders to fully utilize various modalities in video content (visual video, audio, speech, captions, time tags, location tags, text, etc.) and associate various modalities to output video feature representations that can fully express video content information has become one of the cores of current research. An effective existing technology for the application of multiple modalities in video content is the multi-modal transformer (MMT) proposed by Gabeur, V. et al. of Google based on the Transformer technology ("Multi-modal transformer for video retrieval," In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 (pp. 214-229). Springer International Publishing). It uses N pre-trained experts F n , n = 1…N to extract a sequence containing K features from the video, and obtain the cumulative embedding for the feature sequence of each expert, so as to be able to learn effective representations from different modalities in the video; to process cross-modal information, learn N embeddings En , where \(n = 1,\ldots,N\) is used to distinguish between different expert embeddings; finally, a temporal embedding \(T\) is provided; for the feature \(F(v)\), the expert embedding \(E(v)\) and the temporal embedding \(T(v)\), the video embedding \(\Omega(v)=F(v)+E(v)+T(v)\) is obtained. The video embedding is input into the MMT transformer to obtain the video representation. The transformer framework used by MMT is the attention-based transformer architecture proposed by Vaswani, A. et al. (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I., "Attention is all you need" In Advances in neural information processing systems (pp. 5998 - 6008)).

[0015] In the video retrieval frameworks in the prior art, especially the design of video encoders, there is still a great need for improvement, especially in terms of how to handle multiple modalities of videos. For example, the MMT model has the problem of overfitting, mainly because a single video annotation dataset contains too little data to effectively train the transformer structure. There are two research directions to solve this problem. One is to directly increase the amount of data. The other way is to reduce the number of parameters that the model needs to train. Summary of the Invention

[0016] To solve the above technical problems, this paper proposes a novel video encoder for video retrieval.

[0017] In one aspect, considering the overfitting phenomenon in the MMT model, this paper proposes to use NetVLAD in the video encoder to replace the transformer in MMT to fuse video features.

[0018] In another aspect, aiming at the overfitting phenomenon in the MMT model, this paper proposes that a general text encoder can be trained with a large dataset on other tasks and then introduced into the video retrieval model to reduce the number of parameters for model training. As mentioned above, CLIP aims to learn visual representations under the supervision of text. It is trained with a dataset of 400 million pairs of text - images and has achieved excellent performance. If a picture is regarded as a video with 1 frame, then the CLIP model can also be considered as learning text representations under the supervision of videos during training. Therefore, this paper proposes to introduce the text encoder of CLIP into the video retrieval model to reduce the number of parameters that the model needs to train.

[0019] When using a pre-trained CLIP model for the text encoder, there is a problem of how to adjust it to make it suitable for video retrieval tasks. One method is to fine-tune the model during training. However, since the initial parameters of the video encoder are randomly initialized, when fine-tuning the CLIP model, these randomly initialized parameters will greatly affect the parameters trained by CLIP, thus weakening the performance of the text encoder.

[0020] In a further aspect, this paper proposes a two-step training method for the training of the entire video retrieval framework (i.e., similarity learning) in the case of using a pre-trained CLIP model for the text encoder. Since the parameters of the text encoder are fixed, the semantic embedding representation obtained in the first step may not be suitable for the video retrieval task. Therefore, in the second step of training, all model parameters will be fine-tuned, so that an embedding representation suitable for the video-text task can be obtained.

[0021] According to one aspect of the present invention, a video retrieval device is proposed, which may include:

[0022] A video encoder, which can receive an input video and obtain a video feature representation of the input video. The video encoder includes:

[0023] Multiple NetVLAD networks, each NetVLAD network including a convolutional neural network (CNN) and a NetVLAD layer. The CNN is used to extract a specific modality among multiple modalities in the input video, and the NetVLAD layer is used to fuse multiple features in the corresponding modality provided by the corresponding CNN.

[0024] A connector, which receives the outputs of the multiple NetVLAD networks.

[0025] A fully connected network, which receives the output of the connector, and

[0026] A text encoder, which obtains a text feature representation of the input text; and

[0027] A similarity calculation unit, which calculates the similarity between the video feature representation and the text feature representation for determining the match between the video and the text.

[0028] According to a further aspect, the fully connected network is a two-layer fully connected network.

[0029] According to a further aspect, the text encoder adopts a CLIP text encoder.

[0030] According to a further aspect, cosine similarity is used to calculate the similarity.

[0031] According to a further aspect, the fully-connected network outputs a video feature representation of the input video, and the video feature representation has a reduced dimension.

[0032] According to a further aspect, the video encoder includes: a gating module that receives the output of the fully-connected network and outputs a video feature representation of the input video, where the gating module is used to perform non-linear interaction between multiple dimensions of the features from the fully-connected network, use a self-gating mechanism to reactivate different features, and perform L2 normalization.

[0033] According to another aspect of the present invention, there is provided a method for retrieving a video using the video retrieval device, including:

[0034] Using a video encoder to obtain a video feature representation of an input video, the video encoder includes: a plurality of NetVLAD networks, each NetVLAD network including a convolutional neural network (CNN) and a NetVLAD layer, where the CNN is used to extract a specific modality among multiple modalities in the input video, and the NetVLAD layer is used to fuse multiple features in the corresponding modality provided by the corresponding CNN; a connector that receives the outputs of the plurality of NetVLAD networks; a fully-connected network that receives the output of the connector;

[0035] Using a text encoder to obtain a text feature representation of the input text;

[0036] Calculating the similarity between the video feature representation and the text feature representation.

[0037] According to another aspect of the present invention, the video retrieval device is a software module implemented by code, and when the code is executed, it implements the video retrieval device or executes the method for retrieving a video.

[0038] According to another aspect of the present invention, the video retrieval device is a hardware module implemented by a processor dedicated to neural networks.

[0039] According to another aspect of the present invention, the video retrieval device is a software and hardware architecture implemented by executable code in combination with a hardware processor dedicated to neural networks, where the hardware processor dedicated to neural networks can implement a part of the architecture of the video retrieval device in the form of a hardware functional module, such as a fully-connected network, a convolutional neural network (CNN), etc.

[0040] According to another aspect of the present invention, there is provided a two-step training method for training a video retrieval device, including:

[0041] In the first step, freeze the parameters of the text encoder and train only the parameters of the video encoder using the training set; and

[0042] In the second step, use the training set to fine-tune the parameters of the text encoder and the parameters of the video encoder.

[0043] According to a further aspect, the text encoder employs a pre-trained text encoder, and the video encoder is randomly initialized before training.

[0044] According to a further aspect, the video retrieval framework includes:

[0045] The video encoder, which obtains a video feature representation of the input video;

[0046] The text encoder, which obtains a text feature representation of the input text;

[0047] A similarity calculation unit, which calculates the similarity between the video feature representation and the text feature representation. Description of the Drawings

[0048] Figure 1 Shows a schematic diagram of a video retrieval framework according to an embodiment of the present invention.

[0049] Figure 2 Shows a schematic diagram of the structure of the NetVLAD network according to an embodiment of the present invention.

[0050] Figure 3 Shows a schematic diagram of a two-step training method according to an embodiment of the present invention.

[0051] Figure 4 Shows a schematic diagram of a video retrieval device for implementing an embodiment of the present invention. Detailed Description of the Embodiment

[0052] Now, various solutions will be described with reference to the drawings. In the following description, for the purpose of explanation, a number of specific details are set forth in order to provide a thorough understanding of one or more solutions. However, it is obvious that these solutions can also be implemented without these specific details.

[0053] As used in this application, "devices", "frameworks", "encoders", "modules", "units", etc. related to artificial intelligence herein are intended to refer to computer-related entities, such as, but not limited to, hardware, software executed by a general or special-purpose processor, firmware, software, or any combination thereof. For example, these "devices", "frameworks", "encoders", "modules", "units" can be, but are not limited to: processes running on a processor, processors, objects, executables, execution threads, programs, and / or computers. For example, an application running on a computing device and the computing device itself can both be these "devices", "frameworks", "encoders", "modules", "units". One or more of these "devices", "frameworks", "encoders", "modules", "units" can be located within an execution process and / or an execution thread, and these "devices", "frameworks", "encoders", "modules", "units" can be located on one computer and / or distributed across two or more computers. Additionally, these "devices", "frameworks", "encoders", "modules", "units" can execute from various computer-readable media having various data structures stored thereon. Components can communicate by means of local and / or remote processes, such as in accordance with signals having one or more data packets, e.g., data from a component that interacts with another component in a local system, a distributed system by means of signals and / or interacts with other systems over a network such as the Internet by means of signals.

[0054] When implemented in hardware, these "devices", "frameworks", "encoders", "modules", "units" can be implemented or executed using a general-purpose processor, a neural network processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but alternatively, the processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors and a DSP core, or any other such configuration. Additionally, at least one processor can include one or more modules operable to execute one or more of the above steps and / or operations. A neural network processor can be used to implement various basic modules of a neural network, e.g., fully connected networks, convolutional neural networks (CNNs), etc.

[0055] When implemented in hardware, these "devices", "frameworks", "encoders", "modules", "units" are implemented on a system-on-chip (SOC).

[0056] When implementing these "devices", "frameworks", "encoders", "modules", "units" using hardware circuits such as ASICs, FPGAs, etc., they may include various circuit blocks configured to perform various functions. Those skilled in the art can design and implement these circuits in various ways according to various constraints imposed on the entire system to achieve the various functions disclosed in the present invention.

[0057] When implemented in software, these "devices", "frameworks", "encoders", "modules", "units" may be processors or computer-executable code, which can be stored on a computer or processor-readable storage medium, or stored in the cloud (server cluster) of a network, and when the code is executed, it can implement these "devices", "frameworks", "encoders", "modules", "units", or implement the methods using these "devices", "frameworks", "encoders", "modules", "units".

[0058] The present invention relates to video retrieval based on neural networks. The video retrieval framework herein can be applied to two tasks: "title-to-video" and "video-to-title". In these two tasks, the general difference lies in how to form (video, title) pairs. In the "title-to-video" task, the specific retrieval is the input title, and the retrieval target is the relevant video. Therefore, the input to the video retrieval framework is multiple (video candidate, title) pairs; while in the "video-to-title" task, the specific retrieval is the input video, and the retrieval target is the relevant title or relevant text description. Therefore, the input to the video retrieval framework is multiple (video, title candidate) pairs.

[0059] The "video retrieval framework", "video retrieval device", "video retrieval module", "video retrieval unit" mentioned herein all refer to any video retrieval function that can be implemented as software, hardware, or a combination of software and hardware, and these terms can be used interchangeably herein. Additionally, "similarity estimation", "similarity calculation", "similarity determination" are used interchangeably, representing the calculation of two inputs (such as video feature representation and "text feature representation") implemented by hardware, software, or a combination of hardware and software. Additionally, the terms "training" and "learning" are generally used interchangeably herein.

[0060] As described above, the present invention is mainly directed to the technical problem of overfitting in the MMT ("Multi-modal transformer for video retrieval," In ComputerVision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 (pp. 214-229). Springer International Publishing) proposed by Gabeur, V. et al. of Google based on the Transformer technology. The MMT framework is an encoder for video retrieval, which is proposed based on the attention-based Transformer architecture proposed by Vaswani, A. et al. of Google (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I., "Attention is all you need" In Advances in neuralinformation processing systems (pp. 5998-6008)). Therefore, these two papers and multiple documents cited therein are incorporated herein by reference as part of the present disclosure to make the disclosure herein complete.

[0061] According to one aspect of the present invention, a novel video retrieval framework is proposed.

[0062] According to another aspect of the present invention, a two-step training method for training the novel video retrieval framework is proposed.

[0063] Figure 1 A schematic diagram of a video retrieval framework according to an embodiment of the present invention is shown. As shown in the figure, the video retrieval framework herein includes a video encoder, a text encoder, and a similarity estimation unit.

[0064] The video encoder receives a video input, passes the video input through a neural network, and obtains a video feature representation of the video input at the output end.

[0065] In a specific embodiment, the video encoder includes: a plurality of NetVLAD networks, each NetVLAD network including a convolutional neural network (CNN) and a NetVLAD layer; a connector that receives the outputs of the plurality of NetVLAD networks and connects them; and a fully-connected network that receives the output of the connector, i.e., the connection of the outputs of the plurality of NetVLAD networks.

[0066] Figure 2 FIG. shows a schematic structural diagram of a NetVLAD network according to an embodiment of the present invention. A NetVLAD network consists of a convolutional neural network (CNN) and a NetVLAD layer. The output of the CNN network is an HxWxD map. The CNN is used to extract multiple features of a specific modality among multiple modalities in the input video, while the NetVLAD layer is used to fuse the multiple features of a specific modality.

[0067] Specifically, for video v, we can use N specific models (i.e., CNNs) to extract N multi-modalities {I n} n∈1...N from a video, and then N NetVLADs will be used to fuse all the features of the N multi-modalities in the time dimension, and then obtain the embedded representation of each multi-modality

[0068]

[0069] In the video encoder, the connector receives the outputs of the plurality of NetVLAD networks and connects them. The dimension of each multi-modal embedded representation is d n , and then all the multi-modal representations are connected to form a vector I c , and the dimension of I c is:

[0070]

[0071] In the video encoder, the fully-connected network receives the output of the connector, i.e., the connection of the outputs of the plurality of NetVLAD networks. The fully-connected network is used to obtain the final video feature representation. The input of the fully-connected network is I c , and the dimension is d c . In a preferred embodiment, the fully-connected network can be a two-layer fully-connected network. One achievable function of the fully-connected network is to reduce the dimension of the final video feature representation.

[0072] In a preferred embodiment, a gating module is added after the fully-connected network, and the output Z of this module is regarded as the representation of the video:

[0073] Z = f(FC(I c )) (5)

[0074] The gating module is used to implement three functions:

[0075] First, perform non-linear interactions between multiple dimensions of the input features;

[0076] Second, recalibrate different activations of the input features through a certain self-gating mechanism;

[0077] Third, perform L2 normalization.

[0078] In one embodiment, the gating module Z = f(Z0) can be implemented as follows:

[0079] Z1 = W1Z0 + b1,

[0080]

[0081]

[0082] where the learnable parameters W1, W2, b1, and b2 are as follows:

[0083]

[0084] The calculation of Z1 realizes spatial mapping. The calculation of Z2 realizes context gating, where each dimension of Z1 is re-weighted using the learned gating weight σ(W2Z1 + b2) with values between 0 and 1. Finally, L2 normalization is performed.

[0085] The video feature representation of the video input directly output by the fully connected network or output via the gating module is provided to the similarity estimation unit for similarity calculation with the text feature representation.

[0086] In a specific embodiment, the text encoder receives the text input and obtains the text feature representation of the text input. In the "title-to-video" task, the text input can be a query, while in the "video-to-title", the text input can come from the candidate titles in a pre-set candidate title library.

[0087] As described above, since text analysis is relatively simple and research on text retrieval has been carried out for decades, text editors for generating text feature representations have been relatively mature and efficient. For example, Radford et al. used a CLIP text encoder and an image encoder based on the Transformer architecture in "Learning transferable visual models from natural language supervision" arXiv preprint arXiv: 2103.00020. The authors collected 400 million pairs of text-image data pairs, where the image and text pairs are semantically matched. After obtaining the features through the text encoder and the image encoder, the similarity is calculated, and then the contrast loss function is used for training. Finally, the similarity between the matching text and pictures is high, and the similarity between the non-matching ones is low. The CLIP text encoder trained as described above is available online. A specific embodiment of the text encoder may adopt the CLIP text encoder of Radford et al. Therefore, these two papers and multiple documents cited therein are incorporated herein by reference as part of the present disclosure to make the disclosure herein complete. Of course, the present invention is not limited thereto, and any text encoder can be used. The text encoder used may be a text encoder initialized during the training of the present video retrieval framework, or a text encoder that has been pre-trained with a large amount of training data. When using a text encoder that has been pre-trained with a large amount of training data, a two-step training method for training the video retrieval framework according to one aspect of the present invention can be used to obtain the best text encoder and video retrieval framework.

[0088] The NetVLAD network is a network for place recognition based on the CNN architecture proposed by Arandjelovic, R. et al. (Arandjelovic, R., Gronat, P., Torii, A., Pajdla, T., & Sivic, J., "NetVLAD: CNN architecture for weakly supervised place recognition," In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5297-5307)). However, Arandjelovic, R. et al. only used a single NetVLAD network to process images for place recognition and did not use it to process each modality in the multi-modalities for videos. Therefore, this paper and multiple cited documents are incorporated herein by reference as part of the present disclosure to make the disclosure herein complete.

[0089] When using a pre-trained CLIP model for the text encoder, there is a problem of how to adjust it to suit the video retrieval task. One method to solve this problem is to fine-tune the model during training. However, since the initial parameters of the video encoder are randomly initialized, when fine-tuning the CLIP model, these randomly initialized parameters will greatly affect the parameters trained by CLIP, thus weakening the performance of the text encoder.

[0090] In a further aspect, this paper proposes a two-step training method for the training (i.e., similarity learning) of the entire video retrieval framework when using a pre-trained CLIP model for the text encoder.

[0091] Figure 3 A schematic diagram of the two-step training method according to an embodiment of the present invention is shown. Combining Figure 1 and Figure 3 , in the first stage, the parameters of the text encoder are fixed (i.e., frozen), and only the parameters of the video encoder are updated. After such training, the semantic embedding representations obtained by the video encoder and the text encoder are similar. Specifically, when the video encoder and the text encoder are output for image-text matching, their similarity is very high, so it is said that the representations are similar. This step is to train the video encoder using the training set so that it learns knowledge to a certain extent. Because updating some parameters can reduce the loss, and the reduction of the loss indicates that the retrieval performance of the model is improved, which also means that when performing video-text matching, the outputs of the text encoder and the video encoder are highly similar, that is, the embedded representations of the two outputs are considered similar.

[0092] Since the parameters of the text encoder are fixed, the semantic embedding representation obtained in the first step may not be suitable for the video retrieval task. Therefore, in the second-stage training, all model parameters will be fine-tuned so that we can obtain an embedding representation suitable for the video text task.

[0093] The principle of the two-step training method of the present invention is that after the adjustment in the first stage, the outputs of the video encoder and the text encoder are already relatively similar to a certain extent. Thus, the training in the second stage will only cause fine-tuning of the two, rather than significant changes to the text encoder.

[0094] In a specific embodiment of similarity learning, assume that X represents the set of videos for training, and Y represents the relevant captions of all videos (also referred to as "text" in this article). Given a learning database of B pairs of data {(v1, c1), …, (v i , c i ), …, (v B , c B )}, where v i ∈X and c i ∈Y, similarity learning is to find the video feature representation F v and the text feature representation F c , and find the matching videos and texts by comparing the similarity scores.

[0095] For the B pairs of data {(v1, c1), …, (v i , c i ), …, (v B , c B )}, through the video encoder F v and the text encoder F c , calculate the representation {(Z1, T1), …, (Z i , T i ), …, (Z B , T B )} (Z and T are the video feature representation and the text feature representation respectively). Then a similarity matrix of dimension B×B can be obtained:

[0096] s i,j = d(Z i , T j ) (8)

[0097] This matrix contains all the distances d between Z and T, and the elements on the diagonal of the matrix are the similarities of the matching texts and videos.

[0098] Use the following contrastive learning loss function to optimize the encoder:

[0099]

[0100]

[0101]

[0102] Where τ is a pre-selected temperature hyperparameter used to adjust the influence of the learning process on the encoder.

[0103] In other words, the elements of the similarity matrix are the similarity values of each text and video, and the elements on the diagonal are the similarities of the matching text and video. When each element value on the diagonal is the largest in its row and column, a completely correct video retrieval can be achieved. The contrastive learning loss function updates the model parameters by calculating the gradient and backpropagation to achieve the above-mentioned purpose. During training, we will observe the calculated loss value. When the loss is very small and no longer decreasing, it is considered that the training is completed.

[0104] The above specific embodiments of similarity learning are only one specific embodiment that can implement the two-step training method of the present invention. It is easy for those skilled in the art to understand that other examples can be used to practice the two-step training method of the present invention. For example, a loss function different from the above can be used.

[0105] Figure 4 The schematic diagram of a video retrieval device for implementing an embodiment of the present invention is shown. The device includes a processor and a memory. As described above, the processor can be a general-purpose processor, a neural network processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any circuit module designed to execute the functions described herein. The processor includes various processors located in a remote computer or server or server cluster or cloud.

[0106] The memory can be any type of memory capable of storing programs, codes, and data to be operated on. The memory includes various storages located in a remote computer or server or server cluster or cloud. The memory contains instructions that, when executed by the processor, can implement the methods described herein or implement the devices described herein. The memory can also store the data manipulated by the methods and devices described herein, such as video input, text input, intermediate data, various model parameters, neural network parameters, and so on.

[0107] The device can also be implemented in a system-on-chip (SOC).

[0108] The above aspects of the present invention can be summarized into the following non-limiting aspects. It is easy to understand that these aspects can be combined, deleted, and replaced in any form.

[0109] Aspect 1. A method for retrieving a video, comprising:

[0110] Using a video encoder to obtain a video feature representation of an input video, the video encoder comprising:

[0111] A plurality of NetVLAD networks, each NetVLAD network comprising a convolutional neural network (CNN) and a NetVLAD layer, the CNN being used to extract a specific modality among multiple modalities in the input video, and the NetVLAD layer being used to fuse multiple features in the corresponding modality provided by the corresponding CNN,

[0112] A connector that receives the outputs of the plurality of NetVLAD networks,

[0113] A fully connected network that receives the output of the connector;

[0114] Using a text encoder to obtain a text feature representation of an input text;

[0115] Calculating a similarity between the video feature representation and the text feature representation.

[0116] Aspect 2. A video retrieval device, comprising:

[0117] A video encoder that obtains a video feature representation of an input video, comprising:

[0118] A plurality of NetVLAD networks, each NetVLAD network comprising a convolutional neural network (CNN) and a NetVLAD layer,

[0119] A connector that receives the outputs of the plurality of NetVLAD networks,

[0120] A fully connected network that receives the output of the connector;

[0121] A text encoder that obtains a text feature representation of an input text;

[0122] A similarity calculation unit that calculates a similarity between the video feature representation and the text feature representation for determining a match between the video and the text.

[0123] Aspect 3. The method or device according to Aspect 1 and 2, wherein the fully connected network is a two-layer fully connected network.

[0124] Aspect 4. The method or device according to Aspect 1-3, wherein the text encoder employs a CLIP text encoder.

[0125] Aspect 5. The method or device according to Aspect 1-4, wherein a cosine similarity is used to calculate the similarity.

[0126] Aspect 6. The method or device according to Aspects 1-5, wherein the video encoder further comprises:

[0127] A gating module that receives the output of the fully connected network and outputs the video feature representation, wherein the gating module is configured to:

[0128] Perform a non-linear interaction between multiple dimensions of the features in the output from the fully connected network,

[0129] Use a self-gating mechanism to reactivate different features, and

[0130] Perform L2 normalization.

[0131] Aspect 7. A computer-readable storage medium storing code for performing video retrieval, which when executed can implement the method or device according to Aspects 1-5.

[0132] Aspect 8. A two-step training method for training a video retrieval framework, comprising:

[0133] In the first step, freeze the parameters of the text encoder and train only the parameters of the video encoder using the training set; and

[0134] In the second step, use the training set to fine-tune the parameters of the text encoder and the parameters of the video encoder.

[0135] Aspect 9. The method according to Aspect 8, wherein the text encoder uses a pre-trained text encoder, and the video encoder is randomly initialized before training.

[0136] Aspect 10. The method according to Aspects 8 and 9, wherein the video retrieval framework comprises:

[0137] The video encoder that obtains the video feature representation of the input video;

[0138] The text encoder that obtains the text feature representation of the input text; and

[0139] A similarity calculation unit that calculates the similarity between the video feature representation and the text feature representation.

[0140] Aspect 11. The method according to Aspects 8-10, wherein the text encoder comprises a pre-trained text encoder.

[0141] Aspect 12. The method according to Aspects 8-11, wherein the video encoder comprises:

[0142] Multiple NetVLAD networks, each NetVLAD network including a Convolutional Neural Network (CNN) and a NetVLAD layer,

[0143] A connector that receives the outputs of the multiple NetVLAD networks,

[0144] A fully connected network that receives the output of the connector.

[0145] Aspect 13. The method according to aspects 8 - 12, wherein the video encoder further comprises:

[0146] A gating module that receives the output of the fully connected network and outputs the video feature representation, wherein the gating module is configured to:

[0147] Perform non - linear interaction between multiple dimensions of the features in the output from the fully connected network,

[0148] Use a self - gating mechanism to re - weight different excitations of the features, and

[0149] Perform L2 normalization.

[0150] Aspect 14. The method according to aspects 12 - 13, wherein the fully connected network is a two - layer fully connected network.

[0151] Aspect 15. The method according to aspects 10 - 14, wherein the text encoder employs a CLIP text encoder.

[0152] Aspect 16. The method according to aspects 10 - 14, wherein cosine similarity is used to calculate the similarity.

[0153] Aspect 17. A computer - readable storage medium that stores code for training for video retrieval, the code when executed being capable of implementing the method according to aspects 8 - 16.

[0154] Aspect 18. The device according to aspects 2 - 6, the device being trained using the method according to aspects 8 and 11.

[0155] Although the foregoing disclosure discusses exemplary scenarios and / or embodiments, it should be noted that many changes and modifications can be made herein without departing from the scope of the described scenarios and / or embodiments as defined by the claims. Moreover, although the elements of the described scenarios and / or embodiments are described or claimed in the singular, the plural case can also be contemplated unless explicitly stated to be limited to the singular. Additionally, all or part of any scenario and / or embodiment can be used in combination with all or part of any other scenario and / or embodiment, unless otherwise indicated.

Claims

1. A method for retrieving videos, comprising: Using a video encoder to obtain a video feature representation of an input video, the video encoder comprising: A plurality of NetVLAD networks, each NetVLAD network comprising a convolutional neural network (CNN) and a NetVLAD layer, the CNN being used to extract a specific modality among multiple modalities in the input video, and the NetVLAD layer being used to fuse multiple features in the corresponding modality provided by the corresponding CNN, A connector that receives the outputs of the plurality of NetVLAD networks, A fully-connected network that receives the output of the connector; Using a text encoder to obtain a text feature representation of the input text, wherein the text encoder employs a CLIP text encoder; Calculating the similarity between the video feature representation and the text feature representation, Wherein, the video retrieval framework is trained using the following two-step training method: In the first step, freeze the parameters of the text encoder and use the training set to train only the parameters of the video encoder; and In the second step, use the training set to fine-tune the parameters of the text encoder and the parameters of the video encoder.

2. The method according to claim 1, wherein, The fully-connected network is a two-layer fully-connected network.

3. The method according to claim 1, wherein, Cosine similarity is used to calculate the similarity.

4. The method according to claim 1, wherein, The video encoder further comprises: A gating module that receives the output of the fully-connected network and outputs the video feature representation, wherein the gating module is used to: Perform a non-linear interaction between multiple dimensions of the features in the output from the fully-connected network, Use a self-gating mechanism to reactivate different features, and Perform L2 normalization.

5. The method according to claim 1, wherein, The text encoder employs a pre-trained text encoder, and the video encoder is randomly initialized before training.

6. A video retrieval system, comprising: A video encoder that obtains a video feature representation of an input video, comprising: A plurality of NetVLAD networks, each NetVLAD network comprising a convolutional neural network (CNN) and a NetVLAD layer, A connector that receives the outputs of the plurality of NetVLAD networks, A fully-connected network that receives the output of the connector; A text encoder that obtains a text feature representation of the input text, wherein the text encoder employs a CLIP text encoder; A similarity calculation unit that calculates the similarity between the video feature representation and the text feature representation for determining the match between the video and the text, Wherein, the video retrieval framework is trained using the following two-step training method: In the first step, freeze the parameters of the text encoder and use the training set to train only the parameters of the video encoder; and In the second step, use the training set to fine-tune the parameters of the text encoder and the parameters of the video encoder.

7. The system according to claim 6, wherein, The fully-connected network is a two-layer fully-connected network.

8. The system according to claim 6, wherein, Cosine similarity is used to calculate the similarity.

9. The system according to claim 6, wherein, The video encoder further comprises: A gating module that receives the output of the fully-connected network and outputs the video feature representation, wherein the gating module is configured to: perform non-linear interaction between multiple dimensions of features in the output from the fully-connected network, use a self-gating mechanism to reactivate different features, and perform L2 normalization.

10. The system according to claim 6, wherein, The text encoder employs a pre-trained text encoder, and the video encoder is randomly initialized before training.

11. A computer-readable storage medium that stores code for performing video retrieval, which when executed can implement the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Cross-modal image text retrieval method of hybrid fusion model

    CN112784092A

  • Cross-modal retrieval model based on strong representation deep hash

    CN113641846A