Video text retrieval model training method and device, equipment and storage medium
By performing feature interaction and aggregation processing on the video text retrieval model, and combining multi-head and self-attention mechanisms for training, the inaccurate retrieval problem caused by similar videos and single-language text is solved, thus improving the accuracy of the video text retrieval model.
Patent Information
- Application Number
- CN202311369350.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-10-20
AI Technical Summary
In existing technologies, due to the influence of similar videos and monolingual text, video text retrieval models have poor retrieval performance and inaccurate feature extraction.
By acquiring training samples, including original videos and second text data obtained by language conversion based on first text data, feature extraction is performed. Then, a multi-head attention mechanism is used for feature interaction and aggregation, contextual information is added, and a self-attention mechanism is used to reconstruct and train video features and text features, thereby improving the accuracy of the video text retrieval model.
Adding contextual information during training helps to uncover the connections between video frames, reduce interference from polysemous words and expression habits, and improve the retrieval accuracy of the video text retrieval model.
Smart Images

Figure CN119862302B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of learning and video retrieval, and particularly relates to a video text retrieval model training method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rise and popularity of various video platforms, the number of videos on the network also presents an explosive growth, and video text retrieval has become a new demand for people to efficiently find relevant videos. Among them, the video text retrieval task aims to automatically mine the potential semantic information of the corresponding modal according to the given text or video query value, and then retrieve the single or multiple videos or texts with the highest similarity in the data set.
[0003] When performing video text retrieval, a feature extraction module is usually used to extract features of videos and texts respectively, and then a similarity function is designed to calculate the similarity of the features of videos and texts. However, in actual application, there may be similar videos between different videos, which are similar in picture composition or plot, resulting in inaccurate extraction of video features. At the same time, single language text may not be accurate in feature extraction due to the existence of polysemous words or differences in expression habits, thereby causing poor retrieval matching effect of the video text retrieval model. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a video text retrieval model training method, device, electronic device and storage medium, to solve the technical problem of poor video text retrieval effect caused by similar videos and single language text in related technologies.
[0005] In a first aspect, the embodiments of the present application provide a video text retrieval model training method, comprising:
[0006] Obtaining a training sample, wherein the training sample comprises an original video, first text data and second text data, and the second text data is obtained by language conversion based on the first text data;
[0007] Performing feature extraction on the training sample to obtain initial video features, first text features and second text features;
[0008] Performing feature interaction and aggregation processing on the initial video features to obtain video features carrying context information;
[0009] Training a video text retrieval model according to the first text features, the second text features and the video features carrying context information, and obtaining a trained video text retrieval model when it is determined that the training is completed.
[0010] In a second aspect, an embodiment of the present application provides a device for training a video-text retrieval model, comprising:
[0011] a data obtaining module configured to obtain training samples, wherein the training samples comprise an original video, first text data, and second text data, and the second text data is obtained by language conversion based on the first text data;
[0012] a feature extraction module configured to perform feature extraction on the training samples to obtain initial video features, first text features, and second text features;
[0013] a feature processing module configured to perform feature interaction and aggregation processing on the initial video features to obtain video features carrying context information;
[0014] a model training module configured to train a video-text retrieval model according to the first text features, the second text features, and the video features carrying context information, and obtain a trained video-text retrieval model when it is determined that the training is completed.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, which comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and the processor implements the steps in the method for training a video-text retrieval model according to any one of the above aspects when executing the computer program.
[0016] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps in the method for training a video-text retrieval model according to any one of the above aspects when executed by a processor.
[0017] The embodiment of the present application provides a video text retrieval model training method and device, an electronic device and a storage medium. When a constructed video text retrieval model is trained, a training sample is acquired, wherein the training sample comprises an original video, first text data corresponding to the original video and second text data obtained by performing language conversion on the first text data. Then, feature extraction is performed on the training sample to obtain corresponding initial video features, first text features and second text features. Meanwhile, feature interaction and aggregation processing are performed on the initial video features to obtain video features carrying context information. Finally, the video text retrieval model is trained according to the first text features, the second text features and the video features, so as to obtain a trained video text retrieval model when the training is completed. The context information is added to the video features in the training process, the association between video frames is mined, the accuracy of video feature extraction is improved, meanwhile, the original text data is subjected to language conversion, the semantic alignment is performed by using multilingual text, the interference of polysemous words and expression habits is reduced, the accuracy of text features is improved, and the retrieval accuracy of the video text retrieval model is further improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 FIG. 1 is a flowchart of a video text retrieval model training method provided by the embodiment of the present application;
[0019] Figure 2 FIG. 2 is a flowchart of a step of obtaining video features carrying context information provided by the embodiment of the present application;
[0020] Figure 3 FIG. 3 is a flowchart of a step of training a video text retrieval model provided by the embodiment of the present application;
[0021] Figure 4 FIG. 4 is a block diagram of reconstructing video features provided by the embodiment of the present application;
[0022] Figure 5 FIG. 5 is a block diagram of training a video text retrieval model provided by the embodiment of the present application;
[0023] Figure 6 FIG. 6 is a structural diagram of a video text retrieval model training device provided by the embodiment of the present application;
[0024] Figure 7 FIG. 7 is a structural diagram of an electronic device provided by the embodiment of the present application;
[0025] Figure 8 FIG. 8 is another structural diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0026] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of the present application.
[0027] It should be understood that each step described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0028] The term "comprising" and variations thereof as used herein are open-ended, that is "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions are given throughout the description.
[0029] In the related art, when training a video text retrieval model, features of the video and the text are extracted, and then a similarity function is designed to calculate the similarity of the features of the video and the text. However, due to the existence of video frames with similar pictures, the feature extraction is not accurate when extracting the features of the video. At the same time, for a single language, there are ambiguous words or different expression habits, which leads to inaccurate extraction of text features, and further leads to low accuracy of the use of the learned and optimized video text retrieval model.
[0030] To solve the technical problems in the related art, the present embodiment provides a training method of a video text retrieval model, please refer to Figure 1 , Figure 1 is a flowchart of a training method of a video text retrieval model provided by the present embodiment, which includes steps 101 to 104.
[0031] Step 101, obtaining a training sample, wherein the training sample includes an original video, first text data and second text data, and the second text data is obtained by language conversion based on the first text data.
[0032] In an embodiment, before training, a training sample for training is obtained, and the obtained training sample includes an original video, first text data and second text data, and the second text data is obtained by language conversion based on the first text data.
[0033] Exemplarily, when the training sample is acquired, the corresponding video data and the text data corresponding to the video data are acquired first, and in order to avoid the influence of polysemy or expression habit in a single language, the first text data is subjected to language conversion at this time to obtain text data in another language. Generally, when processing, the first text data is directly translated or the like to obtain the second text data.
[0034] In step 102, feature extraction is performed on the training sample to obtain initial video features, first text features and second text features.
[0035] In an embodiment, after the training sample is acquired, the training sample is input into the video-text retrieval model for training, and then feature extraction is performed on the training sample. Specifically, feature extraction is performed on the original data, the first text data and the second text data respectively to obtain the initial video features, the first text features and the second text features.
[0036] Further, when feature extraction is performed on the original video in the training sample, the feature extraction includes: video frame sampling is performed on the original video to obtain video key frames used for feature extraction, and feature extraction is performed on the video key frames to obtain initial video features of the original video; feature extraction is performed on the first text data to obtain first text features, and feature extraction is performed on the second text data to obtain second text features.
[0037] Exemplarily, when feature extraction is performed on the original video, the original video can be processed by using CLIP (Contrastive Language-Image Pre-training) to obtain video features. Specifically, when performing feature extraction, video frame sampling is first performed on the original video to obtain a plurality of video frames used for feature extraction, and then a ViT-B / 32 image encoder in the CLIP is used to perform feature extraction on the sampled video frames to obtain features of each video frame, which constitute the initial video features of the original video. For example, when video frame sampling is performed, 12 video frames can be randomly sampled, and then feature extraction is performed on the 12 video frames.
[0038] In addition, when feature extraction is performed on the first text data and the second text data, the feature extraction includes: the first text data is subjected to word segmentation processing to obtain a first text sequence; the first text sequence is filled according to position information corresponding to the first text sequence; and the filled first text sequence is subjected to encoding processing to obtain first text features of the first text data.
[0039] Specifically, when performing feature extraction on the first text data, the first text data is subjected to word segmentation processing to obtain a first text sequence corresponding to the first text data, and then the first text sequence is subjected to padding processing according to position information of each word in the first text sequence, and then the first text sequence after the padding processing and reverse encoding processing are performed to obtain the first text feature.
[0040] For example, when processing the first text data, the word segmentation processing can be performed on the sentences in the first text data using the byte pair encoding (BPE) method, and then the first text sequence is obtained according to the result of the word segmentation processing. For the first text data, after the word segmentation processing, each word corresponds to a position in the first text sequence, so the position information corresponding to the first text sequence can be obtained at this time, and then the position information is fused with the first text sequence, and an [EOS] vector is embedded at the tail of the processed first text sequence as an end symbol, and finally the encoding processing is performed to obtain the corresponding first text feature.
[0041] Further, when processing the second text data, the following steps are included: performing word segmentation processing on the second text data to obtain a second text sequence; performing padding on the second text sequence according to position information corresponding to the second text sequence; and performing encoding processing on the padded second text sequence to obtain a second text feature of the second text data.
[0042] Since the feature extraction method for the second text data is the same as the feature extraction method for the first text data, detailed description is not given, and specific embodiments of the feature extraction method for the first text data can be referred to.
[0043] In step 103, the initial video feature is subjected to feature interaction and aggregation processing to obtain a video feature carrying context information.
[0044] In order to improve the accuracy of feature extraction of the video, after the feature extraction of the original video is completed, the video feature carrying context information is obtained by performing corresponding processing, and then the video text retrieval model is trained using the video feature carrying context information.
[0045] In an embodiment, when the video feature carrying context information is obtained, the initial video feature is subjected to feature interaction and aggregation processing, and then the video feature carrying context information is obtained after the related processing is completed.
[0046] When the initial video feature is processed to obtain the video feature carrying context information, the following embodiments can be referred to. Figure 2 , Figure 2is a flowchart of a step of obtaining a video feature carrying context information provided by the embodiment of the present application, wherein the step comprises steps 201 to 203.
[0047] Step 201, performing feature interaction on each feature in the initial video feature based on the multi-head attention mechanism to obtain the corresponding interaction feature;
[0048] Step 202, performing size mapping processing on the interaction feature to obtain the interaction video feature, wherein the interaction video feature has the same feature size as the initial video feature;
[0049] Step 203, performing feature aggregation processing on the interaction video feature and the initial video feature to obtain the video feature carrying the context information.
[0050] Specifically, when processing, first, the multi-head attention mechanism is used to perform feature interaction processing on each feature in the initial video feature to obtain the corresponding interaction feature, then size mapping processing is performed on the interaction feature to obtain the interaction video feature having the same feature size as the initial video feature, and finally, the interaction video feature and the initial video feature are aggregated to obtain the video feature carrying the context information.
[0051] Exemplarily, when the context feature of the original video is extracted, the initial video feature obtained by feature extraction is input into the context information extraction module, the initial video feature is interacted and learned through the multi-head attention mechanism algorithm, and then the initial video feature after interaction and the initial video feature are aggregated to obtain the video feature carrying the context information.
[0052] For the obtained initial video feature, each video frame in the video frame sampled contains a feature, at this time, when performing interaction processing, the correlation between one feature and the remaining features is calculated, and then each feature in the initial video feature is updated according to the correlation.
[0053] Among them, when performing interaction processing, it includes: selecting a first initial video feature in the initial video feature, and calculating an interaction similarity matrix of the first initial video feature and a second initial video feature based on the multi-head attention mechanism, wherein the initial video feature is composed of the first initial video feature and the second initial video feature, and the number of the first initial video feature is one; aggregating the similarity matrix with each video feature in the second initial video feature, and collecting the obtained aggregation result to obtain the interaction feature corresponding to the first initial video feature; under the condition that each feature in the initial video feature is processed, the interaction feature corresponding to the initial video feature is obtained.
[0054] That is, when the interaction feature is obtained by processing, each feature in the initial video feature is updated, and when the update is performed, the obtained interaction feature is used as the updated feature. Taking one of the features as an example, a first initial video feature is selected from the initial video feature, a similarity matrix between the first initial video feature and a second initial video feature is calculated based on the multi-head attention mechanism, the second initial video feature is a set of all features in the initial video feature except the first initial video feature, after the similarity matrix is obtained, the similarity matrix is aggregated with each video feature in the second initial video feature, and each aggregation process obtains a feature, and finally all the features are summarized to obtain the interaction feature corresponding to the first initial video feature. After the interaction of each feature in the initial video feature is completed, the interaction feature corresponding to the initial video feature is obtained.
[0055] After obtaining the interaction feature, in order to ensure the consistency of the feature size, the interaction feature is adjusted in size to obtain an interaction video feature, and finally the interaction video feature is aggregated with the initial video feature again. The video feature obtained after processing will carry context information.
[0056] In the actual processing process, when the context information of the original video is extracted and processed to obtain a video feature carrying context information, first, one feature in the initial video feature can be converted into a query value Q v , a key value K v and a real value V v , and the specific conversion method is as follows:
[0057]
[0058] K v =LayerNorm(F v )W2,
[0059] V v =LayerNorm(F v )W3,
[0060] Where F v is the initial video feature, F v · is the transposed initial video feature, W1, W2, and W3 represent weight matrices respectively.
[0061] At this time, the correlation weight (similarity matrix) between each feature and other features is calculated by using the multi-head self-attention mechanism, and then the correlation weight is used for feature aggregation to obtain the feature used for replacement.
[0062] For the multi-head attention mechanism used, the number of parallel layers can be set as H, and 1 < h < H, then the output feature calculation formula of the h-th layer is:
[0063]
[0064] When each feature processed is obtained, the outputs of the same feature in all parallel layers are spliced, such as head-to-tail, to obtain the spliced feature:
[0065] attn_out=MultiHead(Q v ,K v ,V v )=concat(head1,…,head1)W o ;
[0066] Wherein, W0 represents a weight matrix.
[0067] Then, the spliced feature is mapped back to the original size to obtain the interactive video feature after interaction:
[0068] Z v =LayerNorm(FC(attn_out)+LayerNorm(attn_out)),
[0069] Wherein, FC represents a full connection layer.
[0070] Finally, when the video feature carrying context information is obtained, the obtained interactive video feature and the corresponding initial video feature are aggregated to obtain the video frame feature carrying context information:
[0071] C v =Conv(LayerNorm(F v +Z v );
[0072] Wherein, Conv is a convolution layer with a kernel of 1x1.
[0073] Step 104, training the video text retrieval model according to the first text feature, the second text feature and the video feature carrying context information, and obtaining the trained video text retrieval model when it is determined that the training is completed.
[0074] In an embodiment, after the feature extraction and feature processing are completed, the first text feature, the second text feature and the video feature carrying the context information are obtained, and the video-text retrieval model is trained based on the first text feature, the second text feature and the video feature. Specifically, the first text feature, the second text feature and the video feature are input into the video-text retrieval model, and whether the training is completed is determined according to the loss value generated during the training, and then the trained video-text retrieval model is obtained when the training is completed.
[0075] Specifically, when the video-text retrieval model is trained, the first text feature, the second text feature and the video feature can be referred to Figure 3 , Figure 3 is a flowchart of the step of training the video-text retrieval model provided by the embodiments of the present application, wherein the step includes steps 301 to 303.
[0076] Step 301, reconstructing the video feature according to the first text feature to obtain the first reconstructed video feature, and reconstructing the video feature according to the second text feature to obtain the second reconstructed video feature;
[0077] Step 302, training the video-text retrieval model according to the first text feature and the first reconstructed video feature to obtain the first loss value, and training the video-text retrieval model according to the second text feature and the second reconstructed video feature to obtain the second loss value;
[0078] Step 303, obtaining the trained video-text retrieval model in the case that the training is completed according to the first loss value and the second loss value.
[0079] In an embodiment, when the training is performed based on the first text feature, the second text feature and the video feature carrying the context information, first, the reconstruction of the video feature is performed to obtain the first reconstructed video feature associated with the first text feature and the second reconstructed video feature associated with the second text feature, then the first text feature and the first reconstructed video feature are taken as a group to train the video-text retrieval model, and the second text feature and the second reconstructed video feature are taken as a group to train the video-text retrieval model, and in each training stage, the corresponding training loss is generated, specifically including the first loss value and the second loss value, and then the trained video-text retrieval model is obtained when the training is completed according to the first loss value and the second loss value.
[0080] In the actual training process, video features need to be reconstructed to align with text features. This video reconstruction requires reconstructing video features separately for the first and second text features. Specifically, the process includes: calculating the similarity weight between each feature in the video features and the first text features; updating each feature in the video features based on the similarity weights, and performing feature mapping on the updated features to obtain the mapped video features; and aggregating the mapped video features to obtain the first reconstructed video features associated with the first text data.
[0081] In other words, by inputting video features carrying contextual information and text features together into the cross-modal interaction module, the correlation weights of text features and video features are calculated using a self-attention mechanism algorithm. Then, feature aggregation is performed using the correlation weights to obtain reconstructed video features that are related to the text.
[0082] For modules used for video feature reconstruction, please refer to... Figure 4 Based on Figure 4 When the interactive processing of the block diagram shown completes the reconstruction of video features, the first and second text features are converted into corresponding query values Q. Simultaneously, the video features carrying contextual information are converted into key values K and real values V. Then, a self-attention algorithm is used to calculate the similarity between each feature in the first text features and the video features carrying contextual information. Based on the similarity, the video features carrying contextual information are aggregated to obtain video features related to the first text data, thus completing the reconstruction of video features based on the first text data. Simultaneously, the reconstruction of video features based on the second text data is also completed. Based on the method used to obtain the first reconstructed video features, the second reconstructed video features can be obtained to complete modality alignment.
[0083] For example, taking the first reconstructed video feature associated with the first text feature as an example, during processing, the first text feature is first converted into the query value q. t The video features carrying contextual information are converted into key-value pairs k. v and real value v v The self-attention algorithm is used to calculate similarity for feature aggregation. Finally, the aggregated video features are processed by a linear layer and mapped to the same dimension as the video features carrying contextual information.
[0084] r v|t =LayerNorm(Attention(q) t ,k v ,v v W t );
[0085] wherein W t is an adaptive weight matrix.
[0086] Finally, when performing the aggregation processing, the obtained first reconstructed video feature associated with the first text feature is:
[0087]
[0088] Similarly, when reconstructing the video feature based on the second text feature, the second reconstructed video feature associated with the second text feature can be obtained in the same way as the first reconstructed video feature, and thus will not be described in detail here.
[0089] Further, after completing the reconstruction of the video feature, the video-text retrieval model is trained and optimized based on the reconstructed video feature, the first text feature, and the second text feature. When training, the training can be divided into two stages. The first stage is training based on the first text feature and the first reconstructed video feature. The second stage is training based on the second text feature and the second reconstructed video feature. Each training stage produces a corresponding loss value.
[0090] In an embodiment, the process of obtaining the loss value of each stage includes: inputting the first text feature and the first reconstructed video feature into the video-text retrieval model, calculating a first similarity matrix of the first text feature and the first reconstructed video feature, and according to the first similarity matrix, calculating a first loss value of text retrieval video and a first loss of video retrieval text based on the training sample; inputting the second text feature and the second reconstructed video feature into the video-text retrieval model, calculating a second similarity matrix of the second text feature and the second reconstructed video feature, and according to the second similarity matrix, calculating a second loss value of text retrieval video and a second loss of video retrieval text based on the training sample.
[0091] As described above, the first stage of training and the second stage of training are the same, with the difference being the data used for training. When obtaining the loss value of each stage, the similarity matrix of the text feature and the reconstructed video feature is calculated, and then the loss value generated during training of different stages is calculated according to the similarity matrix.
[0092] Taking the first loss value of the first stage as an example, when training the video-text retrieval model, the first reconstructed video feature associated with the first text feature is calculated by cosine similarity with the first text feature, and a first similarity matrix is obtained:
[0093]
[0094] wherein, is the first reconstructed video feature, is the transposed first reconstructed video feature, c t is the first text feature.
[0095] Then, when calculating the loss value at this stage, the loss of the text retrieval video task and the loss of the video retrieval text task are calculated, and then the loss summation is performed to obtain the first loss value when training at this stage. Specifically, when the number of training samples is B, and belongs to batch Ω, the loss L t2v of the text retrieval video task can be obtained according to the obtained first similarity matrix, and the loss L v2t of the video retrieval text task is calculated as follows:
[0096]
[0097]
[0098] And the first loss value at this time is: L1=L t2v +L v2t .
[0099] Similarly, when the second loss value L2 is obtained, the method for obtaining the first loss value is also used.
[0100] Further, when the trained video text retrieval model is obtained, the first loss value and the second loss value are summed to obtain a total loss value, and then whether the training is completed is determined based on the total loss value, specifically including: adding the first loss value and the second loss value to obtain the total loss value when training the video text retrieval model; in the case that the total loss value is less than a preset threshold, it is determined that the training is completed, and the trained video text retrieval model is obtained.
[0101] That is, after the total loss value is summed, the total loss value is compared with the preset threshold to determine whether the training is completed, wherein when the total loss value is less than the preset threshold, it is determined that the training is completed, and the trained video text retrieval model is obtained.
[0102] Further, referring to Figure 5 , Figure 5 is a block diagram schematic diagram provided by the embodiment of the present application for training the video text retrieval model.
[0103] As Figure 5As shown, when training the constructed video-text retrieval model, the input video data and the corresponding first text data are determined, and the first text data can be translated by using the language translation function to obtain second text data. Then, the first text data, the second text data, and the video data are respectively extracted to obtain first text features, second text features, and video features carrying context information. Then, the video features are reconstructed in two parallel processes to obtain video features associated with the first text data / second text data. Then, semantic interaction processing is performed, similarity matrices at different stages are calculated, and finally, the loss values at different stages are determined according to the similarity matrices to obtain the total loss value in the training process, which is used to determine whether the training is completed.
[0104] Further, after the training of the constructed video-text retrieval model is completed, when used, the video data to be queried is input into the trained video-text retrieval model, and the final matching processing is completed by calculating the similarity between the global features of the video and the text features.
[0105] Specifically, when the video-text retrieval model performs video retrieval, it includes: receiving an input video to be retrieved; frame sampling the video to be retrieved to obtain target video frames; extracting features from the target video frames to obtain target video features of the video to be retrieved; inputting the target video features into the trained video-text retrieval model to obtain similarity values between the target video features and each text feature in the database, and outputting the target text data corresponding to the maximum similarity value as the data matched with the video to be retrieved, wherein the text data and the text features are one-to-one corresponding.
[0106] Exemplarily, when the video to be retrieved is received, frame sampling is performed on the video to be retrieved to obtain target video frames for feature extraction. Then, the target video frames are extracted to obtain target video features. Further, the target video features are input into the trained video-text retrieval model. Through the processing of the model, the similarity values between the target video features and the text features of each text data can be calculated, and the text data corresponding to the maximum similarity value is the target text matched with the video to be retrieved.
[0107] To sum up, the application discloses a training method of a video text retrieval model. When training the constructed video text retrieval model, a training sample is acquired, wherein the training sample includes an original video, first text data corresponding to the original video, and second text data obtained by performing language conversion on the first text data. Then, feature extraction is performed on the training sample to obtain corresponding initial video features, first text features, and second text features. Meanwhile, feature interaction and aggregation processing are performed on the initial video features to obtain video features carrying context information. Finally, the video text retrieval model is trained according to the first text features, the second text features, and the video features, so as to obtain a trained video text retrieval model when the training is completed. The context information is added to the video features in the training process, the association between video frames is mined, the accuracy of video feature extraction is improved, the original text data is subjected to language conversion, the semantic alignment is performed by using multilingual text, the interference of polysemous words and expression habits is reduced, the accuracy of text features is improved, and the retrieval accuracy of the video text retrieval model is improved.
[0108] According to the method described in the above embodiment, the present embodiment will be further described from the perspective of a video text retrieval model training device. The video text retrieval model training device can be implemented as an independent entity or integrated in an electronic device, such as a terminal, which can include a mobile phone, a tablet computer, etc.
[0109] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of a video text retrieval model training device provided by the application, as Figure 6 indicated, the video text retrieval model training device 600 provided by the application includes:
[0110] The data acquisition module 601 is configured to acquire a training sample, wherein the training sample includes an original video, first text data, and second text data, and the second text data is obtained by performing language conversion on the first text data.
[0111] The feature extraction module 602 is configured to perform feature extraction on the training sample to obtain initial video features, first text features, and second text features.
[0112] The feature processing module 603 is configured to perform feature interaction and aggregation processing on the initial video features to obtain video features carrying context information.
[0113] The model training module 604 is configured to train the video text retrieval model according to the first text features, the second text features, and the video features carrying context information, and obtain a trained video text retrieval model when it is determined that the training is completed.
[0114] In an embodiment, the feature extraction module 602 is further configured to:
[0115] sample video frames from the original video to obtain video key frames for feature extraction, and perform feature extraction on the video key frames to obtain initial video features of the original video;
[0116] perform feature extraction on the first text data to obtain first text features, and perform feature extraction on the second text data to obtain second text features.
[0117] In an embodiment, the feature extraction module 602 is further configured to:
[0118] perform word segmentation processing on the first text data to obtain a first text sequence;
[0119] fill the first text sequence according to position information corresponding to the first text sequence;
[0120] perform encoding processing on the filled first text sequence to obtain first text features of the first text data;
[0121] and,
[0122] perform word segmentation processing on the second text data to obtain a second text sequence;
[0123] fill the second text sequence according to position information corresponding to the second text sequence;
[0124] perform encoding processing on the filled second text sequence to obtain second text features of the second text data.
[0125] In an embodiment, the feature processing module 603 is further configured to:
[0126] perform feature interaction on each feature in the initial video features based on a multi-head attention mechanism to obtain corresponding interaction features;
[0127] perform size mapping processing on the interaction features to obtain interaction video features, wherein the interaction video features have the same feature size as the initial video features;
[0128] perform feature aggregation processing on the interaction video features and the initial video features to obtain video features carrying context information.
[0129] In an embodiment, the feature processing module 603 is further configured to:
[0130] select a first initial video feature from the initial video features, and calculate an interaction similarity matrix of the first initial video feature and a second initial video feature based on a multi-head attention mechanism, wherein the initial video features consist of the first initial video feature and the second initial video feature, and the number of the first initial video features is one;
[0131] aggregate the similarity matrix with each video feature in the second initial video features, and aggregate the aggregation results of all the video features to obtain an interaction feature corresponding to the first initial video feature;
[0132] In the case that each feature in the initial video features is processed, the interaction feature corresponding to the initial video features is obtained.
[0133] In an embodiment, the model training module 604 is further configured to:
[0134] reconstruct the video features according to the first text features to obtain first reconstructed video features, and reconstruct the video features according to the second text features to obtain second reconstructed video features;
[0135] train the video-text retrieval model according to the first text features and the first reconstructed video features to obtain a first loss value, and train the video-text retrieval model according to the second text features and the second reconstructed video features to obtain a second loss value;
[0136] In the case that the training is determined to be completed according to the first loss value and the second loss value, the trained video-text retrieval model is obtained.
[0137] In an embodiment, the model training module 604 is further configured to:
[0138] calculate a similarity weight of each feature in the video features and the first text features;
[0139] update the features in the video features according to the similarity weights, and perform feature mapping on the updated features to obtain mapped video features;
[0140] perform feature aggregation on the mapped video features to obtain first reconstructed video features associated with the first text data.
[0141] In an embodiment, the model training module 604 is further configured to:
[0142] input the first text features and the first reconstructed video features into the video-text retrieval model, calculate a first similarity matrix of the first text features and the first reconstructed video features, and calculate a first loss value of the text retrieval video and a first loss of the video retrieval text based on the first similarity matrix, which are trained based on the training samples;
[0143] The second text feature and the second reconstructed video feature are input into the video-text retrieval model, a second similarity matrix of the second text feature and the second reconstructed video feature is calculated, and a second loss value of the text retrieval video and a second loss value of the video retrieval text are calculated based on the second similarity matrix.
[0144] In an embodiment, the model training module 604 is further configured to:
[0145] The first loss value and the second loss value are added to obtain a total loss value of the video-text retrieval model during training.
[0146] In a case where the total loss value is less than a preset threshold, it is determined that the training is completed, and a trained video-text retrieval model is obtained.
[0147] In an embodiment, the training device 800 of the video-text retrieval model further comprises a retrieval matching module configured to:
[0148] receive an input video to be retrieved;
[0149] frame sampling is performed on the video to be retrieved to obtain a target video frame;
[0150] feature extraction is performed on the target video frame to obtain a target video feature of the video to be retrieved;
[0151] The target video feature is input into the trained video-text retrieval model to obtain a similarity value between the target video feature and each text feature in the database, and the target text data corresponding to the maximum similarity value is output as the data matched with the video to be retrieved, wherein the text data and the text feature are one-to-one corresponding.
[0152] In addition, please refer to Figure 7 , Figure 7 is a structural schematic diagram of an electronic device provided by the embodiment of the present application. The electronic device can be a mobile terminal such as a smart phone, a tablet computer, etc. As shown in Figure 7 , the electronic device 700 comprises a processor 701 and a memory 702. The processor 701 is electrically connected with the memory 702.
[0153] The processor 701 is the control center of the electronic device 700, and connects each part of the whole electronic device through various interfaces and lines. By running or loading the application program stored in the memory 702 and calling the data stored in the memory 702, the processor 701 executes various functions and processes data of the electronic device 700, so as to monitor the whole electronic device 700.
[0154] In the embodiment, the processor 701 in the electronic device 700 loads the instructions corresponding to the processes of one or more application programs into the memory 702, and runs the application programs stored in the memory 702 by the processor 701, so as to implement any step in the training method of the video text retrieval model provided in the above embodiments.
[0155] The electronic device 700 can implement the steps in any embodiment of the training method of the video text retrieval model provided in the embodiments, and thus can implement the beneficial effects of any training method of the video text retrieval model provided in the embodiments. Details are described in the above embodiments, and thus will not be described here.
[0156] Please refer to Figure 8 , Figure 8 is another structural schematic diagram of the electronic device provided in the embodiments, as Figure 8 shown, Figure 8 a specific structural block diagram of the electronic device provided in the embodiments is shown, which can be used to implement the training method of the video text retrieval model provided in the above embodiments. The electronic device 800 can be a mobile terminal such as a smart phone or a notebook computer.
[0157] The RF circuit 810 is used to receive and send electromagnetic waves, and to convert electromagnetic waves and electrical signals to each other, so as to communicate with a communication network or other devices. The RF circuit 810 can include various existing circuit elements for performing these functions, for example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and the like. The RF circuit 810 can communicate with various networks, such as the Internet, an intranet, a wireless network, or other devices through the wireless network. The wireless network can include a cellular telephone network, a wireless local area network or a metropolitan area network. The wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers (IEEE) standards 802.11a, 802.11b, 802.11g and / or 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols used for mail, instant messaging and short message service, and any other suitable communication protocol, and can even include protocols that have yet to be developed.
[0158] The memory 820 can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method of the video text retrieval model in the above embodiments. The processor 880 executes various functions and the training method of the video text retrieval model by running the software programs and modules stored in the memory 820.
[0159] Memory 820 can include high-speed random access memory and can also include nonvolatile memory, such as one or more magnetic data storage disks, flash memory, or other nonvolatile solid-state memory. In some examples, memory 820 can further include memory that is remote from processor 880, such as the memory memory of a remote server that is connected to electronic device 800 through a network. Examples of such networks include, without limitation, the Internet, an enterprise intranet, a local area network, a wide area network, a mobile communications network, and combinations thereof.
[0160] Input unit 830 can be used to receive input of numbers or character information, and to generate a keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control. Specifically, input unit 830 can include a touch-sensitive surface 831 and other input devices 832. Touch-sensitive surface 831, also known as a touch display screen or touchpad, can collect touch operations (such as a user's operation on or near the touch-sensitive surface 831 using a finger, a stylus, or any suitable object or accessory) of the user on or near the touch-sensitive surface 831 and drive the corresponding connection device according to the pre-set program. Optionally, touch-sensitive surface 831 can include two parts of touch detection device and touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device and converts it into touch coordinates, and then sends it to processor 880, and can receive the command from processor 880 and execute it. In addition, the touch-sensitive surface 831 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to touch-sensitive surface 831, input unit 830 can also include other input devices 832. Specifically, other input devices 832 can include one or more of, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc.
[0161] The display unit 840 can be used to display information input by a user or information provided to the user, as well as various graphical user interfaces of the electronic device 800, which can be composed of graphics, text, icons, video, and any combination thereof. The display unit 840 can include a display panel 841, which can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like, optionally. Further, the touch-sensitive surface 831 can cover the display panel 841, and when the touch-sensitive surface 831 detects a touch operation thereon or adjacent thereto, transmit to the processor 880 to determine the type of touch event, and then the processor 880 provides corresponding visual output on the display panel 841 according to the type of touch event. Although in the figure, the touch-sensitive surface 831 and the display panel 841 are implemented as two independent components to realize the input and output functions, in some embodiments, the touch-sensitive surface 831 and the display panel 841 can be integrated to realize the input and output functions.
[0162] The electronic device 800 can further include at least one sensor 850, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor can include an ambient light sensor that can adjust the brightness of the display panel 841 according to the brightness of ambient light, and a proximity sensor that can generate an interrupt when the cover is closed or turned off. As one of the motion sensors, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and when at rest, can detect the magnitude and direction of gravity, which can be used for applications such as identifying the posture of the mobile phone (such as switching between landscape and portrait screens, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometers, tapping), and the like; as well as other sensors that the electronic device 800 can be configured, such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, and the like, which will not be described here.
[0163] The audio circuit 860, the speaker 861, and the microphone 862 can provide an audio interface between the user and the electronic device 800. The audio circuit 860 can convert received audio data into an electrical signal, transmit the electrical signal to the speaker 861, and convert the electrical signal into a sound signal by the speaker 861 for output; on the other hand, the microphone 862 converts the collected sound signal into an electrical signal, which is received by the audio circuit 860 and converted into audio data, and then output to the processor 880 for processing, and then transmitted to another terminal via the RF circuit 810, or output to the memory 820 for further processing. The audio circuit 860 can also include an earphone jack to provide communication between an external earphone and the electronic device 800.
[0164] The electronic device 800 can help the user to receive requests, send information, etc. through the transmission module 870 (e.g., a Wi-Fi module), which provides the user with wireless broadband Internet access. Although the transmission module 870 is shown in the figure, it can be understood that it does not belong to the necessary components of the electronic device 800, and can be omitted as needed without changing the essence of the application.
[0165] The processor 880 is the control center of the electronic device 800, which connects various parts of the entire mobile phone through various interfaces and lines, executes various functions of the electronic device 800 and processes data by running or executing software programs and / or modules stored in the memory 820 and calling data stored in the memory 820, thereby monitoring the entire electronic device. Optionally, the processor 880 can include one or more processing cores; in some embodiments, the processor 880 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 880.
[0166] The electronic device 800 further includes a power supply 890 (such as a battery) for supplying power to various components, and in some embodiments, the power supply can be logically connected to the processor 880 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 890 can also include one or more direct or alternating power supplies, recharging systems, power failure detection circuits, power converters or inverters, power state indicators, and any other components.
[0167] Although not shown, the electronic device 800 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be described here. In this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors to implement any one of the steps in the training method of the video text retrieval model provided by the above embodiments.
[0168] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined as the same or several entities, and the specific implementation of each of the above modules can be referred to the method embodiments described above, which will not be described here.
[0169] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware by the instructions, and the instructions can be stored in a computer readable storage medium and loaded and executed by a processor. To this end, the embodiments of the present application provide a storage medium, wherein a plurality of instructions are stored, and the instructions can be executed by a processor to implement any step in the training method of the video text retrieval model provided by the above embodiments.
[0170] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0171] Since the instructions stored in the storage medium can execute the steps in any embodiment of the training method of the video text retrieval model provided by the embodiments of the present application, the beneficial effects that can be achieved by any training method of the video text retrieval model provided by the embodiments of the present application can be achieved, which are described in detail in the foregoing embodiments and will not be described here.
[0172] The above describes in detail the training method, device, electronic device and storage medium of a video text retrieval model provided by the embodiments of the present application. The principle and implementation manner of the present application are described by applying specific examples. The above embodiment is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation manner and application range can be changed; in conclusion, the content of the specification should not be understood as a limitation of the present application. Moreover, for those skilled in the art, without departing from the principle of the present application, a number of improvements and refinements can be made, which are also regarded as the protection scope of the present application.
Claims
1. A method for training a video-text retrieval model, characterized in that, The method comprises: obtaining training samples, wherein the training samples comprise original videos, first text data and second text data, and the second text data is obtained by language conversion based on the first text data; performing feature extraction on the training samples to obtain initial video features, first text features and second text features; performing feature interaction and aggregation processing on the initial video features to obtain video features carrying context information; training a video-text retrieval model according to the first text features, the second text features and the video features carrying context information, and obtaining a trained video-text retrieval model when it is determined that the training is completed; wherein the feature interaction and aggregation processing on the initial video features to obtain video features carrying context information comprises: performing feature interaction on each feature in the initial video features based on a multi-head attention mechanism to obtain corresponding interaction features; performing size mapping processing on the interaction features to obtain interaction video features, wherein the interaction video features have the same feature size as the initial video features; performing feature aggregation processing on the interaction video features and the initial video features to obtain video features carrying context information; and the feature interaction on each feature in the initial video features based on a multi-head attention mechanism to obtain corresponding interaction features comprises: selecting a first initial video feature from the initial video features, and calculating an interaction similarity matrix of the first initial video feature and a second initial video feature based on a multi-head attention mechanism, wherein the initial video features comprise the first initial video feature and the second initial video feature, and the number of the first initial video features is one; aggregating the similarity matrix and each video feature in the second initial video features, and summarizing the aggregation results to obtain interaction features corresponding to the first initial video feature; obtaining the interaction features corresponding to each feature in the initial video features after the processing is completed.
2. The method of claim 1, wherein, The feature extraction on the training samples to obtain initial video features, first text features and second text features comprises: performing video frame sampling on the original videos to obtain video key frames for feature extraction, and performing feature extraction on the video key frames to obtain initial video features of the original videos; performing feature extraction on the first text data to obtain first text features, and performing feature extraction on the second text data to obtain second text features.
3. The method of claim 2, wherein, The feature extraction on the first text data to obtain first text features comprises: performing word segmentation processing on the first text data to obtain first text sequences; filling the first text sequences according to position information corresponding to the first text sequences; performing encoding processing on the filled first text sequences to obtain first text features of the first text data; The feature extraction on the second text data to obtain second text features comprises: The second text data is subjected to word segmentation processing to obtain a second text sequence; The second text sequence is filled according to position information corresponding to the second text sequence; The filled second text sequence is subjected to encoding processing to obtain a second text feature of the second text data.
4. The method of claim 1, wherein, The video text retrieval model is trained according to the first text feature, the second text feature and the video feature carrying context information, and the trained video text retrieval model is obtained when it is determined that the training is completed, and the trained video text retrieval model comprises: The video feature carrying context information is reconstructed according to the first text feature to obtain a first reconstructed video feature, and the video feature carrying context information is reconstructed according to the second text feature to obtain a second reconstructed video feature; The video text retrieval model is trained according to the first text feature and the first reconstructed video feature to obtain a first loss value, and the video text retrieval model is trained according to the second text feature and the second reconstructed video feature to obtain a second loss value; The trained video text retrieval model is obtained when it is determined that the training is completed according to the first loss value and the second loss value.
5. The method of claim 4, wherein, The video feature carrying context information is reconstructed according to the first text feature to obtain a first reconstructed video feature, and the video feature carrying context information is reconstructed according to the second text feature to obtain a second reconstructed video feature; The similarity weight of each feature in the video feature carrying context information and the first text feature is calculated; The features in the video feature carrying context information are updated according to the similarity weight, and the updated features are mapped to obtain the mapped video feature; The mapped video feature is aggregated to obtain the first reconstructed video feature associated with the first text data.
6. The method of claim 4, wherein, The video text retrieval model is trained according to the first text feature and the first reconstructed video feature to obtain a first loss value, and the video text retrieval model is trained according to the second text feature and the second reconstructed video feature to obtain a second loss value, and the video text retrieval model comprises: The first text feature and the first reconstructed video feature are input into the video text retrieval model, the first similarity matrix of the first text feature and the first reconstructed video feature is calculated, and the first loss value of the text retrieval video and the first loss of the video retrieval text based on the training sample are calculated according to the first similarity matrix; The second text feature and the second reconstructed video feature are input into the video text retrieval model, the second similarity matrix of the second text feature and the second reconstructed video feature is calculated, and the second loss value of the text retrieval video and the second loss of the video retrieval text based on the training sample are calculated according to the second similarity matrix.
7. The method of claim 4, wherein, The trained video text retrieval model is obtained when it is determined that the training is completed according to the first loss value and the second loss value, and the trained video text retrieval model comprises: The first loss value and the second loss value are added to obtain the total loss value of the video text retrieval model during training; In a case where the total loss value is less than a preset threshold, it is determined that the training is completed, and a trained video-text retrieval model is obtained.
8. The method of claim 1, wherein, After the trained video-text retrieval model is obtained in a case where it is determined that the training is completed, the method further includes: receiving an inputted video to be retrieved; frame sampling on the video to be retrieved to obtain a target video frame; feature extraction on the target video frame to obtain a target video feature of the video to be retrieved; inputting the target video feature into the trained video-text retrieval model to obtain a similarity value between the target video feature and each text feature in a database, and outputting a target text data corresponding to a maximum similarity value as data matched with the video to be retrieved, wherein each text data in the database corresponds to a text feature.
9. A training device for a video text retrieval model, characterized in that, comprises: a data acquisition module configured to acquire training samples, wherein the training samples comprise an original video, first text data, and second text data, and the second text data is obtained by language conversion based on the first text data; a feature extraction module configured to perform feature extraction on the training samples to obtain initial video features, first text features, and second text features; a feature processing module configured to perform feature interaction and aggregation processing on the initial video features to obtain video features carrying context information; a model training module configured to train a video-text retrieval model according to the first text features, the second text features, and the video features carrying context information, and obtain a trained video-text retrieval model in a case where it is determined that the training is completed; wherein the feature processing module is further configured to perform feature interaction on each feature in the initial video features based on a multi-head attention mechanism to obtain corresponding interaction features; perform size mapping processing on the interaction features to obtain interaction video features, wherein the interaction video features have the same feature size as the initial video features; perform feature aggregation processing on the interaction video features and the initial video features to obtain the video features carrying context information; and the feature processing module is further configured to: select a first initial video feature from the initial video features, and calculate an interaction similarity matrix of the first initial video feature and a second initial video feature based on a multi-head attention mechanism, wherein the initial video features consist of the first initial video feature and the second initial video feature, and the number of the first initial video features is one; aggregate the similarity matrix and each video feature in the second initial video features, and summarize the obtained aggregation results to obtain interaction features corresponding to the first initial video feature; obtain the interaction features corresponding to each feature in the initial video features in a case where the processing on each feature in the initial video features is completed.
10. An electronic device, comprising: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the method of any one of claims 1 to 8 when executing the computer program.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Cross-modal retrieval model training method and device, equipment and storage medium
CN115168638A
Cross-modal retrieval model processing method and device, equipment, product and medium
CN115994243A