Training Method, Retrieval Method and Related Devices for Video Text Retrieval Model

Through the pre-training of the video feature extraction model and the step-by-step training method of the text feature extraction model, the similarity loss value of the video feature and text feature is used to adjust the model parameters, which solves the problem of poor performance of the video text retrieval model under limited training resources, and achieves efficient video text retrieval effect.

CN115599953BActive Publication Date: 2025-07-04BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211183287.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-04
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Under limited training resources, the performance of existing video text retrieval models is poor and cannot effectively use multimodal data for efficient training.

Method used

By pre-training the video feature extraction model and the text feature extraction model to be trained, the model parameter adjustment is used to adjust the similarity loss value of the video feature and text feature, and the video text search model is trained step by step to reduce the number of model parameter adjustments and improve model performance.

Benefits of technology

Under limited training resources, the performance of the video text retrieval model is improved through step-by-step training method, reduced video memory consumption, and allowed the text feature extraction model to be trained with a larger number of samples, ensuring that the model is easier to converge and achieving efficient video text retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599953B_ABST
    Figure CN115599953B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, a retrieval method and related devices for a video text retrieval model. The training method includes: inputting a first video sample into a pre-trained video feature extraction model to obtain a first video feature; inputting the description text data of the first video sample into a text feature extraction model to be trained to obtain a first text feature; determining a second video feature and a second text feature, as well as a third video feature and a third text feature from the first video feature and the first text feature; determining a first loss value according to the second video feature and the second text feature, and determining a second loss value according to the third video feature and the third text feature; adjusting the model parameters of the text feature extraction model to be trained based on the first loss value and the second loss value to obtain a trained text feature extraction model; using the pre-trained video feature extraction model and the trained text feature extraction model as a video text retrieval model, and the performance of the video text retrieval model is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to Internet application technologies, and in particular, to a training method, a retrieval method, and related devices for a video text retrieval model. Background Art

[0002] A lot of data exists in the form of modalities. For the same event, it can be represented by various modalities of data such as images, videos, audios, and texts. With the continuous emergence of various video platforms, the demand for video text retrieval is getting higher and higher. Among them, video text retrieval means retrieving the corresponding text according to a video, or retrieving the corresponding video according to a text. Currently, when obtaining a retrieval model for video text retrieval, a large sample size of data is required, so the requirements for training resources are relatively high during the training process. Under limited training resources, the performance of the trained model is poor. Summary of the Invention

[0003] The present disclosure provides a training method, a retrieval method, and related devices for a video text retrieval model, so as to at least solve the technical problem that the performance of the trained model is poor under limited training resources in the related art. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a training method for a video text retrieval model is provided, including the following steps:

[0005] Obtain a first video sample and the description text data corresponding to the first video sample;

[0006] Input the first video sample into a pre-trained video feature extraction model to obtain a first video feature;

[0007] Input the description text data corresponding to the first video sample into a text feature extraction model to be trained to obtain a first text feature;

[0008] In the first video feature and the first text feature, determine a second video feature and a second text feature originating from the same first video sample, and a third video feature and a third text feature originating from different first video samples; according to the second video feature and the second text feature, determine at least one first loss value, and according to the third video feature and the third text feature, determine at least one second loss value; based on the at least one first loss value and the at least one second loss value, adjust the model parameters of the text feature extraction model to be trained to obtain a trained text feature extraction model;

[0009] Use the pre-trained video feature extraction model and the trained text feature extraction model as a video-text retrieval model, which is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data.

[0010] Optionally, the at least one first loss value includes a first video-text loss value and a first text-video loss value. Determining the at least one first loss value according to the second video feature and the second text feature includes:

[0011] Determine a first video-text similarity between the second video feature and the second text feature, and a first text-video similarity between the second text feature and the second video feature; the first video-text similarity is used to represent the result of comparing the similarity between the second video feature and the second text feature with the second video feature first; the first text-video similarity is used to represent the result of comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0012] Based on the first video-text similarity and a first preset video-text label, determine the first video-text loss value; based on the first text-video similarity and a first preset text-video label, determine the first text-video loss value; the first preset video-text label is used to represent the annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent the annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0013] Use at least one of the first video-text loss value and the first text-video loss value as the first loss value.

[0014] Optionally, determining the first video-text similarity between the second video feature and the second text feature, and the first text-video similarity between the second text feature and the second video feature includes:

[0015] Perform regularization processing on the second video feature to obtain a regularized second video feature;

[0016] Perform regularization processing on the second text feature to obtain a regularized second text feature;

[0017] Perform a dot product on the regularized second video feature and the regularized second text feature to determine the first video-text similarity between the second video feature and the second text feature; perform a dot product on the regularized second text feature and the regularized second video feature to determine the first text-video similarity between the second text feature and the second video feature.

[0018] Optionally, the at least one second loss value includes a second video-text loss value and a second text-video loss value. Determining the at least one second loss value according to the third video feature and the third text feature includes:

[0019] Determine the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature; the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0020] Based on the second video-text similarity and the second preset video-text label, determine the second video-text loss value; based on the second text-video similarity and the second preset text-video label, determine the second text-video loss value; the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0021] Use at least one of the second video-text loss value and the second text-video loss value as the second loss value.

[0022] Optionally, before the step of inputting the first video sample into the pre-trained video feature extraction model to obtain the first video feature, the method further includes:

[0023] Obtain the image data, content text data, and preset classification label corresponding to the second video sample respectively;

[0024] Input the image data and the content text data into the video feature extraction model to be trained to obtain the video feature during pre-training;

[0025] Based on the third loss value corresponding to the video features in the pre-training and the preset classification labels, adjust the model parameters of the video feature extraction model to be trained to obtain a pre-trained video feature extraction model.

[0026] Optionally, the inputting the image data and the content text data into the video feature extraction model to be trained to obtain the video features in the pre-training includes:

[0027] Input the image data and the content text data into the video feature extraction model to be trained to obtain corresponding image features and text features in the pre-training respectively;

[0028] Perform feature fusion on the image features and the text features in the pre-training to obtain the video features in the pre-training.

[0029] Optionally, the obtaining the content text data corresponding to the second video sample respectively includes:

[0030] Obtain the video speech recognition result and the video image text recognition result of each second video sample;

[0031] Based on the video speech recognition result and the video image text recognition result, obtain the content text data of the second video sample.

[0032] According to the second aspect of the embodiments of the present disclosure, there is provided a video text retrieval method, including:

[0033] Obtain the data to be retrieved, where the data to be retrieved is video retrieval data or text retrieval data;

[0034] Input the data to be retrieved into the video text retrieval model to obtain the target retrieval data; when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data;

[0035] Wherein, the video text retrieval model is obtained according to the training method of the video text retrieval model described above.

[0036] According to the third aspect of the embodiments of the present disclosure, there is provided a training device for a video text retrieval model, including the following modules:

[0037] A data acquisition module, configured to acquire a first video sample and the description text data corresponding to the first video sample;

[0038] A first extraction module, configured to input the first video sample into the pre-trained video feature extraction model to obtain a first video feature;

[0039] A second extraction module, configured to input the description text data corresponding to the first video sample into a text feature extraction model to be trained, and obtain first text features;

[0040] A loss determination module, configured to determine, from the first video features and the first text features, second video features and second text features that originate from the same first video sample, and third video features and third text features that originate from different first video samples; determine at least one first loss value according to the second video features and the second text features, and determine at least one second loss value according to the third video features and the third text features; based on the at least one first loss value and the at least one second loss value, adjust the model parameters of the text feature extraction model to be trained to obtain a trained text feature extraction model;

[0041] A model acquisition module, configured to use the pre-trained video feature extraction model and the trained text feature extraction model as a video-text retrieval model, and the video-text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data.

[0042] Optionally, the loss determination module includes:

[0043] A first similarity determination unit, configured to determine a first video-text similarity between the second video feature and the second text feature, and a first text-video similarity between the second text feature and the second video feature; the first video-text similarity is used to represent the result of comparing the similarity between the second video feature and the second text feature with the second video feature first; the first text-video similarity is used to represent the result of comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0044] A first loss value determination unit, configured to determine a first video-text loss value based on the first video-text similarity and a first preset video-text label; determine a first text-video loss value based on the first text-video similarity and a first preset text-video label; the first preset video-text label is used to represent the annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent the annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0045] A first loss value selection unit, configured to use at least one of the first video-text loss value and the first text-video loss value as the first loss value.

[0046] Optionally, the first similarity determination unit includes:

[0047] A first regularization subunit, configured to perform regularization processing on the second video feature to obtain a regularized second video feature;

[0048] A second regularization subunit, configured to perform regularization processing on the second text feature to obtain a regularized second text feature;

[0049] A similarity determination subunit, configured to perform a dot product on the regularized second video feature and the regularized second text feature to determine a first video-text similarity between the second video feature and the second text feature; perform a dot product on the regularized second text feature and the regularized second video feature to determine a first text-video similarity between the second text feature and the second video feature. Optionally, the loss determination module includes:

[0050] A second similarity determination unit, configured to determine a second video-text similarity between the third video feature and the third text feature, and a second text-video similarity between the third text feature and the third video feature; the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0051] A second loss value determination unit, configured to determine a second video-text loss value based on the second video-text similarity and a second preset video-text label; determine a second text-video loss value based on the second text-video similarity and a second preset text-video label; the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0052] A second loss value selection unit, configured to use at least one of the second video-text loss value and the second text-video loss value as the second loss value.

[0053] Optionally, the apparatus further includes: a pre-training module, and the pre-training module includes:

[0054] A data acquisition unit, configured to acquire image data, content text data, and a preset classification label respectively corresponding to a second video sample;

[0055] A feature extraction unit, configured to input the image data and the content text data into a video feature extraction model to be trained, and obtain video features during pre-training;

[0056] A training processing unit, configured to adjust model parameters of the video feature extraction model to be trained based on a third loss value corresponding to the video features during pre-training and the preset classification label, and obtain a pre-trained video feature extraction model.

[0057] Optionally, the feature extraction unit includes:

[0058] An extraction processing subunit, configured to input the image data and the content text data into a video feature extraction model to be trained, and obtain image features and text features during pre-training;

[0059] A fusion processing subunit, configured to perform feature fusion on the image features and the text features during pre-training to obtain video features during pre-training.

[0060] Optionally, the data acquisition unit includes:

[0061] An identification processing subunit, configured to obtain a video speech recognition result and a video image text recognition result of each second video sample;

[0062] A data acquisition subunit, configured to obtain the content text data of the second video sample based on the video speech recognition result and the video image text recognition result.

[0063] According to a fourth aspect of the embodiments of the present disclosure, a video text retrieval device is provided, including:

[0064] A data acquisition module, configured to acquire data to be retrieved, where the data to be retrieved is video retrieval data or text retrieval data;

[0065] A data retrieval module, configured to input the data to be retrieved into a video text retrieval model to obtain target retrieval data; when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data;

[0066] Wherein, the video text retrieval model is obtained according to the training method of the video text retrieval model described above.

[0067] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0068] A processor;

[0069] A memory for storing the processor-executable instructions;

[0070] Wherein, the processor is configured to execute the instructions to implement the training method of the video text retrieval model as described in the first aspect, or to implement the video text retrieval method as described in the second aspect.

[0071] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the training method of the video text retrieval model as described in the first aspect, or implement the video text retrieval method as described in the second aspect.

[0072] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product. The computer program product includes readable program code. When the readable program code is executed by a processor of an electronic device, the electronic device can execute the training method of the video text retrieval model as described in the first aspect, or implement the video text retrieval method as described in the second aspect.

[0073] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0074] The present disclosure pre-trains a video feature extraction model to obtain a pre-trained video feature extraction model. After obtaining a first video sample and the description text data corresponding to the first video sample, the first video sample is input into the pre-trained video feature extraction model to obtain a first video feature; and the description text data is input into a text feature extraction model to be trained to obtain a first text feature. Further, among the first video feature and the first text feature, a second video feature and a second text feature originating from the same first video sample, and a third video feature and a third text feature originating from different first video samples are determined. At least one first loss value is determined according to the second video feature and the second text feature, and at least one second loss value is determined according to the third video feature and the third text feature. The model parameters of the text feature extraction model to be trained are adjusted by using at least one first loss value and at least one second loss value, and a trained text feature extraction model is determined. The pre-trained video feature extraction model and the trained text feature extraction model are used as a video-text retrieval model, and the video-text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data. In the technical solution provided by the present disclosure, the video-text retrieval model is trained step by step, that is, the video feature extraction model is pre-trained. When training the video feature extraction model, the number of model parameters to be adjusted is small, so it can be ensured that the pre-trained video feature extraction model has high performance; then when training the text feature extraction model, the model parameters of the pre-trained video feature extraction model are fixed, and the gradient of the model parameters of the text feature extraction model to be trained is solved and updated by using at least one accurate first loss value and at least one second loss value, effectively reducing the number of model parameters to be adjusted at the same time and reducing the consumption of video memory. Thus, when the training resources are limited, a larger number of samples can be used to train the text feature extraction model, making the text feature extraction model easier to converge and ensuring that the trained text extraction model has high performance. Based on the pre-trained video feature extraction model and the trained text feature extraction model, a video-text retrieval model with high performance can be obtained.

[0075] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0077] Figure 1 is a flowchart of a method for training a video-text retrieval model according to an exemplary embodiment;

[0078] Figure 2 is a flowchart of a video text retrieval method shown according to an exemplary embodiment;

[0079] Figure 3 is a block diagram of a training device for a video text retrieval model shown according to an exemplary embodiment;

[0080] Figure 4 is a block diagram of a video text retrieval device shown according to an exemplary embodiment;

[0081] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0082] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0083] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0084] Figure 1 is a flowchart of a training method for a video text retrieval model shown according to an exemplary embodiment. The training method for the video text retrieval model is used for the server side. Specifically, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The method specifically includes the following steps.

[0085] In step S11, a first video sample and the description text data corresponding to the first video sample are obtained.

[0086] In this step, the first video sample is the video data obtained for model training. Each first video sample corresponds to descriptive text data, which is used to reflect the main content of the first video sample and can be words obtained by summarizing the content of the first video sample. For example, there is a first video sample A, and the descriptive text data of the first video sample A can be "pet". Among them, the first video sample A in the form of a video is a modal data, and the "pet" in the form of text is a modal data, and the two modal data have a one-to-one correspondence relationship.

[0087] In step S12, the first video sample is input into the pre-trained video feature extraction model to obtain the first video feature.

[0088] In this step, the video feature extraction model is pre-trained to obtain the pre-trained video feature extraction model. In the pre-training stage, the training resources can be fully utilized to obtain a relatively accurate pre-trained video feature extraction model. Thus, when the first video sample is input into the pre-trained video feature extraction model, a first video feature with higher accuracy can be obtained, providing a prerequisite for obtaining a video-text retrieval model with higher performance.

[0089] Specifically, the image data and content text data of the first video sample are obtained; the image data and content text data are input into the pre-trained video feature extraction model to obtain the first video feature. Among them, the image data is the data obtained by frame-dividing the first video sample, and the content text data is the data obtained by performing video speech recognition and video image text recognition on the first video sample. By comprehensively considering the image data and content text data of the first video sample, a first video feature with higher accuracy is obtained. It should be noted that since the first video feature and the first text feature belong to different modalities, in order to accurately determine the similarity between the first video feature and the first text feature subsequently, the first video feature and the first text feature need to be in the same space, that is, the video-text space. Therefore, the output layer of the pre-trained video feature extraction model outputs the first video feature in the video-text space.

[0090] In step S13, the descriptive text data corresponding to the first video sample is input into the text feature extraction model to be trained to obtain the first text feature.

[0091] In this step, the text feature extraction model to be trained is a pre-built neural network model, which is used for text feature extraction. Each descriptive text data is input into the text feature extraction model to be trained, and the output layer of the preset text feature extraction module outputs the first text feature in the video-text space. Specifically, the text feature extraction model to be trained can be a bert model.

[0092] In step 14, among the first video features and the first text features, second video features and second text features originating from the same first video sample, and third video features and third text features originating from different first video samples are determined; based on the second video features and the second text features, at least one first loss value is determined, and based on the third video features and the third text features, at least one second loss value is determined; based on the at least one first loss value and the at least one second loss value, the model parameters of the text feature extraction model to be trained are adjusted to obtain a trained text feature extraction model.

[0093] In this step, at least one first loss value is used to evaluate the difference between the second video features and the second text features originating from the same first video sample, and at least one second loss value is used to evaluate the difference between the third video features and the third text features originating from different first video samples. After obtaining at least one first loss value and at least one second loss value, with the model parameters of the pre-trained video feature extraction model fixed, the model parameters of the text feature extraction model to be trained are adjusted to obtain a trained text feature extraction model. The trained text feature extraction model is a model that has been trained and meets the training end condition, and the training end condition can be model convergence. Thus, there is no need to adjust the model parameters of the pre-trained video feature extraction model, and the training resources are fully utilized to train the text feature extraction model to be trained, obtaining a text feature extraction model with high performance.

[0094] In one embodiment, the at least one first loss value includes a first video-text loss value and a first text-video loss value. The step 14 of determining at least one first loss value based on the first video features and the first text features includes:

[0095] Determine the first video-text similarity between the second video features and the second text features, and the first text-video similarity between the second text features and the second video features; the first video-text similarity is used to represent the result of comparing the similarity between the second video features and the second text features with the second video features first; the first text-video similarity is used to represent the result of comparing the similarity between the second text features and the second video features with the second text features first;

[0096] Determine a first video-text loss value based on the first video-text similarity and the first preset video-text label; determine a first text-video loss value based on the first text-video similarity and the first preset text-video label; the first preset video-text label is used to represent annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0097] Use at least one of the first video-text loss value and the first text-video loss value as the first loss value.

[0098] In this embodiment, the second video feature and the second text feature derived from the same first video sample are the matching second video feature and second text feature, that is, the second video feature and the second text feature should have a relatively high similarity. For example, input the i-th first video sample into a pre-trained video feature extraction model to obtain the i-th first video feature, and input the i-th description text data corresponding to the i-th first video sample into a text feature extraction model to be trained to obtain the i-th first text feature. The i-th first video feature and the i-th first text feature are the matching second video feature and second text feature, both derived from the i-th first video sample. Although their modalities are different, their description objects are the same, so they should have a relatively high similarity.

[0099] There are two possible cases for the similarity between the second video feature and the second text feature derived from the same first video sample, namely the first video-text similarity between the second video feature and the second text feature, and the first text-video similarity between the second text feature and the second video feature. The first video-text similarity is used to represent the result of comparing the similarity between the second video feature and the second text feature with the second video feature first, and the first text-video similarity is used to represent the result of comparing the similarity between the second text feature and the second video feature with the second text feature first. A first preset video-text label and a first preset text-video label are preset. Among them, the first preset video-text label is used to represent the annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent the annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first. Using the first video-text similarity as the predicted value and the first preset video-text label as the true value, the first video-text loss value is accurately determined; using the first text-video similarity as the predicted value and the first preset text-video label as the true value, the first text-video loss value is accurately determined. At least one of the first video-text loss value and the first text-video loss value is used as the first loss value.

[0100] For example, determine the second video feature and the second text feature derived from the i-th first video sample, that is, the i-th first video feature and the i-th first text feature. The i-th first video feature is should be the most similar to the i-th first text feature, which is Multiply with The dot product result is denoted as is The similarity score between and is the first video-text similarity. Set the corresponding annotation data, that is, the first preset video-text label, to 1. According to and the corresponding annotation data, the first video-text loss value is determined. Correspondingly, the i-th first text feature, which is should be the most similar to the i-th first video feature, which is Multiply with The dot product result is denoted as is The similarity score between and is the first text-video similarity. Set the corresponding annotation data, that is, the first preset text-video label, to 1. According to and Determine the first text video loss value based on the corresponding tags.

[0101] In one embodiment, determining the first video text similarity between the second video feature and the second text feature, and the first text video similarity between the second text feature and the second video feature includes:

[0102] Perform regularization processing on the second video feature to obtain the regularized second video feature;

[0103] Perform regularization processing on the second text feature to obtain the regularized second text feature;

[0104] Perform dot product on the regularized second video feature and the regularized second text feature to determine the first video text similarity between the second video feature and the second text feature; perform dot product on the regularized second text feature and the regularized second video feature to determine the first text video similarity between the second text feature and the second video feature.

[0105] In this embodiment, before determining the first video text similarity and the first text video similarity, perform regularization processing on the second video feature and the second text feature respectively to reduce the amount of data in the second video feature and the second text feature, prevent overfitting, and effectively improve the utilization efficiency of training resources while retaining important features. Specifically, perform L2 regularization on the second video feature to obtain the regularized second video feature, and perform L2 regularization processing on the second text feature to obtain the regularized second text feature. Perform dot product on the regularized second video feature and the regularized second text feature to determine the first video text similarity between the second video feature and the second text feature; perform dot product on the regularized second text feature and the regularized second video feature to determine the first text video similarity between the second text feature and the second video feature.

[0106] In one embodiment, the at least one second loss value includes a second video text loss value and a second text video loss value. The step of determining the at least one second loss value according to the third video feature and the third text feature in step 14 includes:

[0107] Determine the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature; the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0108] Based on the second video-text similarity and the second preset video-text label, determine the second video-text loss value; based on the second text-video similarity and the second preset text-video label, determine the second text-video loss value; the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0109] Take at least one of the second video-text loss value and the second text-video loss value as the second loss value.

[0110] In this embodiment, the first video feature and the first text feature from different first video samples are the mismatched third video feature and third text feature, that is, the third video feature and the third text feature should have a relatively low similarity. That is to say, the third video feature corresponds to one first video sample, and the third text feature corresponds to another first video sample. For example, input the i-th first video sample into the pre-trained video feature extraction model to obtain the i-th first video feature, and input the j-th description text data corresponding to the j-th first video sample into the text feature extraction model to be trained to obtain the j-th first text feature. The i-th first video feature and the j-th first text feature are the mismatched third video feature and third text feature. They belong to different modalities and have different description objects, so they should have a relatively low similarity.

[0111] There are two possible situations for the similarity between the third video feature and the third text feature derived from different first video samples, namely the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature. Among them, the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first. A second preset video-text label and a second preset text-video label are preset in advance. Among them, the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first. Using the second video-text similarity as the predicted value and the second preset video-text label as the true value, the second video-text loss value is accurately determined; using the second text-video similarity as the predicted value and the second preset text-video label as the true value, the second text-video loss value is accurately determined. At least one of the second video-text loss value and the second text-video loss value is used as the second loss value.

[0112] For example, determine the third video feature and the third text feature of the first video samples from different sources, such as the i-th first video feature and the j-th first text feature. The i-th first video feature is and the first text features other than are not similar. Multiply and The dot product result is denoted as s j,i , s j,i is and The similarity score, that is, the second video-text similarity. Set the annotation data corresponding to s j,i as the second preset video-text label to 0. According to s j,i and the annotation data corresponding to s j,i , the second video-text loss value is determined. Correspondingly, the i-th first text feature is should be and the first video features other than are not similar. Multiply and The dot product result is denoted as s i,j , s i,j is and The similarity score is the second video-text similarity. Let s i,j The corresponding annotation data, i.e., the second preset video-text label, is set to 0. Based on s i,j and s i,j and their corresponding annotation data, determine the second video-text loss value.

[0113] In one embodiment, determining the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature includes:

[0114] Perform a dot product on the regularized third video feature and the regularized third text feature to determine the second video-text similarity between the third video feature and the third text feature; perform a dot product on the regularized third text feature and the regularized third video feature to determine the second text-video similarity between the third text feature and the third video feature.

[0115] In this embodiment, before determining the second video-text similarity and the second text-video similarity, first perform regularization processing on the third video feature and the third text feature respectively to reduce the amount of data in the third video feature and the third text feature, prevent the occurrence of overfitting, and effectively improve the usage efficiency of training resources while retaining important features. Specifically, perform L2 regularization on the third video feature to obtain the regularized third video feature, and perform L2 regularization on the third text feature to obtain the regularized third text feature. Perform a dot product on the regularized third video feature and the regularized third text feature to determine the second video-text similarity between the third video feature and the third text feature; perform a dot product on the regularized third text feature and the regularized third video feature to determine the second text-video similarity between the third text feature and the third video feature.

[0116] In one embodiment, after determining the first video-text loss value, the first text-video loss value, the second video-text loss value, and the second text-video loss value, adjust the model parameters of the text feature extraction model to be trained according to the first video-text loss value, the first text-video loss value, the second video-text loss value, and the second text-video loss value, and obtain the trained text feature extraction model. Specifically, the first cross-entropy loss value can be determined according to the first video-text loss value and the second video-text loss value, the second cross-entropy loss value can be determined according to the first text-video loss value and the second text-video loss value, and the model parameters of the text feature extraction model to be trained can be adjusted according to the first cross-entropy loss value and the second cross-entropy loss value to obtain the trained text feature extraction model.

[0117] In step S15, the pre-trained video feature extraction model and the trained text feature extraction model are used as a video-text retrieval model, which is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data.

[0118] In this step, the pre-trained video feature extraction model and the trained text feature extraction model are jointly used as a video-text retrieval model. The pre-trained video feature extraction model in the video-text retrieval model can be used for video feature extraction, and the trained text feature extraction model in the video-text retrieval model can be used for text feature extraction.

[0119] In the above embodiments, the video feature extraction model is pre-trained to obtain a pre-trained video feature extraction model. After obtaining the first video sample and the description text data corresponding to the first video sample, the first video sample is input into the pre-trained video feature extraction model to obtain the first video feature; and the description text data is input into the text feature extraction model to be trained to obtain the first text feature. Further, among the first video feature and the first text feature, the second video feature and the second text feature originating from the same first video sample, and the third video feature and the third text feature originating from different first video samples are determined. According to the second video feature and the second text feature, at least one first loss value is determined, and according to the third video feature and the third text feature, at least one second loss value is determined. The model parameters of the text feature extraction model to be trained are adjusted by using at least one first loss value and at least one second loss value, and the trained text feature extraction model is determined. The pre-trained video feature extraction model and the trained text feature extraction model are used as a video-text retrieval model, and this video-text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data. In the technical solution provided by the present disclosure, the video-text retrieval model is trained step by step, that is, the video feature extraction model is pre-trained. When training the video feature extraction model, the number of model parameters to be adjusted is small, so it can be ensured that the pre-trained video feature extraction model has high performance; then when training the text feature extraction model, the model parameters of the pre-trained video feature extraction model are fixed, and the gradient of the model parameters of the text feature extraction model to be trained is solved and updated by using the relatively accurate at least one first loss value and at least one second loss value, effectively reducing the number of model parameters to be adjusted at the same time and effectively reducing the consumption of video memory. Thus, when the training resources are limited, a larger number of samples can be used to train the text feature extraction model, making the text feature extraction model easier to converge, ensuring that the trained text extraction model has high performance. Based on the pre-trained video feature extraction model and the trained text feature extraction model, a video-text retrieval model with high performance can be obtained.

[0120] In one embodiment, before the step S12 of inputting the first video sample into the pre-trained video feature extraction model to obtain the first video feature, the method further includes:

[0121] In step 16, the image data, the content text data, and the preset classification label corresponding to the second video sample are obtained.

[0122] In this step, the second video sample is training data for training the video feature extraction model to be trained. The image data is data obtained by frame-dividing the second video sample, and the image data carries rich image information. The content text data is various text contents in the second video sample. The preset classification label is the video classification text pre-screened for each second video sample and is labeled data.

[0123] In one embodiment, obtaining the content text data corresponding to the second video sample in step 16 includes:

[0124] In step 161, obtain the video speech recognition result and the video image text recognition result of each second video sample.

[0125] In step 162, based on the video speech recognition result and the video image text recognition result, obtain the content text data of the second video sample.

[0126] In this embodiment, perform video speech recognition on the second video sample to obtain the video speech recognition result, and perform text recognition on the image obtained by frame-dividing the second video sample to obtain the video image text recognition result. Concatenate the video speech recognition result and the video image text recognition result to obtain the content text data of the second video sample. This content text data is rich in content, fully considering various possible text information and avoiding omission of important text information.

[0127] Step 17, input the image data and the content text data into the video feature extraction model to be trained to obtain the video features in pre-training.

[0128] In this step, fully consider the image data and the content text data to obtain relatively accurate video features in pre-training.

[0129] In one embodiment, step 17 inputs the image data and the content text data into the video feature extraction model to be trained to obtain the video features in pre-training, including:

[0130] In step 171, input the image data and the content text data into the video feature extraction model to be trained to obtain image features and text features in pre-training.

[0131] Specifically, the video feature extraction model to be trained includes a video feature extraction module and a text feature extraction module. Input the image data into the video feature extraction module to obtain image features, and input the content text data into the text feature extraction module to obtain text features in pre-training. Among them, the video feature extraction module can be resnet-50 (a residual network structure), and the text feature extraction module can be a bert network.

[0132] In step 172, the image features and the text features in the pre-training are subjected to feature fusion to obtain video features in the pre-training.

[0133] In this step, the image features and the text features in the pre-training are subjected to feature fusion to obtain video features in the pre-training that integrate multi-modal features. Specifically, a multi-head attention module is used to fuse the image features and the text features in the pre-training to obtain video features in the pre-training.

[0134] Step 18: Based on the third loss value corresponding to the video features in the pre-training and the preset classification label, the model parameters of the video feature extraction model to be trained are adjusted to obtain a pre-trained video feature extraction model.

[0135] In this step, the preset classification label is the video classification text pre-screened, and the video feature extraction model to be trained is a pre-built neural network model, which is used for video feature extraction. The third loss value is determined using the video features in the pre-training and the preset classification label. Specifically, the loss function can be a cross-entropy function. This third loss value can accurately represent the difference between the video features in the pre-training and the preset classification label. Based on the third loss value, the model parameters of the video feature extraction model to be trained are adjusted to obtain a pre-trained video feature extraction model. In this embodiment, the training resources can be fully utilized to train the video feature extraction model, and a video feature extraction model with high performance can be trained.

[0136] Figure 2 It is a flowchart of a video text retrieval method shown according to an exemplary embodiment. The method includes the following steps:

[0137] In step 21, retrieve data to be retrieved is obtained. The retrieve data to be retrieved is video retrieval data or text retrieval data.

[0138] In step 22, the retrieve data to be retrieved is input into the video text retrieval model to obtain target retrieval data; when the retrieve data to be retrieved is video retrieval data, the target retrieval data is target text data; when the retrieve data to be retrieved is text retrieval data, the target retrieval data is target video data.

[0139] Among them, the video text retrieval model is obtained according to the above training method of the video text retrieval model.

[0140] In this embodiment, the data to be retrieved is the input content of the user. There are two possible forms of the data to be retrieved, namely video retrieval data and text retrieval data. By inputting the data to be retrieved into the video-text retrieval model, target retrieval data with a modality different from that of the data to be retrieved can be obtained. That is to say, when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data.

[0141] In a possible implementation manner, when the data to be retrieved is video retrieval data, the data to be retrieved is input into the video feature extraction model in the video-text retrieval model to obtain a first video feature. The candidate text data is input into the text feature extraction model in the video-text retrieval model to obtain a first text feature. Based on the similarity information between the first video feature and the first text feature, the target retrieval data is determined from the candidate text data. For example, the candidate text data with similarity information greater than a set similarity threshold or the candidate text data with the maximum similarity information is used as the target retrieval data, so as to retrieve text data using video data.

[0142] In a possible implementation manner, when the data to be retrieved is text retrieval data, the data to be retrieved is input into the text feature extraction model in the video-text retrieval model to obtain a second text feature. The candidate video data is input into the video feature extraction model in the video-text retrieval model to obtain a second video feature. Based on the similarity information between the second text feature and the second video feature, the target retrieval data is determined from the candidate video data. For example, the candidate video data with similarity information greater than a set similarity threshold or the candidate video data with the maximum similarity information is used as the target retrieval data, so as to retrieve video data using text data.

[0143] Figure 3 It is a block diagram of a training device for a video-text retrieval model shown according to an exemplary embodiment. The device includes a data acquisition module 31, a first extraction module 32, a second extraction module 33, a loss determination module 34, and a model acquisition module 35.

[0144] The data acquisition module 31 is configured to acquire a first video sample and the description text data corresponding to the first video sample;

[0145] The first extraction module 32 is configured to input the first video sample into a pre-trained video feature extraction model to obtain a first video feature;

[0146] The second extraction module 33 is configured to input the description text data corresponding to the first video sample into a text feature extraction model to be trained to obtain a first text feature;

[0147] A loss determination module 34, configured to determine, from the first video feature and the first text feature, a second video feature and a second text feature derived from the same first video sample, and a third video feature and a third text feature derived from different first video samples; determine at least one first loss value according to the second video feature and the second text feature, and determine at least one second loss value according to the third video feature and the third text feature; adjust model parameters of the text feature extraction model to be trained based on the at least one first loss value and the at least one second loss value, to obtain a trained text feature extraction model;

[0148] A model acquisition module 35, configured to use the pre-trained video feature extraction model and the trained text feature extraction model as a video-text retrieval model, where the video-text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data.

[0149] In an exemplary embodiment of the present disclosure, the loss determination module includes:

[0150] A first similarity determination unit, configured to determine a first video-text similarity between the second video feature and the second text feature, and a first text-video similarity between the second text feature and the second video feature; the first video-text similarity is used to represent a result of comparing the similarity between the second video feature and the second text feature with the second video feature first; the first text-video similarity is used to represent a result of comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0151] A first loss value determination unit, configured to determine a first video-text loss value based on the first video-text similarity and a first preset video-text label; determine a first text-video loss value based on the first text-video similarity and a first preset text-video label; the first preset video-text label is used to represent annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first;

[0152] A first loss value selection unit, configured to use at least one of the first video-text loss value and the first text-video loss value as the first loss value.

[0153] In an exemplary embodiment of the present disclosure, the first similarity determination unit includes:

[0154] The first regularization subunit is configured to perform regularization processing on the second video feature to obtain the regularized second video feature;

[0155] The second regularization subunit is configured to perform regularization processing on the second text feature to obtain the regularized second text feature;

[0156] The similarity determination subunit is configured to perform dot multiplication on the regularized second video feature and the regularized second text feature to determine the first video-text similarity between the second video feature and the second text feature; perform dot multiplication on the regularized second text feature and the regularized second video feature to determine the first text-video similarity between the second text feature and the second video feature.

[0157] In an exemplary embodiment of the present disclosure, the loss determination module includes:

[0158] The second similarity determination unit is configured to determine the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature; the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0159] The second loss value determination unit is configured to determine the second video-text loss value based on the second video-text similarity and the second preset video-text label; determine the second text-video loss value based on the second text-video similarity and the second preset text-video label; the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first;

[0160] The second loss value selection unit is configured to use at least one of the second video-text loss value and the second text-video loss value as the second loss value.

[0161] In an exemplary embodiment of the present disclosure, the device further includes: a pre-training module, and the pre-training module includes:

[0162] The data acquisition unit is configured to acquire the image data, content text data, and preset classification label corresponding to the second video sample respectively;

[0163] A feature extraction unit, configured to input the image data and the content text data into a video feature extraction model to be trained, and obtain video features during pre-training;

[0164] A training processing unit, configured to adjust model parameters of the video feature extraction model to be trained based on a third loss value corresponding to the video features during pre-training and the preset classification label, and obtain a pre-trained video feature extraction model.

[0165] In an exemplary embodiment of the present disclosure, the feature extraction unit includes:

[0166] An extraction processing subunit, configured to input the image data and the content text data into a video feature extraction model to be trained, and obtain image features and text features during pre-training;

[0167] A fusion processing subunit, configured to perform feature fusion on the image features and the text features during pre-training to obtain video features during pre-training.

[0168] In an exemplary embodiment of the present disclosure, the data acquisition unit includes:

[0169] An identification processing subunit, configured to obtain a video speech recognition result and a video image text recognition result of each second video sample;

[0170] A data acquisition subunit, configured to obtain content text data of the second video sample based on the video speech recognition result and the video image text recognition result.

[0171] Figure 4 It is a block diagram of a video text retrieval device shown according to an exemplary embodiment. The device includes a data acquisition module and a data retrieval module.

[0172] The data acquisition module 41 is configured to acquire data to be retrieved, and the data to be retrieved is video retrieval data or text retrieval data;

[0173] The data retrieval module 42 is configured to input the data to be retrieved into a video text retrieval model to obtain target retrieval data; when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data;

[0174] Wherein, the video text retrieval model is obtained according to the training method of the video text retrieval model described above.

[0175] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0176] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. The electronic device can be a server or a similar computing device. Referring to Figure 5 , the electronic device 500 includes a processing component 522, which further includes one or more processors, and memory resources represented by a memory 532 for storing instructions executable by the processing component 522, such as application programs. The application programs stored in the memory 532 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 522 is configured to execute instructions to perform the above-described training method or video text retrieval method of the video text retrieval model.

[0177] The electronic device 500 may further include a power supply component 526 configured to perform power management of the electronic device 500, a wired or wireless network interface 550 configured to connect the electronic device 500 to a network, and an input / output (I / O) interface 558. The electronic device 500 can operate based on an operating system stored in the memory 532, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0178] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as the memory 532 including instructions, and the above instructions can be executed by the processing component 522 of the electronic device 500 to complete the above-described training method or video text retrieval method of the video text retrieval model. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0179] In an exemplary embodiment, there is also provided a computer program product including a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the above-described training method or video text retrieval method of the video text retrieval model is implemented.

[0180] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0181] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A training method for a video text retrieval model, characterized in that The method includes the following steps: Obtain a first video sample and the description text data corresponding to the first video sample; Input the first video sample into a pre-trained video feature extraction model to obtain a first video feature; Input the description text data corresponding to the first video sample into a text feature extraction model to be trained to obtain a first text feature; Among the first video feature and the first text feature, determine a second video feature and a second text feature that originate from the same first video sample, and a third video feature and a third text feature that originate from different first video samples; determine at least one first loss value according to the second video feature and the second text feature, and determine at least one second loss value according to the third video feature and the third text feature; Based on the at least one first loss value and the at least one second loss value, adjust the model parameters of the text feature extraction model to be trained to obtain a trained text feature extraction model; Use the pre-trained video feature extraction model and the trained text feature extraction model as a video-text retrieval model, where the video-text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data; Before the step of inputting the first video sample into the pre-trained video feature extraction model to obtain a first video feature, the method further includes: Obtain the image data, content text data, and preset classification labels respectively corresponding to a second video sample; Input the image data and the content text data into a video feature extraction model to be trained to obtain a video feature during pre-training; Based on the third loss value corresponding to the video feature during pre-training and the preset classification label, adjust the model parameters of the video feature extraction model to be trained to obtain a pre-trained video feature extraction model.

2. The method according to claim 1, wherein The at least one first loss value includes a first video-text loss value and a first text-video loss value. The determining of at least one first loss value according to the second video feature and the second text feature includes: Determine a first video-text similarity between the second video feature and the second text feature, and a first text-video similarity between the second text feature and the second video feature; the first video-text similarity is used to represent the result of comparing the similarity between the second video feature and the second text feature with the second video feature first; the first text-video similarity is used to represent the result of comparing the similarity between the second text feature and the second video feature with the second text feature first; Determine a first video-text loss value based on the first video-text similarity and the first preset video-text label; determine a first text-video loss value based on the first text-video similarity and the first preset text-video label; the first preset video-text label is used to represent the annotation data for comparing the similarity between the second video feature and the second text feature with the second video feature first; the first preset text-video label is used to represent the annotation data for comparing the similarity between the second text feature and the second video feature with the second text feature first; Use at least one of the first video-text loss value and the first text-video loss value as the first loss value.

3. The method according to claim 2, wherein The determining of the first video-text similarity between the second video feature and the second text feature, and the first text-video similarity between the second text feature and the second video feature includes: Perform regularization processing on the second video feature to obtain the regularized second video feature; Perform regularization processing on the second text feature to obtain the regularized second text feature; Perform a dot product on the regularized second video feature and the regularized second text feature to determine the first video-text similarity between the second video feature and the second text feature; perform a dot product on the regularized second text feature and the regularized second video feature to determine the first text-video similarity between the second text feature and the second video feature.

4. The method according to claim 1, wherein The at least one second loss value includes a second video-text loss value and a second text-video loss value. The determining of at least one second loss value according to the third video feature and the third text feature includes: Determine the second video-text similarity between the third video feature and the third text feature, and the second text-video similarity between the third text feature and the third video feature; the second video-text similarity is used to represent the result of comparing the similarity between the third video feature and the third text feature with the third video feature first; the second text-video similarity is used to represent the result of comparing the similarity between the third text feature and the third video feature with the third text feature first; Based on the second video-text similarity and the second preset video-text label, determine a second video-text loss value; based on the second text-video similarity and the second preset text-video label, determine a second text-video loss value; the second preset video-text label is used to represent the annotation data for comparing the similarity between the third video feature and the third text feature with the third video feature first; the second preset text-video label is used to represent the annotation data for comparing the similarity between the third text feature and the third video feature with the third text feature first; Use at least one of the second video-text loss value and the second text-video loss value as the second loss value.

5. The method according to claim 1, characterized in that Inputting the image data and the content text data into a video feature extraction model to be trained to obtain video features in pre-training includes: Inputting the image data and the content text data into a video feature extraction model to be trained to respectively obtain corresponding image features and text features in pre-training; Performing feature fusion on the image features and the text features in pre-training to obtain video features in pre-training.

6. The method according to claim 1, wherein Obtaining content text data respectively corresponding to second video samples includes: Obtaining the video speech recognition result and the video image text recognition result of each second video sample; Based on the video speech recognition result and the video image text recognition result, obtaining the content text data of the second video sample.

7. A video text retrieval method, characterized in that, Including: Obtaining data to be retrieved, where the data to be retrieved is video retrieval data or text retrieval data; Inputting the data to be retrieved into a video text retrieval model to obtain target retrieval data; when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data; Wherein, the video text retrieval model is obtained according to the training method of the video text retrieval model according to any one of claims 1-6.

8. A training device for a video text retrieval model, characterized in that Including the following modules: A data acquisition module configured to acquire a first video sample and description text data corresponding to the first video sample; A first extraction module configured to input the first video sample into a pre-trained video feature extraction model to obtain first video features; A second extraction module configured to input the description text data corresponding to the first video sample into a text feature extraction model to be trained to obtain first text features; A loss determination module configured to determine, from the first video features and the first text features, second video features and second text features originating from the same first video sample, and third video features and third text features originating from different first video samples; determining at least one first loss value according to the second video features and the second text features, and determining at least one second loss value according to the third video features and the third text features; Adjusting the model parameters of the text feature extraction model to be trained based on the at least one first loss value and the at least one second loss value to obtain a trained text feature extraction model; A model acquisition module configured to use the pre-trained video feature extraction model and the trained text feature extraction model as a video text retrieval model, where the video text retrieval model is used to retrieve target text data based on video retrieval data and retrieve target video data based on text retrieval data; The apparatus further includes: a pre-training module, and the pre-training module includes: A data acquisition unit configured to acquire image data, content text data, and a preset classification label respectively corresponding to a second video sample; A feature extraction unit, configured to input the image data and the content text data into a video feature extraction model to be trained, and obtain video features during pre-training; A training processing unit, configured to adjust model parameters of the video feature extraction model to be trained based on a third loss value corresponding to the video features during pre-training and the preset classification label, and obtain a pre-trained video feature extraction model.

9. A video text retrieval device, characterized in that, Comprising: A data acquisition module, configured to acquire data to be retrieved, where the data to be retrieved is video retrieval data or text retrieval data; A data retrieval module, configured to input the data to be retrieved into a video-text retrieval model to obtain target retrieval data; when the data to be retrieved is video retrieval data, the target retrieval data is target text data; when the data to be retrieved is text retrieval data, the target retrieval data is target video data; Wherein, the video-text retrieval model is obtained according to the training method of the video-text retrieval model according to any one of claims 1-6.

10. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the training method of the video-text retrieval model according to any one of claims 1 to 6, or the video-text retrieval method according to claim 7.

11. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of the video-text retrieval model according to any one of claims 1 to 6, or the video-text retrieval method according to claim 7.

Citation Information

Patent Citations

  • Pre-training language model generation method and device, electronic equipment and storage medium

    CN113705187A

  • Video processing method and apparatus, and storage medium and device

    WO2022171067A1