Text-to-video model training method, text-to-video generation method and device

By introducing a video detection model into the text-generated video model for detection and training, the problem of generated videos not conforming to text descriptions is solved, thereby improving video quality and satisfaction.

CN119516419BActive Publication Date: 2026-01-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410667838.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2026-01-13
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

In existing text-to-video technology, the generated video may not match the description of the input text, resulting in low video quality.

Method used

By inputting descriptive text into the text-generated video model, and using a video detection model to detect the generated video, the text-generated video model is trained and optimized based on the detection results to improve video quality.

Benefits of technology

It improves the relevance and aesthetics of the generated videos, thereby increasing video satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516419B_ABST
    Figure CN119516419B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text-to-video model training method and a text-to-video method and device, and relates to the technical field of computers, in particular to the fields of artificial intelligence such as large models, reinforcement learning, supervised learning, etc. The specific implementation scheme is as follows: inputting a description text into a first text-to-video model to obtain a to-be-detected video; inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video; training the first text-to-video model according to the detection result and a training sample to obtain a second text-to-video model. According to the present disclosure, the video detection model can detect the to-be-detected video generated by the text-to-video model, and then the text-to-video model is retrained according to the detection result, so as to improve the quality of the video generated by the model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of computer technology, and particularly to the fields of artificial intelligence such as large models, reinforcement learning, supervised learning, etc. BACKGROUND

[0002] Text-to-video (T2V) technology is a technology that uses artificial intelligence to convert text content into visual and auditory effects. In T2V technology, a video can be generated by inputting a text into a model. However, the video generated by the model may not conform to the description of the text. SUMMARY

[0003] The present disclosure provides a text-to-video model training method, a text-to-video generation method, an apparatus, a device, and a storage medium.

[0004] According to an aspect of the present disclosure, a text-to-video model training method is provided, comprising:

[0005] inputting a description text into a first text-to-video model to obtain a to-be-detected video;

[0006] inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video;

[0007] training the first text-to-video model according to the detection result to obtain a second text-to-video model.

[0008] According to another aspect of the present disclosure, a text-to-video generation method is provided, comprising:

[0009] inputting a description text into a text-to-video model to obtain a to-be-detected video; wherein the text-to-video model is trained based on the text-to-video model training method described above.

[0010] According to another aspect of the present disclosure, a text-to-video model training apparatus is provided, comprising:

[0011] a text processing module configured to input a description text into a first text-to-video model to obtain a to-be-detected video;

[0012] a video processing module configured to input the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video;

[0013] a training module configured to train the first text-to-video model according to the detection result to obtain a second text-to-video model.

[0014] According to another aspect of the present disclosure, a text-to-video generation apparatus is provided, comprising:

[0015] The text processing module is configured to input the description text into the text-to-video model to obtain a to-be-detected video.

[0016] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0017] at least one processor; and

[0018] a memory in communication with the at least one processor; wherein

[0019] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any embodiment of the present disclosure.

[0020] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to make the computer perform the method according to any embodiment of the present disclosure.

[0021] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any embodiment of the present disclosure.

[0022] The embodiments of the present disclosure can detect the to-be-detected video generated by the text-to-video model through the video detection model, and then retrain the text-to-video model according to the detection result, thereby improving the quality of the video generated by the model.

[0023] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:

[0025] Figure 1 is a flowchart of a text-to-video model training method according to an embodiment of the present disclosure;

[0026] Figure 2 is a flowchart of a text-to-video model training method according to another embodiment of the present disclosure;

[0027] Figure 3 is a flowchart of a text-to-video model training method according to another embodiment of the present disclosure;

[0028] Figure 4is a flowchart of a text-to-video model training method according to another embodiment of the present disclosure;

[0029] Figure 5 is a general flowchart of optimizing a text-to-video model based on reinforcement learning;

[0030] Figure 6 is a general flowchart of optimizing a text-to-video model and a rewriting model based on reinforcement learning;

[0031] Figure 7 is a schematic block diagram of a text-to-video model training apparatus according to an embodiment of the present disclosure;

[0032] Figure 8 is a schematic block diagram of a text-to-video model training apparatus according to another embodiment of the present disclosure;

[0033] Figure 9 is a schematic block diagram of a text-to-video generation apparatus according to an embodiment of the present disclosure;

[0034] Figure 10 is a schematic block diagram of a text-to-video generation apparatus according to another embodiment of the present disclosure;

[0035] Figure 11 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0037] The "and / or" of the embodiments of the present disclosure means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first", "second", etc. herein mean to refer to and distinguish a plurality of similar technical terms, and do not mean to limit the order or mean to limit to only two, for example, a first feature and a second feature mean to refer to two categories / two features, and the first feature can be one or more and the second feature can also be one or more.

[0038] The training process of the text-to-video model includes training of a text-side model and training of a video-side model. For example, the text-side can use a pre-trained Transformer model such as Text to Text Transfer Transformer (T5), Generative Pre-Training (GPT), etc. to model the text input. The video-side model structure can use a Diffusion Transformer structure and be modeled as an input of length L according to the Tuplet splitting method of Video Vision Transformer (ViViT). The training method can be a diffusion modeling process, which can enable the model to be trained to restore the original video based on text guidance and according to random noise. The text-side model and the video-side model are associated through a cross-attention mechanism.

[0039] Figure 1 is a flowchart of a text-to-video model training method according to an embodiment of the present disclosure, including:

[0040] S110, inputting the description text into a first text-to-video model to obtain a to-be-detected video;

[0041] S120, inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video;

[0042] S130, training the first text-to-video model according to the detection result to obtain a second text-to-video model.

[0043] According to an embodiment of the present disclosure, the description text can be input text or text sampled from historical data. The first text-to-video model can be a text-to-video model that needs to be trained or a pre-trained text-to-video model. The text-to-video model can be a large model or a model dedicated to text-to-video. For example, Sora, Pika, Stable Video Diffusion, Transformer, etc. The first text-to-video model can be pre-trained using supervised learning, and the first text-to-video model can be used as a hot start model for further reinforcement training to obtain the second text-to-video model.

[0044] According to an embodiment of the present disclosure, after the description text is input into the first text-to-video model, a plurality of random videos, i.e., a plurality of to-be-detected videos, can be generated using different random seeds. For example, the different random seeds can include different dropout layer (Dropout) positions in a Transformer model. Dropout is a regularization method in a neural network, which can randomly delete some neurons during the training process of the model. Other schemes can also be used to generate the plurality of to-be-detected videos, for example, using uncertainty sampling when the text-to-video model performs result sampling to obtain the plurality of to-be-detected videos.

[0045] According to an embodiment of the present disclosure, the video detection model can be pre-trained as a reward model of the pre-trained first text-to-video model. The video detection model can be trained using a model such as ViViT, Encoder, Convolutional Neural Network (CNN), or Transformer. The training data of the video detection model can include training samples related to video quality, such as video text relevance, video aesthetics, video clarity, and the like.

[0046] According to an embodiment of the present disclosure, the second text-to-video model can be obtained by iteratively training the first text-to-video model according to the detection results output by the video detection model. For example, the training process of the text-to-video model can use a Reinforcement Learning with human feedback (RLHF) method to iteratively train the text-to-video model. RLHF is a method of training artificial intelligence systems by combining reinforcement learning with human feedback. By using human feedback to create a reward signal, the signal is then used to improve the optimization function of the model through reinforcement learning, so that the model can better capture complex human preferences.

[0047] According to an embodiment of the present disclosure, if the description text is input into the first text-to-video model, a plurality of to-be-detected videos can be obtained, each to-be-detected video can be input into the video detection model to obtain a corresponding detection result. Then, according to the detection result corresponding to each to-be-detected video, a corresponding optimization function can be obtained, or an optimization function can be obtained comprehensively. Then, the parameters of the first text-to-video model are adjusted according to the optimization function. The model after each adjustment is used as the first text-to-video model for the next training, and the second text-to-video model can be obtained after the optimization function converges.

[0048] Through the embodiments of the present disclosure, the video detection model can be used to detect the to-be-detected video generated by the text-to-video model, and then the text-to-video model is retrained according to the detection result, so as to improve the quality of the video generated by the model.

[0049] Figure 2 is a flowchart of a text-to-video model training method according to another embodiment of the present disclosure. The method can include one or more features of the above method embodiments. In an implementation, the video detection model includes a video-text relevance model and a video aesthetic model; step S120 inputs the to-be-detected video into the video detection model to obtain a video detection result corresponding to the to-be-detected video, including:

[0050] S210, input the description text and the to-be-detected video into the video-text relevance model to obtain the relevance of the to-be-detected video; and / or

[0051] S220, input the to-be-detected video into the video aesthetic model to obtain the aesthetic degree of the to-be-detected video.

[0052] According to the embodiments of the present disclosure, the video detection model can be used as a reward model of the text-to-video model.

[0053] In some examples, the video detection model can include a video-text relevance model. The video-text pair can be used as the training data of the video-text relevance model. The text end can be modeled using structures such as Bidirectional Encoder Representation from Transformers (BERT), and the video end can be modeled through Encoder structures such as ViViT, CNN or Transformer. The classification (CLS) feature vectors of the video end and the text end are extracted respectively as the classification vectors of the video end and the text end, and contrastive learning training is performed. For example, a certain video A and its corresponding text are used as positive examples for contrastive training, indicating that the relevance between video A and the text is high. Samples in the same batch that are not related to video A are used as negative examples, and the negative examples indicate that video A has no relevance with the text. One or more to-be-detected videos and description texts are input into the trained video-text relevance model, and the cosine similarity between the CLS classification of the video end and the CLS classification of the text end is obtained. The relevance of the to-be-detected video is obtained. The relevance can be represented by a score, so it can be called a relevance score. For example, the relevance score of video A1 and text T1 is 0.7, and the relevance score of video A2 and text T1 is 0.5.

[0054] In some examples, the video detection model can include a video aesthetics model. The video aesthetics model can be modeled using an Encoder structure such as ViViT. Video end CLS classification can be used to represent regression prediction of the aesthetics score. By inputting one or more videos to be detected into the trained video aesthetics model, the aesthetics of each video to be detected can be obtained. The aesthetics can be represented by a score, so it can be referred to as an aesthetics score. For example, the aesthetics score of video A1 is 8, and the aesthetics score of video A2 is 9.

[0055] If the text-to-video model outputs multiple videos to be detected according to the description text, the multiple videos to be detected and the description text can be combined into multiple text-video pairs respectively. By inputting these multiple text-video pairs into the trained video-text relevance model, the video-text relevance of each video to be detected can be obtained. Moreover, the multiple videos to be detected can be input into the trained video aesthetics model to obtain the aesthetics of each video to be detected. In S130, the first text-to-video model is trained according to the relevance and aesthetics of the video to be detected to obtain the second text-to-video model. For example, an optimization function can be calculated according to the relevance and aesthetics of each video to be detected, and the parameters of the first text-to-video model are adjusted according to the optimization function. The model after each adjustment is used as the first text-to-video model for the next training, and the second text-to-video model can be obtained after the optimization function converges.

[0056] By the embodiments of the present disclosure, the relevance and / or the aesthetics of the video and the text generated by the text-to-video model can be detected, the video with high relevance and / or high aesthetics can be selected for presentation, and the satisfaction with the video generated by the model can be improved.

[0057] In an embodiment, S130 trains the first text-to-video model according to the detection result to obtain a second text-to-video model, including: adjusting parameters of the first text-to-video model according to a first optimization function to obtain the second text-to-video model; wherein the first optimization function is determined according to the detection result, a training sample, and a first probability of generating the to-be-detected video based on the description text. This step can be performed multiple times, and each time the calculation result of the optimization function can be used to determine whether the model needs to be updated. If the model needs to be updated, the parameters of the first text-to-video model are adjusted. The adjusted model is used as the first text-to-video model for the next training, and the training continues. If the model does not need to be updated, the last updated model can be used as the trained second text-to-video model. The first optimization function includes the first probability of generating the to-be-detected video based on the description text. Training the text-to-video model by the first optimization function can improve the probability of generating a video with high aesthetic score and similarity score, and reduce the probability of generating a video with low aesthetic score and similarity score. The first probability of generating the to-be-detected video based on the description text can be an output result of the text-to-video model. For example, the text-to-video model can output: video A and its corresponding first probability, video B and its corresponding first probability, based on the input text T.

[0058] In an embodiment, the training sample includes a reinforcement learning sample and a supervised learning sample, the reinforcement learning sample including unlabeled text, and the supervised learning sample including labeled data of text and video.

[0059] According to the embodiments of the present disclosure, the training sample used in the first optimization function can include a reinforcement learning sample and a supervised learning sample. An example of a first optimization function is shown as follows:

[0060]

[0061] wherein x represents a video generated by the first text-to-video model; z represents a description text input into the first text-to-video model; D unlabeled represents a sample for reinforcement learning (i.e., a reinforcement learning sample, which does not require corresponding video labeling); D labeled represents a sample for supervised learning (i.e., a supervised learning sample, which is used to keep the output result of the model stable for the supervised sample). θ represents a model parameter to be optimized; p θ (x|z) represents a probability of generating x based on z under the current model (i.e., the first probability). r simφ (x, z) represents a correlation score output by a video-text correlation model, r aesφ (x) represents an aesthetic score output by a video aesthetic model. γ and β are both adjustment coefficients. The average value is used to determine the overall optimization target. The adjustment coefficient is generally a hyperparameter of the model training process.

[0062] D unlabeled The sample can include a text sample, such as text in historical data, and z can be sampled from D unlabeled D labeled The sample can include a video-text pair, which can be training data used by a pre-trained first text-to-video model, data sampled from the training data, or other supervised learning samples.

[0063] According to an embodiment of the present disclosure, the training sample, the first probability, and the detection result output by the video detection model are substituted into the first optimization function to calculate a value of the first optimization function. Whether the model needs to be updated can be determined according to the value of the first optimization function. For example, in the case where the value of the first optimization function does not converge, the model needs to be updated. In the case where the value of the first optimization function converges, a second text-to-video model can be obtained.

[0064] According to an embodiment of the present disclosure, the text-to-video model can be iteratively trained by the first optimization function, so as to increase the probability of generating a video with a high detection result score and decrease the probability of generating a video with a low detection result score, thereby improving the satisfaction of the generated video.

[0065] In an implementation, as shown in Figure 3 The method further includes:

[0066] S310, inputting the initial text into the rewriting model to obtain a description text.

[0067] According to an embodiment of the present disclosure, the initial text can be an input text or a text sampled from historical data. The rewriting model can be added to rewrite the initial text, for example, to modify the syntax, logic, and the like of the initial text, to expand according to the keywords of the initial text, and the like. In this way, the rewriting model can be used to generate a more detailed, more reasonable, and more fluent description text, to improve the rationality, accuracy, or richness of the input text of the text-to-video model, and the like, to input the rewritten description text into the text-to-video model, to improve the model training efficiency, and to improve the quality of the video output by the model.

[0068] In an implementation, according to the detection result, the first text-to-video model is trained to obtain a second text-to-video model, including: according to the detection result, the first text-to-video model and the rewriting model are jointly trained to obtain a second text-to-video model and an updated rewriting model.

[0069] According to an embodiment of the present disclosure, the rewriting model can be trained using the training data of the first text-to-video model. The rewriting model can be modeled using a model such as a Transformer model.

[0070] According to the embodiments of the present disclosure, the rewriting model can be pre-trained, and the training samples of the rewriting model can be the same as the training samples of the text-to-video model on the text side, or other training samples can be used. After the initial text is input into the pre-trained rewriting model, the rewriting model can output the rewritten description text. After the description text is input into the first text-to-video model, the first text-to-video model can output a plurality of to-be-detected videos. After the plurality of to-be-detected videos are input into the video detection model respectively, the video detection model can output a detection result of each to-be-detected video. According to the detection result output by the video detection model, an optimization function can be calculated, and according to the optimization function, the parameters of the first text-to-video model and the rewriting model can be modified. For example, according to the correlation score output by the video text correlation model and the aesthetic score output by the video aesthetic model, the total score of the detection result of the video can be obtained. The total score can be equal to the sum or weighted sum of the correlation score and the aesthetic score. According to the total score, an optimization function for joint training can be calculated. The order of modifying the parameters of the first text-to-video model and the rewriting model using the optimization function is not limited, which can be processed in parallel or in sequence. Joint training of the two models using the same detection result can improve the training efficiency and improve the correlation between the text rewritten by the rewriting model and the video generated by the text-to-video model.

[0071] In an embodiment, according to the detection result, the first text-to-video model and the rewriting model are jointly trained to obtain a second text-to-video model and an updated rewriting model, including: adjusting the parameters of the first text-to-video model according to a second optimization function to obtain the second text-to-video model and the updated rewriting model; wherein the second optimization function is based on the detection result, the training sample, a second probability of generating the to-be-detected video based on the description text, and a third probability of generating the rewritten text based on the description text. This step can be performed multiple times, each time first determining whether the model needs to be updated according to the calculation result of the second optimization function, and if the model needs to be updated, adjusting the parameters of the first text-to-video model and / or the rewriting model. The adjusted model is used as the first text-to-video model and / or the rewriting model for the next training to continue the training. If the model does not need to be updated, the last updated text-to-video model can be used as the trained second text-to-video model, and / or the last updated rewriting model can be used as the trained rewriting model.

[0072] According to an embodiment of the present disclosure, the second probability of generating a video based on the description text can be an output result of the text-to-video model. For example, the text-to-video model can output video A and its corresponding second probability, video B and its corresponding second probability based on the input text T. The third probability of generating a rewritten text based on the description text can be an output result of the rewriting model. For example, the rewriting model can output the rewritten text T2 and its corresponding third probability based on the input text T1.

[0073] According to an embodiment of the present disclosure, the training samples used in the second optimization function can include reinforcement learning samples and supervised learning samples. An example of jointly training the second optimization function is shown as follows:

[0074]

[0075] wherein x represents a video generated by the second text-to-video model, z represents a description text of the second text-to-video model, D unlabeled represents a sample for reinforcement learning (without corresponding video annotation), D labeled represents a sample for supervised learning (used to keep the output result of the model stable for supervised samples). θ represents model parameters to be optimized; z rew represents a rewritten description text of the rewriting model; z ori represents an initial text; p θ (x|z rew ) represents a probability of generating x based on the rewritten description text z rew under the current model (i.e., the second probability); p θ (z rew |z ori ) represents a probability of generating the rewritten description text z ori based on the description text z rew under the current model (i.e., the third probability); r φ (x,z ori ) represents a score weighted by the aesthetic degree model and the video-text relevance model; r simφ (x|z ori ) represents a relevance score of generating a video x through an initial text; r simφ (x,z ori ) represents a relevance score output by the video-text relevance model; r aesφ (x) represents a score of the video aesthetic degree model. γ and β are both adjustment coefficients. represents an average value. The adjustment coefficient is generally a hyperparameter in the model training process.

[0076] According to the embodiment of the present disclosure, the same training sample is used to calculate the optimization function by the second optimization function iterative training of the text-to-video model and the rewriting model, which can ensure that the rewriting model and the text-to-video model do not deviate from the original intention of the user in the process of joint training, improve the probability of generating a video with a high detection result score, and reduce the probability of generating a video with a low detection result score, thereby improving the satisfaction of the generated video.

[0077] In some examples, after the rewriting model is added, only the text-to-video model can be trained, and the rewriting model is not trained. Specifically, after the description text of the initial text is rewritten by the rewriting model, the description text is input into the first text-to-video model to generate a to-be-detected video. The initial text and the to-be-detected video are input into the video-text correlation model to obtain a correlation score, and the to-be-detected video is input into the aesthetic degree model to obtain an aesthetic degree score. Then, the first optimization function described above is calculated by using the correlation score and the aesthetic degree score, and the first text-to-video model is trained by using the first optimization function to obtain a second text-to-video model.

[0078] In an implementation, as shown in Figure 3 , further comprising:

[0079] S320, labeling the to-be-detected video to obtain a sample for updating the video detection model.

[0080] According to the embodiment of the present disclosure, the detection result of the to-be-detected video sample output by the text-to-video model can be labeled. For example, the videos output by the text-to-video model include videos A, B and C, wherein the aesthetic degree of video A is labeled as 8, the aesthetic degree of video B is labeled as 7, and the aesthetic degree of video C is labeled as 6. The specific labeling basis can refer to the output result of the video aesthetic degree model, or can be labeled according to experience. For example, the output result of the video A input into the video aesthetic degree model is 7.5, and the aesthetic degree of the video A can be labeled as 8. For another example, the aesthetic degree of the video B is labeled as 7 according to experience. The clarity, length, content, etc. of the video can also be referred to for labeling the to-be-detected video. The labeled sample can be used as a new training sample of the video detection model; the video detection model is further trained, the video detection model is updated, the accuracy of the search result of the video detection model is improved, and then the training efficiency of the text-to-video model and / or the rewriting model and the quality of the generated video are improved.

[0081] Figure 4 The method for generating a video based on text according to an embodiment of the present disclosure can include:

[0082] S410, inputting the description text into the text-to-video model to obtain a to-be-detected video; wherein the text-to-video model is trained based on the text-to-video model training method described above.

[0083] According to the embodiments of the present disclosure, a description text can be input into a trained text-to-video model, and the text-to-video model can generate one or more to-be-detected videos A, B, C, …, N according to the description text. The to-be-detected video can also be referred to as a candidate video. For example, the description text input into the model is T. The generated to-be-detected videos A, B, C, …, N can include one or more elements mentioned in the description text.

[0084] According to the embodiments of the present disclosure, inputting the description text into the text-to-video model trained in the above embodiments can generate a high-quality video related to the description text.

[0085] In an implementation, the method further includes:

[0086] S420, inputting the to-be-detected video into a video detection model to obtain a video detection result corresponding to the to-be-detected video.

[0087] According to the embodiments of the present disclosure, the video detection model can be used to detect one or more to-be-detected videos generated by the text-to-video model. The quality of the video can be detected in various ways. For example, a plurality of video detection models can be pre-trained to detect a plurality of detection results of the video, such as text-video relevance, video aesthetics, and video clarity.

[0088] In an implementation, the video detection model includes a text-video relevance model and a video aesthetics model, and step S420 inputs the to-be-detected video into the video detection model to obtain a detection result corresponding to the to-be-detected video, including:

[0089] inputting the description text and the to-be-detected video into the text-video relevance model to obtain relevance of the to-be-detected video; and / or

[0090] inputting the to-be-detected video into the video aesthetics model to obtain aesthetics of the to-be-detected video.

[0091] According to the embodiments of the present disclosure, one or more to-be-detected videos are respectively input into the text-video relevance model and / or the video aesthetics model, and the text-video relevance and the video aesthetics of each to-be-detected video can be obtained, which are used to represent the detection result of the to-be-detected video.

[0092] According to the embodiments of the present disclosure, using the text-video relevance score and the video aesthetics score as the evaluation standard of the to-be-detected video can improve the quality of the video generated by the model, and further improve the satisfaction.

[0093] In an implementation, the method further includes:

[0094] According to the detection result, the plurality of to-be-detected videos are sorted.

[0095] According to the embodiment of the present disclosure, if multiple detection results can be output by multiple video detection models, for example, multiple relevance scores of the multiple videos to be detected are output by the video text relevance model, and the aesthetic scores are output by the video aesthetic model. The multiple videos to be detected can be sorted from high to low according to the relevance scores, or sorted from high to low according to the aesthetic scores, or sorted after calculating the total scores by comprehensively considering the relevance scores and the aesthetic scores. Through the embodiment of the present disclosure, by using the video detection result as the sorting standard of the video, the high-quality video with better detection result can be preferentially displayed in the display interface.

[0096] In an implementation manner, the method further includes:

[0097] The initial text is rewritten into the description text by using the rewriting model.

[0098] According to the embodiment of the present disclosure, the initial text can be rewritten by using the rewriting model when the initial text cannot accurately and specifically generate the video. The rewriting model can also be used to rewrite each initial text. The rewritten text is used as the description text of the input text-to-video model. The rewriting model of the embodiment can use a pre-trained rewriting model, or use the updated rewriting model after joint training in the above embodiment.

[0099] Through the embodiment of the present disclosure, the rewriting model can generate more detailed, more reasonable, and more fluent description texts, and improve the rationality, accuracy, or richness of the text input into the text-to-video model. The quality of the video output by the text-to-video model can be improved by inputting the rewritten description text into the text-to-video model.

[0100] In some application scenarios, the video detection model can be trained as a reward model of the text-to-video model. The video detection model includes a video text relevance model and a video aesthetic model. The two models can be constructed by using techniques such as contrastive learning training, supervised learning training, and video classification (CLS).

[0101] For example, the training data used for training the video text relevance model is a video-text pair. The video end can be modeled by using a ViViT model (or other common video Encoder structure, such as CNN or other variants of Transformer), and the text end can be modeled by using a pre-trained BERT model. The video end CLS classification representation is compared with the text end CLS classification representation for contrastive learning training. The video and the text corresponding to the video are used as positive examples of the training sample, and the video and other texts in the same batch that do not correspond to the video are used as negative examples of the training sample. The Cosine similarity of the CLS of the video end and the CLS of the text end can be used as the relevance score of the video text relevance model.

[0102] For example, the data used for training the video aesthetics model is labeled data (for example, the data is in the form of a video and the aesthetics score of the video given by the labeler). The video aesthetics model can be modeled using a ViViT or other video Encoder structure. The video end CLS classification representation is used for regression prediction of the aesthetics score.

[0103] After the video text relevance model and the video aesthetics model are trained, they are used as reward models for training. The present scheme can use multi-reward model reinforcement learning for training.

[0104] Figure 5 is the overall flowchart of the reinforcement learning-based optimization of the text-to-video model. As shown in Figure 5 , the optimization process can also be understood as a training process. The main models used in the optimization process include the text-to-video model 501, the video text relevance model 502 (also referred to as the text video relevance model, etc.), and the video aesthetics model 503.

[0105] First, the text (request text) is input into the text-to-video model 501. The text-to-video model 501 can generate multiple random videos (or referred to as candidate videos) using different random seeds (for example, controlling different Dropout positions in the Transformer model). The multiple random videos generated by the text-to-video model 501 are input into the video text relevance model 502 and the video aesthetics model 503, respectively, to obtain the text video relevance score and the video aesthetics score. The sum (which can be weighted) of the relevance score and the aesthetics score is used as the overall generated video reward. During the training process, the RLHF training process can be used to iterate the text-to-video model.

[0106] An example of an optimization formula is as follows:

[0107]

[0108] where x represents the video output by the text-to-video model, and z represents the input text. D unlabeled represents the sample for reinforcement learning (i.e., the reinforcement learning sample, which does not require corresponding video labeling); D labeled represents the sample for supervised learning (used to keep the output of the model stable for supervised samples). θ represents the model parameters to be optimized, p θ (x|z) represents the probability of generating x based on z under the current model. r simφ (x, z) represents the text video model relevance score, r aesφ (x) represents the video aesthetics model score. γ and β are both adjustment coefficients.

[0109] The main purpose of the optimization of the model is to make the model generate multiple candidate videos through reinforcement learning, to increase the probability of generating videos with high aesthetic and similarity scores, and to reduce the probability of generating videos with low aesthetic and similarity scores.

[0110] As shown in Figure 6 , the overall process of optimizing the text-to-video model based on reinforcement learning can also be based on Figure 5 adding a rewriting model 601 to rewrite the input initial text to generate a more detailed and more suitable input text for the text-to-video model.

[0111] The rewriting model 601 can be trained using labeled data. The labeled data can include input content and rewritten content that is the same or similar in meaning to the input content. The rewritten content can make the text-to-video model generate better video results. The data is obtained by sampling the text-to-video model. If there is no labeled data, the same input (i.e., the same query before and after rewriting) can be used as the training data for the rewriting model. In one example, the rewriting model can use a Transformer model structure.

[0112] The process of reinforcement learning can optimize not only the text-to-video model itself, but also the rewriting model. An example of an optimization formula is as follows:

[0113]

[0114] where z rew represents the rewritten input content, z ori represents the input content before rewriting. x represents the video generated by the second text-to-video model, z represents the description text of the second text-to-video model, D unlabeled represents the sample for reinforcement learning, D labeled represents the sample for supervised learning. θ represents the model parameters to be optimized; p θ (x|z rew ) represents the probability of generating x based on the rewritten description text z rew under the current model; p θ (z rew |z ori ) represents the probability of generating the rewritten description text z ori based on the description text z rew under the current model; r φ (x,z ori ) represents the score of the aesthetic model and the video text relevance model; r simφ (x|z ori ) represents the relevance score of generating video x from the initial text; r simφ (x,z ori) represents the relevance score output by the video-text relevance model; r aesφ (x) represents the video aesthetics model score. Y and β are both adjustment coefficients. The optimization formula can be used to calculate the similarity score using the original input content and supervised training to ensure that the rewriting model does not deviate from the original intention of the user.

[0115] After introducing the rewriting model, the overall process of optimizing the text-to-video and rewriting model based on reinforcement learning includes: inputting the initial text to the rewriting model 601. The rewriting model rewrites the initial text to obtain a rewritten description text (which can be referred to as a rewritten text). The rewritten text is input to the text-to-video model 501, and a plurality of to-be-detected videos (or a plurality of candidate videos) can be generated. A plurality of video-text pairs composed of the plurality of to-be-detected videos and the initial text are input to the video-text relevance model 502, and the relevance score of each to-be-detected video can be obtained. The plurality of to-be-detected videos are input to the video aesthetics model 503, and the aesthetics score of the to-be-detected video can be obtained. The relevance score and the aesthetics score are gradient returned to the text-to-video model 501 and the rewriting model 601, and the text-to-video model 501 and the rewriting model 601 are trained.

[0116] In some examples, after the video generated by the text-to-video model is labeled, the text-video relevance model and / or the aesthetics model can be trained using the labeled data. That is, the model based on reinforcement learning is further reinforced by learning, realizing a circular iteration. The video detection model is trained as a reward model of the text-to-video model by using the reinforcement learning technology, so that the video generation process can consider the video-text similarity, video aesthetics and other goals, and the satisfaction of the generated video is higher.

[0117] Figure 7 The training device of the text-to-video model according to an embodiment of the present disclosure comprises:

[0118] The text processing module 701 is configured to input the description text into the first text-to-video model to obtain a to-be-detected video.

[0119] The video processing module 702 is configured to input the to-be-detected video into the video detection model to obtain a detection result corresponding to the to-be-detected video.

[0120] The training module 703 is configured to train the first text-to-video model according to the detection result to obtain a second text-to-video model.

[0121] In an embodiment, the video detection model comprises a video-text relevance model and a video aesthetics model, and the video processing module 702 is configured to:

[0122] inputting the description text and the to-be-detected video into the video text correlation model to obtain a correlation of the to-be-detected video; and / or

[0123] inputting the to-be-detected video into the video aesthetic degree model to obtain an aesthetic degree of the to-be-detected video.

[0124] In an embodiment, the training module 703 is configured to:

[0125] adjusting parameters of the first text-to-video model according to a first optimization function to obtain a second text-to-video model, wherein the first optimization function is determined according to the detection result, the training sample, and a first probability of generating the to-be-detected video based on the description text.

[0126] Figure 8 is a schematic block diagram of a training device of a text-to-video model according to another embodiment of the present disclosure. The device can include one or more features of the above-described device. In an embodiment, the device further includes:

[0127] The rewriting module 801 is configured to input the initial text into a rewriting model to obtain the description text.

[0128] In an embodiment, the training module 703 includes:

[0129] The joint training submodule 802 is configured to jointly train the first text-to-video model and the rewriting model according to the detection result to obtain a second text-to-video model and an updated rewriting model.

[0130] In an embodiment, the joint training submodule 802 is configured to:

[0131] adjusting parameters of the first text-to-video model according to a second optimization function to obtain the second text-to-video model and the updated rewriting model, wherein the second optimization function is determined according to the detection result, the training sample, a second probability of generating the to-be-detected video based on the description text, and a third probability of generating the rewritten text based on the description text.

[0132] In an embodiment, the training sample includes a reinforcement learning sample and a supervised learning sample, the reinforcement learning sample includes unlabeled text, and the supervised learning sample includes labeled data of text and video.

[0133] In an embodiment, the device further includes:

[0134] The labeling module 803 is configured to label the to-be-detected video to obtain a sample for updating the video detection model.

[0135] Figure 9is a schematic block diagram of an apparatus for generating a video based on text according to an embodiment of the present disclosure, comprising:

[0136] The text processing module 901 is configured to input the description text into a text-to-video model to obtain a to-be-detected video, wherein the text-to-video model is trained based on the text-to-video model training method in any one of claims 1 to 8.

[0137] Figure 10 is a schematic block diagram of an apparatus for generating a video based on text according to another embodiment of the present disclosure. The apparatus can include one or more features of the apparatus described above. In an implementation, the apparatus further comprises:

[0138] The video processing module 1001 is configured to input the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video.

[0139] In an implementation, the video detection model comprises a text-video correlation model and a video aesthetic degree model, and the video processing module 1001 comprises:

[0140] The correlation processing module 1002 is configured to input the description text and the to-be-detected video into the text-video correlation model to obtain a correlation of the to-be-detected video; and / or

[0141] The aesthetic degree processing module 1003 is configured to input the to-be-detected video into the video aesthetic degree model to obtain an aesthetic degree of the to-be-detected video.

[0142] In an implementation, the apparatus further comprises:

[0143] The sorting module 1004 is configured to sort the plurality of to-be-detected videos according to the detection result.

[0144] In an implementation, the apparatus further comprises:

[0145] The rewriting module 1005 is configured to rewrite an initial text into the description text using a rewriting model.

[0146] The specific functions and examples of each module and sub-module of the apparatus of the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, which will not be described here.

[0147] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0148] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0149] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0150] As shown, the device 1100 includes a computing unit 1101 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded into a random access memory (RAM) 1103 from a storage unit 1108. Various programs and data required for the operation of the device 1100 can also be stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104. Figure 11 Various components in the device 1100 are connected to the I / O interface 1105, including an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0151]

[0152] ​The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processors, controllers, microcontrollers, and the like. The computing unit 1101 performs various methods and processes described above, such as the text-to-video model training method or the text-to-video generation method. For example, in some embodiments, the text-to-video model training method or the text-to-video generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded onto the RAM 1103 and executed by the computing unit 1101, one or more steps of the text-to-video model training method or the text-to-video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 1101 can be configured to perform the text-to-video model training method or the text-to-video generation method by any other appropriate means, such as by means of firmware.

[0153] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0154] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0156] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0157] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0158] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers incorporating blockchain.

[0159] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technology disclosed in the present disclosure, and are not limited herein.

[0160] The specific embodiments described above are not intended to be limiting. One of skill in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments described above without departing from the principles of the present disclosure. Any such modifications, equivalents, and alternatives are intended to be included within the scope of the present disclosure.

Claims

1. A method for training a text-to-video model, comprising: inputting a description text into a first text-to-video model to obtain a to-be-detected video; inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video; training the first text-to-video model according to the detection result to obtain a second text-to-video model, comprising: adjusting parameters of the first text-to-video model according to a first optimization function to obtain the second text-to-video model; wherein the first optimization function is determined according to the detection result, a training sample, and a first probability of generating the to-be-detected video based on the description text; a formula of the first optimization function is as follows: wherein, x represents a video generated by the first text-to-video model; z represents a description text input into the first text-to-video model; D unlabeled represents a reinforcement learning sample; D labeled represents a supervised learning sample; θ represents model parameters to be optimized; p θ represents a first probability of generating x based on z under the current model; represents a correlation score output by the video-text correlation model, represents a beauty score output by the video beauty model; γ and β are both adjustment coefficients; represents an average value.

2. The method of claim 1, wherein, the video detection model comprises a video-text relevance model and a video-aesthetics model, and inputting the to-be-detected video into the video detection model to obtain the detection result corresponding to the to-be-detected video comprises: inputting the description text and the to-be-detected video into the video-text relevance model to obtain relevance of the to-be-detected video; and inputting the to-be-detected video into the video-aesthetics model to obtain aesthetics of the to-be-detected video.

3. The method of claim 1 or 2, wherein, the reinforcement learning sample comprises unlabeled text, and the supervised learning sample comprises labeled data of text and video.

4. The method of claim 1 or 2, wherein, further comprising: annotating the to-be-detected video to obtain a sample for updating the video detection model. 5.A method for training a text-to-video model, comprising: inputting an initial text into a rewriting model to obtain a description text; inputting the description text into a first text-to-video model to obtain a to-be-detected video; inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video; and jointly training the first text-to-video model and the rewriting model according to the detection result to obtain a second text-to-video model and an updated rewriting model, comprising: adjusting parameters of the first text-to-video model according to a second optimization function to obtain the second text-to-video model and the updated rewriting model; wherein the second optimization function is determined according to the detection result, a training sample, a second probability of generating the to-be-detected video based on the description text, and a third probability of generating the rewritten text based on the description text. a formula of the second optimization function is as follows: where x represents the video generated by the second text-to-video model, z represents the description text of the second text-to-video model, D unlabeled represents the reinforcement learning sample, D labeled represents the supervised learning sample; θ represents the model parameters to be optimized; z rew represents the rewritten description text after the rewriting model is rewritten; z ori represents the initial text; p θ (x│z rew ) represents the probability of generating x based on the rewritten description text z rew under the current model; p θ (z rew │z ori ) represents the probability of generating the rewritten description text z ori under the current model based on the description text z rew before rewriting; represents the scoring weighting of the video aesthetics model and the video text relevance model; represents the relevance score of generating video x through the initial text; represents the relevance score output by the video text relevance model; represents the video aesthetics model score; γ and β are both adjustment coefficients; represents the average value.

6. The method of claim 5, wherein, the video detection model comprises a video-text relevance model and a video-aesthetics model, and inputting the to-be-detected video into the video detection model to obtain the detection result corresponding to the to-be-detected video comprises: inputting the description text and the to-be-detected video into the video-text relevance model to obtain relevance of the to-be-detected video; and inputting the to-be-detected video into the video-aesthetics model to obtain aesthetics of the to-be-detected video.

7. The method of claim 5 or 6, wherein, the reinforcement learning sample comprises unlabeled text, and the supervised learning sample comprises labeled data of text and video.

8. The method of any one of claims 5 or 6, wherein, further comprising: annotating the to-be-detected video to obtain a sample for updating the video detection model. 9.A method for generating a video based on a text, comprising: The description text is input into a text-to-video model to obtain a to-be-detected video; wherein the text-to-video model is trained based on the text-to-video model training method in any one of claims 1 to 8.

10. The method of claim 9, further comprising: inputting the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video.

11. The method of claim 10, wherein the video detection model comprises a text-video correlation model and a video aesthetic degree model, and inputting the to-be-detected video into the video detection model to obtain the detection result corresponding to the to-be-detected video comprises: inputting the description text and the to-be-detected video into the text-video correlation model to obtain a correlation of the to-be-detected video; inputting the to-be-detected video into the video aesthetic degree model to obtain an aesthetic degree of the to-be-detected video.

12. The method of claim 10 or 11, further comprising: sorting a plurality of to-be-detected videos according to the detection result.

13. The method of any one of claims 9 to 11, further comprising: rewriting an initial text into the description text using a rewriting model.

14. A training device of a text-to-video model, comprising: a text processing module configured to input a description text into a first text-to-video model to obtain a to-be-detected video; a video processing module configured to input the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video; a training module configured to train the first text-to-video model according to the detection result to obtain a second text-to-video model, and configured to adjust parameters of the first text-to-video model according to a first optimization function to obtain the second text-to-video model, wherein the first optimization function is determined according to the detection result, a training sample, and a first probability of generating the to-be-detected video based on the description text; a formula of the first optimization function is as follows: wherein, x represents a video generated by the first text-to-video model; z represents a description text input into the first text-to-video model; D unlabeled represents a reinforcement learning sample; D labeled represents a supervised learning sample; θ represents model parameters to be optimized; p θ (x|z) represents a first probability of generating x based on z under the current model; represents a correlation score output by the video-text correlation model, represents a beauty score output by the video beauty model; γ and β are both adjustment coefficients; represents an average value.

15. The apparatus of claim 14, wherein, the video detection model comprises a video text correlation model and a video aesthetic degree model, and the video processing module is configured to: input the description text and the to-be-detected video into the video text correlation model to obtain a correlation of the to-be-detected video; and input the to-be-detected video into the video aesthetic degree model to obtain an aesthetic degree of the to-be-detected video.

16. The apparatus of claim 14 or 15, wherein, the reinforcement learning sample comprises unlabeled text, and the supervised learning sample comprises labeled data of text and video.

17. The apparatus of claim 14 or 15, wherein, further comprising: a labeling module configured to label the to-be-detected video to obtain a sample for updating the video detection model.

18. A training device of a text-to-video model, comprising: a rewriting module configured to input an initial text into a rewriting model to obtain a description text; input the description text into a first text-to-video model to obtain a to-be-detected video; input the to-be-detected video into a video detection model to obtain a detection result corresponding to the to-be-detected video; The joint training submodule is configured to perform joint training on the first text-to-video model and the rewriting model according to the detection result, to obtain a second text-to-video model and an updated rewriting model; the joint training submodule is configured to adjust parameters of the first text-to-video model according to a second optimization function, to obtain the second text-to-video model and the updated rewriting model; the second optimization function is determined according to the detection result, a training sample, a second probability of generating the to-be-detected video based on the description text, and a third probability of generating the rewritten text based on the description text. The formula of the second optimization function is as follows: where x represents the video generated by the second text-to-video model, z represents the description text of the second text-to-video model, D unlabeled represents a reinforcement learning sample, D labeled represents a supervised learning sample; θ represents the model parameters to be optimized; z rew represents the description text after being rewritten by the rewriting model; z ori represents the initial text; p θ (x│z rew ) represents the probability of generating x based on the rewritten description text z rew under the current model; p θ (z rew │z ori ) represents the probability of generating the rewritten description text z ori under the current model based on the original description text z rew ; represents the score weighting of the video aesthetics model and the video text relevance model; represents the relevance score of generating video x from the initial text; represents the relevance score output by the video text relevance model; represents the score of the video aesthetics model; γ and β are adjustment coefficients; represents the average value.

19. The apparatus of claim 18, wherein, The video detection model includes a video-text relevance model and a video aesthetic degree model; the to-be-detected video is input into the video detection model, to obtain a detection result corresponding to the to-be-detected video, including: The description text and the to-be-detected video are input into the video-text relevance model, to obtain a relevance of the to-be-detected video; the to-be-detected video is input into the video aesthetic degree model, to obtain an aesthetic degree of the to-be-detected video.

20. The apparatus of claim 18 or 19, wherein, The reinforcement learning sample includes an unlabeled text, and the supervised learning sample includes labeled data of a text and a video.

21. The apparatus of claim 18 or 19, wherein, Further comprising: A labeling module configured to label the to-be-detected video, to obtain a sample for updating the video detection model.

22. An apparatus for generating a video based on a text, comprising: a text processing module configured to input a description text into a text-to-video model, to obtain a to-be-detected video; wherein the text-to-video model is trained based on the text-to-video model training method in any one of claims 1 to 8.

23. The apparatus of claim 22, further comprising: a video processing module configured to input the to-be-detected video into a video detection model, to obtain a detection result corresponding to the to-be-detected video.

24. The apparatus of claim 23, wherein the video detection model includes a text-video relevance model and a video aesthetic degree model, and the video processing module includes: a relevance processing module configured to input the description text and the to-be-detected video into the text-video relevance model, to obtain a relevance of the to-be-detected video; an aesthetic degree processing module configured to input the to-be-detected video into the video aesthetic degree model, to obtain an aesthetic degree of the to-be-detected video.

25. The apparatus of claim 23 or 24, further comprising: a sorting module configured to sort a plurality of to-be-detected videos according to the detection result.

26. The apparatus of any one of claims 22 to 24, further comprising: a rewriting module configured to use a rewriting model to rewrite an initial text into the description text.

27. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.

28. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1-13.

29. A computer program product comprising computer program which, when executed by a processor, implements the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Method and system for generating image model by optimizing text through artificial feedback reinforcement learning

    CN116955972A

  • Method and system for evaluating generation-type image video

    CN117237296A