Video generation method, electronic device, and computer readable storage medium
By aligning the initial video generation model with the preset reward model and generating the target video generation model, the problem of low video generation quality in the prior art is solved, and higher quality and personalized video generation is achieved.
Patent Information
- Application Number
- PCT/CN2024/107881
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-07-26
- Publication Date
- 2025-06-12
AI Technical Summary
The existing video generation model is trained through network data, resulting in poor quality of generated videos and cannot meet the personalized needs of users.
The initial video generation model is aligned with the preset reward model by fine-tuning, the target video generation model is generated, and the video is generated based on the target text.
The quality of the generated video is improved, making the video more in line with the user's aesthetic preferences and expectations, and enhancing the personalized characteristics of the video.
Smart Images

Figure CN2024107881_12062025_PF_FP_ABST
Abstract
Description
Video generation method, electronic device, and computer-readable storage medium Technical Field
[0001] The present disclosure relates to the fields of computer technology and video processing technology, and in particular to a video generation method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the popularity of video content on the Internet, the demand for high-quality and personalized video generation is increasing. As a video generation tool, the video generation model can generate realistic video content based on given input.
[0003] Currently, video generation models usually use data from the Internet for model training. Since most data on the Internet are of varying quality, the videos generated by the trained video generation models are of poor quality and do not meet user expectations.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0005] Summary of the Invention
[0006] The embodiments of the present disclosure provide a video generation method, an electronic device, and a computer-readable storage medium to at least solve the technical problem in related technologies of training a video generation model based on network data, resulting in the trained video generation model generating a video of poor quality that does not meet user expectations.
[0007] According to one aspect of an embodiment of the present disclosure, a video generation method is provided, including: obtaining a target text, wherein the target text is used to describe the content of a video to be generated; performing video generation processing on the target text using a target video generation model to obtain a target video, wherein the target video generation model is a model obtained by fine-tuning an initial video generation model and a preset reward model to align the models.
[0008] According to another aspect of an embodiment of the present disclosure, a video generation method is also provided, which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface at least partially includes a video generation scene, including: in response to a first touch operation applied to the graphical user interface, inputting a target text, wherein the target text is used to describe the video content to be generated; in response to a second touch operation applied to the graphical user interface, performing video generation processing on the target text using a target video generation model to obtain a target video, wherein the target video generation model is a model obtained by aligning an initial video generation model with a preset reward model using a fine-tuning method; and displaying the target video in the graphical user interface.
[0009] According to another aspect of an embodiment of the present disclosure, a video generation method is also provided, including: obtaining a currently input video generation dialogue request, wherein the information carried in the video generation dialogue request includes: a target text, which is used to describe the video content to be generated; in response to the video generation dialogue request, returning a video generation dialogue reply, wherein the information carried in the video generation dialogue reply includes: a target video, which is obtained by performing video generation processing on the target text using a target video generation model, and the target video generation model is a model obtained by aligning an initial video generation model with a preset reward model using a fine-tuning method; and displaying the video generation dialogue reply in a graphical user interface.
[0010] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes any one of the above-mentioned video generation methods when running.
[0011] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is further provided, which includes a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned video generation methods.
[0012] In an embodiment of the present disclosure, a target text for describing the video content to be generated is obtained, and then video generation processing is performed on the target text based on a target video generation model obtained by aligning the initial video generation model with a preset reward model in a fine-tuning manner. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video. This achieves the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in the video generated by the trained video generation model having poor quality and not meeting user expectations.
[0013] It is easy to note that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present disclosure, and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0015] FIG1 is a schematic diagram of an application scenario of a video generation method according to Embodiment 1 of the present disclosure;
[0016] FIG2 is a flow chart of a video generation method according to Embodiment 1 of the present disclosure;
[0017] FIG3 is a schematic diagram of a fine-tuning process according to Example 1 of the present disclosure;
[0018] FIG4 is a schematic diagram of a video generation method according to Embodiment 1 of the present disclosure;
[0019] FIG5 is a flow chart of a video generation method according to Embodiment 2 of the present disclosure;
[0020] FIG6 is a flowchart of a video generation method according to Embodiment 3 of the present disclosure;
[0021] FIG7 is a structural diagram of a video generating device according to Embodiment 4 of the present disclosure;
[0022] FIG8 is a schematic structural diagram of another video generating device according to Embodiment 4 of the present disclosure;
[0023] FIG9 is a schematic structural diagram of another video generating device according to Embodiment 4 of the present disclosure;
[0024] FIG10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:
[0028] Video diffusion models: A deep learning-based generative model used to generate or modify video content. These models generate novel sequences of video frames by simulating the distribution of video data. Video diffusion models are typically trained with large amounts of data to learn how to generate realistic videos.
[0029] Human preference refers to the subjective preferences or choices of human users when reviewing or evaluating content. In the context of AI-generated content, human preference typically refers to user preferences for content quality, style, accuracy, and other aspects.
[0030] Human preference model: A machine learning model designed to capture and mimic human preference judgments. This model learns by analyzing human evaluations of content, enabling it to produce results that better align with user preferences in subsequent generation processes. This model typically requires large amounts of manually annotated data for training.
[0031] Alignment: In the context of AI-generated content, alignment generally refers to the process of adjusting the generative model so that the content produced by the generative model is more in line with specific standards or goals, such as user preferences, the requirements of a specific task, etc.
[0032] Reward model: In machine learning, a reward model is used to evaluate the effectiveness of an action or output, and is typically used in reinforcement learning / reward learning. In the context of video generation in the disclosed embodiments, a reward model can be used to evaluate the quality of generated videos to guide the model to produce higher-quality outputs.
[0033] Reward score: A reward score is a metric used to quantify the quality or conformance of a model's output. In the application of the video generation model in the disclosed embodiments, reward fine-tuning specifically refers to the evaluation score of the generated video based on a reward model (e.g., a human preference model). This score reflects the degree of consistency between the generated content and human preferences, target standards, or desired goals. It is typically used to guide and optimize the model's training process to generate video content that better meets user preferences or is of higher quality.
[0034] U-Net: A deep learning neural network architecture widely used in computer vision, particularly in image segmentation. The U-Net architecture, consisting of an encoder and a decoder, derives its name from its U-shaped structure. The encoder extracts features and reduces the dimensionality of the input image, while the decoder restores the encoded feature maps to their original size and performs pixel-level classification or segmentation.
[0035] Model fine-tuning refers to adjusting the model parameters based on a pre-trained model using a small amount of data or domain-specific data to adapt it to a specific task or dataset. Typically, fine-tuning is performed on a model that has been trained on a large dataset. This model is usually a deep learning model that performs well on general tasks, or a natural language processing model pre-trained on a large text corpus.
[0036] Resampling: A commonly used data processing method used to adjust the size, distribution or time interval of data samples. In statistics and machine learning, resampling is often used to solve problems such as sample imbalance, missing data or inconsistent data acquisition frequency. In the embodiment of the present disclosure, resampling the video means adjusting the sampling rate of the original video to change the playback speed of the video or adapt to different playback devices. Resampling can be to increase the sampling rate to improve video quality, or to reduce the sampling rate to reduce the file size or adapt to specific playback requirements. Resampling usually causes changes in the image quality and smoothness of the video.
[0037] Discriminative Dimensionality Reduction for Imitation Learning (DDIM) sampling: DDIM sampling is a generative method designed to extract and learn important information from data. This method is primarily used in the field of imitation learning to build models that can mimic human behavior. The key idea behind DDIM sampling is to reduce the dimensionality of data by distinguishing between important and unimportant dimensions. DDIM sampling performs discriminative dimensionality reduction on the data to identify the dimensions most useful for model training tasks, and then uses these dimensions for model training and generation. Specifically, DDIM sampling first uses feature selection or feature extraction methods to identify the most useful features for the modeling task. It then learns a discriminative dimensionality reduction model to map the data onto the dimensions of these important features. Finally, these important dimensions are used for model training and generation.
[0038] Time-Decay Reward (TAR): A technique used in reinforcement learning to handle situations where the value of future rewards decreases over time. In reinforcement learning, a discount factor is often used to measure the importance of future rewards. However, in some cases, the value of future rewards decreases over time. For example, in some tasks, earlier rewards may be more important than later rewards. TAR accounts for the influence of time by giving higher weight to earlier rewards. This can be achieved by introducing a time decay function into the reward calculation, such as an exponential decay function or a polynomial decay function.
[0039] Sparse sampling: A data sampling method used to select a subset of samples from a large dataset for analysis or processing. In sparse sampling, only a small subset of the dataset is selected to represent the entire dataset, reducing computational cost and time. Sparse sampling can be achieved through random sampling, stratified sampling, or other sampling methods that ensure that the selected samples are representative of the characteristics of the entire dataset.
[0040] Low Power Wide Area Network (LoRA): A low-power wide area network (LPWAN) technology that enables low-power communication over long distances. LoRA technology uses a spread spectrum modulation technique, which enables long-distance communication at low power. LoRA technology enables more efficient and reliable communication by fine-tuning LoRA device parameters, also known as efficient fine-tuning. These parameters include transmit power, data rate, and receive sensitivity. By properly adjusting these parameters, better performance can be achieved in different application scenarios.
[0041] The related art of training video generation models based on network data has the following defects.
[0042] Defect 1: Since most online data is of varying quality, the quality of videos generated by the trained video generation model is poor, which does not meet user expectations.
[0043] Defect 2: The video diffusion model in related technologies cannot fully consider human aesthetic preferences and content relevance.
[0044] With respect to the above-mentioned defects, no effective solution has been proposed before the present disclosure.
[0045] Example 1
[0046] According to an embodiment of the present disclosure, a video generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0047] The method embodiment provided in the first embodiment of the present disclosure can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a video generation method. As shown in Figure 1, the computer terminal 10 (or mobile device) may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is merely illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0048] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present disclosure, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0049] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video generation method in the embodiment of the present disclosure. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned video generation method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0050] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0051] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0052] Under the above operating environment, the present disclosure provides a video generation method as shown in FIG. 2. FIG. 2 is a flowchart of a video generation method according to Embodiment 1 of the present disclosure. As shown in FIG. 2, the method may include the following steps:
[0053] Step S21, obtaining a target text, where the target text is used to describe the video content to be generated;
[0054] Step S22, performing video generation processing on the target text by using a target video generation model to obtain a target video, where the target video generation model is a model obtained by aligning an initial video generation model and a preset reward model in a fine-tuning manner.
[0055] The target text can be understood as the text input to the video generation model. For example, the text input to the target video generation model in the embodiments of the present disclosure, that is, the input to the target video generation model. The target text is used to describe the video content to be generated, for example, the video content expected to be generated by the user.
[0056] Exemplarily, if the video content expected to be generated by the user is a video of dogwood blossoms blowing in the wind, the target text can be "山茱萸花在风中飘荡" or "Dogwood blossoms are blowing in the wind", etc. It can be understood that the target text can be described in natural language words, such as Chinese, English, Japanese, etc., which is not limited herein.
[0057] The initial video generation model can be understood as an initial model for generating videos based on text. For example, the initial video generation model can be a pre-trained U-net, that is, a U-net model that has been pre-trained on a large-scale dataset. Considering that pre-training on a large-scale dataset can enable the model to learn rich image features and semantic information, thereby improving the model's generalization ability and accuracy on specific tasks, the present disclosure uses a pre-trained U-net model as the initial video generation model to make the video generated by the target video generation model more accurate.
[0058] The preset reward model can be a model used to evaluate the quality of generated videos, thereby guiding the video generation model to produce higher quality outputs. For example, the preset reward model can evaluate and score the generated videos, thereby quantifying the output quality or conformity of the video generation model through reward scores.
[0059] It should be noted that, considering that the final target video generation model should fully consider human aesthetic preferences and content relevance in order to generate videos that meet user expectations, the preset reward model in the embodiment of the present disclosure adopts an image-based human preference model, namely an image reward model. This image reward model can evaluate the quality of the generated video based on human aesthetic preferences and content relevance, so that the video generated by the final target video generation model is more in line with user expectations.
[0060] The target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model in accordance with the embodiment of the present disclosure, that is, a model obtained by fine-tuning the pre-trained U-shaped network and the image reward model, thereby accurately generating video content that meets user preferences and expectations according to the target video generation model.
[0061] In an embodiment of the present disclosure, a target text for describing the video content to be generated is obtained, and then video generation processing is performed on the target text based on a target video generation model obtained by aligning the initial video generation model with a preset reward model in a fine-tuning manner. That is, a pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video. This can make the generated target video more consistent with human aesthetic preferences and the target text content, improve the video quality of the generated target video, that is, obtain a video that better meets user expectations, thereby improving the personalization level of the generated video content, and opening up new application prospects in the field of video generation.
[0062] The above-mentioned video generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving video generation in the fields of e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services, for example: scenarios of generating product display content in e-commerce services, scenarios of generating learning content videos in educational services, scenarios of generating case-related videos in legal services, etc., which are not limited here.
[0063] According to the embodiment of the present disclosure, a target text for describing the content of the video to be generated is obtained, and then a video generation process is performed on the target text based on a target video generation model obtained by aligning the initial video generation model with the preset reward model by fine-tuning. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video. This achieves the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and the target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in poor quality of the video generated by the trained video generation model, which does not meet user expectations.
[0064] In an optional embodiment, the target video generation model is a video diffusion model, and the preset reward model is an image reward model, wherein the image reward model is used to perform preference learning on the video diffusion model.
[0065] In the embodiment of the present disclosure, the target video generation model may be a video diffusion model, an autoregressive model, etc., and the preset reward model may be an image reward model for performing preference learning on the video diffusion model.
[0066] In an optional embodiment, the video generation method further includes the following method steps:
[0067] Step S23: using the training sample to sample the initial video generation model to generate a sample video, wherein the training sample includes: a plurality of video-text pairs, each of the plurality of video-text pairs includes: a training video and a training text, and the training text is used to describe the video content of the training video;
[0068] Step S24, using a preset reward model to calculate rewards for the sampled video to obtain a target reward result;
[0069] Step S25: Adjust the model parameters of the initial video generation model based on the target reward result to generate a target video generation model.
[0070] The training samples can be understood as training samples to be used for model fine-tuning. For example, they can be part of the video text selected from the pre-training data set. The training samples can be selected according to any rules, which is not limited in the embodiments of the present disclosure.
[0071] The training samples include multiple video-text pairs, each of which includes a training video and corresponding training text, wherein the training text is text content used to describe the video content of the training video.
[0072] The sampled video can be understood as a video obtained by sampling the training samples using the initial video generation model, that is, a video generated by sampling the training video and training text using the pre-trained U-shaped network model.
[0073] The target reward result is the reward score of the sampled video obtained by calculating the reward for the sampled video according to the preset reward model. The target reward result is used to reflect whether the sampled video is consistent with human aesthetic preferences and text content, which can be denoted as R.
[0074] In the embodiment of the present disclosure, before the fine-tuning phase begins, some video texts can be selected from the pre-training data set as training samples, and then the initial video generation model is sampled using the training samples, that is, the training videos and training texts in the training samples are sampled using the pre-trained U-shaped network model to generate a sampled video. The preset reward model is then used to calculate the reward for the sampled video, that is, the image reward model is used to calculate the reward for the generated sampled video to obtain the target reward result corresponding to the sampled video. Finally, the model parameters of the initial video generation model are adjusted based on the target reward result, that is, the initial video generation model is fine-tuned based on the target reward result to obtain the target video generation model.
[0075] In an optional embodiment, in step S23, the initial video generation model is sampled using the training sample to generate a sampled video, including the following method steps:
[0076] Step S231, performing noise processing on the training video to obtain a noisy video;
[0077] Step S232: resample the initial video generation model using the noisy video and the training text to generate a sampled video.
[0078] In an embodiment of the present disclosure, when a training sample is used to sample the initial video generation model, the training video in the training sample can be denoised to obtain a noisy video, and then the obtained noisy video and the training text in the training sample are used to resample the initial video generation model to generate a sampled video.
[0079] For example, a diffusion process with noise may be performed on the training video to obtain a noisy video, and then the DDIM sampling is performed on the initial video generation model using the obtained noisy video and the training text in the training sample to generate a sampled video.
[0080] It can be seen that compared with the traditional generation process method, the resampling of videos through DDIM sampling in the present disclosure can more effectively utilize data information, improve the generation and generalization capabilities of the model, and help the model better understand the structure and characteristics of the data, thereby better imitating human behavior.
[0081] In an optional embodiment, in step S231, performing noise processing on the training video to obtain a noisy video includes the following method steps:
[0082] Step S2311: Obtaining the number of noise addition steps and the noise level corresponding to the training video, wherein the number of noise addition steps is used to determine the number of steps to be noised in the training video using a preset noise addition function, and the noise level is used to determine the degree of damage to the training video;
[0083] Step S2312: Noise the training video based on the number of noise addition steps and the noise level to obtain a noisy video.
[0084] The preset noise adding function can be expressed as d(τ, D), which is used to calculate how many steps the training video should be noised to based on the noise level τ and the number of noise adding steps D. The output result is usually between 1 and 1000.
[0085] The number of denoising steps D is used to determine the number of steps to be denoised for the training video by using a preset denoising function d(), and is usually set to 20.
[0086] The noise level τ is used to determine the degree of damage to the training video, and its value range is between 0 and 1.
[0087] For example, if the video generation model used in the present embodiment is a diffusion model, the number of noise addition steps D represents the number of steps used in the diffusion model generation process. The preset noise addition function d() can determine the degree of video noise addition based on the noise level τ and the number of noise addition steps D, so as to generate a video with a certain degree of noise in the diffusion model.
[0088] In the embodiment of the present disclosure, when performing noise processing on a training video, the number of noise adding steps d and the noise level τ corresponding to the training video can be obtained, and then the training video can be noise-processed based on the number of noise adding steps d and the noise level τ, thereby generating a noisy video with a certain degree of noise.
[0089] It can be seen that compared with the related art of starting noise addition processing from text, the embodiment of the present disclosure adopts DDIM sampling. The method of starting noise addition processing from video only requires the τ ratio of the complete generation process calculation amount, so the calculation amount is small, which can effectively save computing resources.
[0090] In an optional embodiment, in step S24, a preset reward model is used to calculate a reward for the sampled video to obtain a target reward result, including the following method steps:
[0091] Step S241: Calculate rewards for the sampled video using a preset reward model to obtain an initial reward result;
[0092] Step S242 , adjusting the initial reward weight corresponding to the initial reward result using a time-decay reward method to generate a target reward result, wherein the initial reward weight is a default reward weight corresponding to the video frame sequence included in the sampled video.
[0093] In an embodiment of the present disclosure, after a sampled video is generated, when a preset reward model is used to calculate a reward for the sampled video, the preset reward model can be used to calculate a reward for the sampled video, that is, an image-based human preference model is used, that is, an image reward model is used to calculate a reward, thereby obtaining an initial reward result, that is, the original output result of the preset reward model.
[0094] For example, the generated sample video and the text input when generating the sample video can be input into the image reward model, thereby obtaining a score value based on the output of the image reward model, that is, obtaining the initial reward result. The score value output by the image reward model can be between 0 and 1, with a higher score indicating better video quality, and this is not limited here.
[0095] In addition, in order to improve the training effect of the model, after obtaining the initial reward result, the time-decay reward (TAR) method can be used to adjust the initial reward weight corresponding to the initial reward result, that is, to adjust the default reward weight corresponding to the video frame sequence contained in the sampled video, so as to generate the target reward result.
[0096] In an optional embodiment, in step S241, a preset reward model is used to calculate a reward for the sampled video to obtain an initial reward result, including the following method steps:
[0097] Step S2411, performing video segment sampling on the video frame sequence to obtain segment sampling results;
[0098] Step S2412: Use a preset reward model to calculate rewards for the segmented sampling results to obtain an initial reward result.
[0099] In the embodiment of the present disclosure, when a preset reward model is used to calculate rewards for sampled videos, in order to improve the efficiency of the learning process, segmented video rewards can be used to perform segmental sampling on the video frame sequence, that is, sparse sampling is performed on the video, the video is split into several segments, and multiple continuous video frames are grouped to obtain segmented sampling results, and then the preset reward model is used to calculate rewards for the segmented sampling results to obtain an initial reward result.
[0100] For example, taking a sampled video with 16 frames as an example, the 16-frame video can be divided into 4 groups with 4 frames in each group. The preset reward model is used to calculate the rewards for the 4 groups to obtain the initial reward results, so as to improve the training effect of the model.
[0101] In an optional embodiment, in step S2411, performing segmented video sampling on the video frame sequence to obtain segmented sampling results includes the following method steps:
[0102] Step S24111, obtaining a feature space representation of a video frame sequence;
[0103] Step S24112, performing video segment sampling on the feature space representation to obtain a color space representation of the segmented video frame;
[0104] Step S24113, determining the segmented sampling result based on the color space representation.
[0105] In the embodiment of the present disclosure, when performing video segmentation sampling on a video frame sequence, the feature space representation z_0 of the video frame sequence can be obtained, and then the feature space representation z_0 can be subjected to video segmentation sampling, thereby obtaining the color (RGB) space representation x_0^g of the segmented video frame, and finally determining the segmentation sampling result based on the color space representation x_0^g.
[0106] In an optional embodiment, in step S242, the initial reward weight corresponding to the initial reward result is adjusted using a time-decay reward method to generate a target reward result, including the following method steps:
[0107] Step S2421: Differentiately adjust the initial reward weight corresponding to the initial reward result using a time-decay reward method to obtain a target reward weight, where the target reward weight is used to indicate that the current reward weight of the first video frame in the video frame sequence is higher than the current reward weight of the second video frame, the first video frame is located in the middle of the video frame sequence, and the second video frame is located at the edge of the video frame sequence;
[0108] Step S2422: Generate a target reward result based on the initial reward result and the target reward weight.
[0109] In the embodiment of the present disclosure, after obtaining the initial reward result, when adjusting the initial reward weight corresponding to the initial reward result using the time decay reward method, the time decay reward method can be used to perform differential adjustment on the initial reward weight corresponding to the initial reward result. By adjusting the weight of the middle frame of the video to a higher value and the weight of the edge frame to a lower value, effective and efficient fine-tuning is achieved, thereby obtaining the target reward weight.
[0110] That is, the present disclosure can use a time decay reward method to adjust the current reward weight of the first video frame located in the middle of the video frame sequence to be higher than the current reward weight of the second video frame located at the edge of the video frame sequence. For example, taking a video with 16 frames as an example, the reward weight of each video frame is adjusted by time decay reward, and the coefficient of the middle video frame among the 16 frames is adjusted to 1, and the coefficient of the side video frame is adjusted to Thus generating the target reward result.
[0111] In an optional embodiment, a graphical user interface is provided by a terminal device, and content displayed by the graphical user interface at least partially includes a video generation scene. The video generation method further includes:
[0112] Step S26, in response to the first touch operation on the graphical user interface, inputting a target text, wherein the target text is used to describe the video content to be generated;
[0113] Step S27, in response to the second touch operation on the graphical user interface, performing video generation processing on the target text using the target video generation model to obtain a target video, wherein the target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model;
[0114] Step S28: Display the target video in the graphical user interface.
[0115] In the embodiments of the present disclosure, the graphical user interface displays at least a video generation scene. A user can perform control operations in the image encoding scene to input target text, which is then processed using a target video generation model to generate a target video. It is understood that the aforementioned video generation scene can include, but is not limited to, application scenarios involving video generation in fields such as e-commerce, education, healthcare, conferencing, social networking, financial products, logistics, and navigation.
[0116] The graphical user interface further includes a first control (or a first touch area). When a first touch operation is detected on the first control (or the first touch area), a target text input by the user can be obtained. The target text can be input by the user into a text box in the graphical user interface through the first touch operation. The first touch operation can be a click, box selection, check, conditional filtering, and other operations, which are not limited here.
[0117] The graphical user interface also includes a second control (or a second touch area). When a second touch operation is detected on the second control (or the second touch area), a target video generation model can be used to generate a target video based on the target text input by the user. The second touch operation can be a click, box selection, checkbox, conditional filtering, or other operation, which is not limited here.
[0118] After obtaining the target video, the target video can be displayed in the graphical user interface to provide feedback to the user.
[0119] It should be noted that both the first touch operation and the second touch operation can be operations in which a user touches the display screen of the terminal device with a finger and touches the terminal device. The touch operation can include single-point touch and multi-point touch, wherein the touch operation of each touch point can include clicking, long pressing, pressing hard, swiping, etc. The first touch operation and the second touch operation can also be touch operations implemented through input devices such as a mouse and keyboard, which are not limited here.
[0120] Figure 3 is a schematic diagram of the fine-tuning process according to Example 1 of the present disclosure, wherein z is the representation of the sampled video in the feature space. c is the text corresponding to the sampled video z. It can be understood that z and c together constitute a corresponding video-text pair. In Figure 3, the text data c is "Cornus flowers are flying in the air" and the video data z is the video corresponding to the text data c as an example. z_d(τ·D) represents the noisy video, wherein d() is a function for calculating how many steps the video should be noisy, D is the number of noisy steps, and τ is the noise level. z_0 is the representation of the generated video in the feature space, x_0^g is the representation of the generated video in the RGB space after sampling, and R is the calculated reward score.
[0121] As shown in Figure 3, before the fine-tuning phase begins, some video-text pairs are first selected from the pre-training data set for fine-tuning. After selecting some video-text pairs, a diffusion process with noise is implemented on the video data z in the video-text pairs, wherein the noise level is set to τ and the number of noise addition steps is D, that is, the video data z is noised by the diffusion model to obtain a noisy video z_d(τ·D). Then, based on this noise addition process, the noisy video z_d(τ·D) and the corresponding text data c are resampled using DDIM sampling to generate a sampled video z_0. After generating the sampled video z_0, the present disclosure adopts an image-based human preference model as a reward model, that is, an image reward model is used for reward calculation, and when calculating the reward score, in order to improve the efficiency of the learning process, a segmented sampling and decoding method is adopted to segment the sampled video z_0 to obtain the color space representation x_0^g of the segmented video frame and the segmented sampling result. Then, the reward of the segmented sampling results is calculated using the image reward model and the corresponding text data c, and the reward weight of each video frame is adjusted through the time-decaying reward (TAR) to achieve effective and efficient fine-tuning, thereby obtaining the reward score R, which is the target reward result.
[0122] It is understandable that after obtaining the reward score R, the noisy video z_d(1) can be obtained through gradient backpropagation based on the reward score R to adjust the network parameters to minimize the error, thereby optimizing the model. I will not go into details here.
[0123] Figure 4 is a schematic diagram of a video generation method according to Example 1 of the present disclosure. As shown in Figure 4, compared to the method of the related art, that is, generating a video by inputting text into a pre-trained U-shaped network with fixed model parameters, the generated video does not meet the user's expectations. The method of the present disclosure, through fine-tuning, aligns the video generation model adopted by the present disclosure with the human preference model of the image based on the text and the noisy video, that is, aligns the pre-trained U-shaped network model with trainable model parameters with the image reward model, and in the fine-tuning process, adopts the efficient fine-tuning technology LoRA to improve the performance of the model by making slight adjustments to the parameters of the model. In addition, it can also perform gradient processing optimization model based on the image reward model to obtain the target video generation model. Finally, the target video generation model obtained by the present disclosure is used to perform video generation processing on the input text, so that the target video that meets the user's expectations can be obtained.
[0124] It can be seen that the method of the embodiment of the present disclosure can use the image reward model to learn human preferences for the video diffusion model without a video reward model (such a model has not yet been trained in the relevant technology), so that after the fine-tuning method designed by the present disclosure is used, the output results of the video diffusion model are more in line with user expectations and more popular with humans.
[0125] It is easy to understand that the beneficial effects of the video generation method provided by the present disclosure include the following points.
[0126] Beneficial effect (1): In order to align the video diffusion model with human preferences, the present disclosure proposes to align the video diffusion model with the human preference model of images, so that the video diffusion model obtained after alignment can fully consider human aesthetic preferences and content relevance, output video content that better meets user expectations, improve the personalization level of content, and open up new application prospects in the field of video generation;
[0127] Beneficial effect (2): This disclosure proposes segmented video rewards and time decay rewards in video reward fine-tuning learning to enable the model to be effectively fine-tuned.
[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0129] In addition, it should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0131] Example 2
[0132] Under the operating environment as in Embodiment 1, the present disclosure provides a video generation method as shown in FIG. 5. FIG. 5 is a flowchart of a video generation method according to Embodiment 2 of the present disclosure. As shown in FIG. 5, the method includes:
[0133] Step S51, obtaining a video generation call request through a first application programming interface. Among them, the request data carried in the video generation call request includes: a target text, which is used to describe the video content to be generated. The video generation call request is used to request to call a target video generation model to perform video generation processing on the target text to obtain a target video. The target video generation model is a model obtained by aligning an initial video generation model and a preset reward model in a fine-tuning manner;
[0134] Step S52, returning a video generation call response through a second application programming interface. Among them, the response data carried in the video generation call response includes: the target video.
[0135] Both the first application programming interface and the second application programming interface can be understood as an application programming interface (Application Programming Interface, API). By calling the first application programming interface, a video generation call request carrying the target text can be obtained. By calling the second application programming interface, a video generation call response can be returned.
[0136] The first application programming interface and the second application programming interface in the embodiments of the present disclosure can be the same application programming interface or different application programming interfaces, which is not limited herein.
[0137] The target text can be understood as the text input to the video generation model. For example, the text input to the target video generation model in the embodiments of the present disclosure, that is, the input to the target video generation model. The target text is used to describe the video content to be generated. For example, it is used to describe the video content expected by the user.
[0138] Exemplarily, if the video content expected by the user is a video of dogwood blossoms blowing in the wind (Dogwood blossoms are blowing in the wind), the target text can be "山茱萸花在风中飘荡" or "Dogwood blossoms are blowing in the wind", etc. It can be understood that the target text can be described in natural language words, such as Chinese, English, Japanese, etc., which is not limited herein.
[0139] The initial video generation model can be understood as an initial model for generating videos based on text. For example, the initial video generation model can be a pre-trained U-net, that is, a U-net model that has been pre-trained on a large-scale dataset. Considering that pre-training on a large-scale dataset can enable the model to learn rich image features and semantic information, thereby improving the model's generalization ability and accuracy on specific tasks, the present disclosure uses a pre-trained U-net model as the initial video generation model to make the video generated by the target video generation model more accurate.
[0140] The preset reward model can be a model used to evaluate the quality of generated videos, thereby guiding the video generation model to produce higher quality outputs. For example, the preset reward model can evaluate and score the generated videos, thereby quantifying the output quality or conformity of the video generation model through reward scores.
[0141] It should be noted that, considering that the final target video generation model should fully consider human aesthetic preferences and content relevance in order to generate videos that meet user expectations, the preset reward model in the embodiment of the present disclosure adopts an image-based human preference model, namely an image reward model. This image reward model can evaluate the quality of the generated video based on human aesthetic preferences and content relevance, so that the video generated by the final target video generation model is more in line with user expectations.
[0142] The target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model in accordance with the embodiment of the present disclosure, that is, a model obtained by fine-tuning the pre-trained U-shaped network and the image reward model, thereby accurately generating video content that meets user preferences and expectations according to the target video generation model.
[0143] In the disclosed embodiment, a video generation call request can be obtained by calling a first application programming interface, wherein the video generation call request carries a target text for describing the content of the video to be generated. Based on the obtained video generation call request, a target video generation model is used to perform video generation processing on the target text, thereby obtaining a target video. Then, a video generation call response including the target video can be returned by calling a second application programming interface. According to this method, the generated target video can be made more consistent with human aesthetic preferences and the content of the target text, thereby improving the video quality of the generated target video, that is, obtaining a video that better meets user expectations, thereby improving the personalization level of the generated video content, and opening up new application prospects in the field of video generation.
[0144] The above-mentioned video generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving video generation in the fields of e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services, for example: scenarios of generating product display content in e-commerce services, scenarios of generating learning content videos in educational services, scenarios of generating case-related videos in legal services, etc., which are not limited here.
[0145] By adopting the embodiment of the present disclosure, a video generation call request can be obtained by calling the first application programming interface, and the video generation call request carries a target text for describing the video content to be generated. The target text is subjected to video generation processing using the target video generation model according to the obtained video generation call request, so that the target video can be obtained. Then, a video generation call response including the target video can be returned by calling the second application programming interface, thereby achieving the purpose of generating a target video that meets the user's expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans, thereby solving the technical problem in the related technology of training video generation models based on network data, resulting in the video generated by the trained video generation model having poor quality and not meeting user expectations.
[0146] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0147] Example 3
[0148] In the operating environment of Example 1, the present disclosure provides a video generation method as shown in FIG6 . FIG6 is a flow chart of a video generation method according to Example 3 of the present disclosure. As shown in FIG6 , the method includes:
[0149] Step S61: obtaining a currently input video generation dialogue request, wherein the information carried in the video generation dialogue request includes: a target text, which is used to describe the video content to be generated;
[0150] Step S62: In response to the video generation dialog request, a video generation dialog reply is returned, wherein the information carried in the video generation dialog reply includes: a target video, which is obtained by performing video generation processing on the target text using a target video generation model, where the target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model;
[0151] Step S63: Display the video-generated dialogue response in the graphical user interface.
[0152] A video generation dialogue request can be understood as a dialogue request (request) initiated by a user to a computer or a robot. The video generation dialogue request carries a target text for describing the video content to be generated.
[0153] The target text can be understood as the text input into the video generation model. For example, the text input into the target video generation model in the embodiments of the present disclosure, that is, it serves as the input of the target video generation model. The target text is used to describe the video content to be generated, for example, to describe the video content expected to be generated by the user.
[0154] Exemplarily, if the video content expected to be generated by the user is a video of dogwood blossoms blowing in the wind (Dogwood blossoms are blowing in the wind), the target text can be "山茱萸花在风中飘荡" or "Dogwood blossoms are blowing in the wind", etc. It can be understood that the target text can be described in natural language words, such as Chinese, English, Japanese, etc., which is not limited here.
[0155] In the embodiments of the present disclosure, in response to the obtained video generation dialogue request, a video generation dialogue reply (response) can be returned. Among them, the video generation dialogue reply carries a target video obtained by performing video generation processing on the target text using the target video generation model, and the target video generation model is a model obtained by aligning the initial video generation model and the preset reward model in a fine-tuning manner.
[0156] The initial video generation model can be understood as an initial model for generating videos based on text. Exemplarily, the initial video generation model can be a pre-trained U-shaped network (Pre-trained UNet), that is, a U-shaped network model pre-trained on a large-scale dataset. Considering that pre-training on a large-scale dataset can enable the model to learn rich image features and semantic information, thereby improving the generalization ability and accuracy of the model in specific tasks. Therefore, adopting the pre-trained U-shaped network model as the initial video generation model in the present disclosure can make the videos generated by the finally obtained target video generation model more accurate.
[0157] The preset reward model can be a model used to evaluate the quality of the generated video, so as to guide the video generation model to produce higher-quality outputs. For example, the preset reward model can evaluate the generated video to obtain a score, thereby quantifying the output quality or compliance of the video generation model through the reward score.
[0158] It should be noted that, considering that the final target video generation model should fully consider human aesthetic preferences and content relevance in order to generate videos that meet user expectations, the preset reward model in the embodiment of the present disclosure adopts an image-based human preference model, namely an image reward model. This image reward model can evaluate the quality of the generated video based on human aesthetic preferences and content relevance, so that the video generated by the final target video generation model is more in line with user expectations.
[0159] The target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model in accordance with the embodiment of the present disclosure, that is, a model obtained by fine-tuning the pre-trained U-shaped network and the image reward model, thereby accurately generating video content that meets user preferences and expectations according to the target video generation model.
[0160] After returning to the video to generate the conversation response, the video to generate the conversation response can be displayed in the graphical user interface.
[0161] The above-mentioned video generation method provided by the embodiments of the present disclosure can be applied to, but is not limited to, application scenarios involving video generation in the fields of e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services and navigation services, for example: scenarios of generating product display content in e-commerce services, scenarios of generating learning content videos in educational services, scenarios of generating case-related videos in legal services, etc., which are not limited here.
[0162] According to the embodiment of the present disclosure, a video generation dialogue request carrying a target text for describing the content of the video to be generated is obtained. Then, in response to the video generation dialogue request, video generation processing is performed on the target text based on a target video generation model obtained by fine-tuning the initial video generation model and the preset reward model. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video, that is, a video generation dialogue response carrying the target video. Finally, the video generation dialogue response is displayed in a graphical user interface, thereby achieving the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and the content of the target text, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in poor quality of the video generated by the trained video generation model that does not meet user expectations.
[0163] It should be noted that the preferred implementation of this embodiment can be found in the relevant description in Example 1 and will not be repeated here.
[0164] Example 4
[0165] According to an embodiment of the present disclosure, an embodiment of a device for implementing the above-mentioned video generation method is also provided. FIG7 is a structural diagram of a video generation device according to embodiment 4 of the present disclosure. As shown in FIG7 , the device includes:
[0166] An acquisition module 701 is configured to acquire a target text, wherein the target text is used to describe the video content to be generated;
[0167] The processing module 702 is configured to perform video generation processing on the target text using a target video generation model to obtain a target video, wherein the target video generation model is a model obtained by aligning the initial video generation model with the preset reward model using a fine-tuning method.
[0168] Optionally, it also includes: a training module, which is configured to use training samples to sample the initial video generation model to generate a sampled video, wherein the training samples include: multiple video-text pairs, and the multiple video-text pairs all include: training videos and training texts, and the training texts are used to describe the video content of the training videos; a preset reward model is used to calculate rewards for the sampled videos to obtain target reward results; based on the target reward results, the model parameters of the initial video generation model are adjusted to generate a target video generation model.
[0169] Optionally, the training module is further configured to: perform noise processing on the training video to obtain a noisy video; and use the noisy video and the training text to perform video resampling on the initial video generation model to generate a sampled video.
[0170] Optionally, the above-mentioned training module is also configured to: obtain the number of noise addition steps and the noise level corresponding to the training video, wherein the number of noise addition steps is used to determine the number of steps to be noised in the training video through a preset noise addition function, and the noise level is used to determine the degree of damage to the training video; and perform noise addition processing on the training video based on the number of noise addition steps and the noise level to obtain a noisy video.
[0171] Optionally, the above-mentioned training module is also configured to: use a preset reward model to calculate the reward for the sampled video to obtain an initial reward result; use a time-decay reward method to adjust the initial reward weight corresponding to the initial reward result to generate a target reward result, wherein the initial reward weight is the default reward weight corresponding to the video frame sequence contained in the sampled video.
[0172] Optionally, the above training module is further configured to: perform video segment sampling on the video frame sequence to obtain segment sampling results; and use a preset reward model to perform reward calculation on the segment sampling results to obtain an initial reward result.
[0173] Optionally, the above training module is further configured to: obtain a feature space representation of a video frame sequence; perform video segmentation sampling on the feature space representation to obtain a color space representation of the segmented video frame; and determine the segmentation sampling result based on the color space representation.
[0174] Optionally, the above-mentioned training module is further configured to: use a time-decay reward method to differentially adjust the initial reward weight corresponding to the initial reward result to obtain a target reward weight, wherein the target reward weight is used to indicate that the current reward weight of the first video frame in the video frame sequence is higher than the current reward weight of the second video frame, the first video frame is located in the middle position of the video frame sequence, and the second video frame is located at the edge position of the video frame sequence; generate the target reward result based on the initial reward result and the target reward weight.
[0175] Optionally, the target video generation model is a video diffusion model, and the preset reward model is an image reward model, wherein the image reward model is used to perform preference learning on the video diffusion model.
[0176] According to the embodiment of the present disclosure, a target text for describing the content of the video to be generated is obtained, and then a video generation process is performed on the target text based on a target video generation model obtained by aligning the initial video generation model with the preset reward model by fine-tuning. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video. This achieves the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and the target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in poor quality of the video generated by the trained video generation model, which does not meet user expectations.
[0177] It should be noted that the acquisition module 701 and the processing module 702 correspond to steps S21 and S22 in Example 1. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of the device and can be run in the computer terminal 10 provided in Example 1.
[0178] According to an embodiment of the present disclosure, another embodiment of a device for implementing the above-mentioned video generation method is also provided. FIG8 is a structural schematic diagram of another video generation device according to embodiment 4 of the present disclosure. As shown in FIG8 , the device includes:
[0179] An acquisition module 801 is configured to acquire a video generation call request through a first application programming interface, wherein the request data carried in the video generation call request includes: a target text, the target text being used to describe the video content to be generated, the video generation call request being used to request calling a target video generation model to perform video generation processing on the target text to obtain a target video, wherein the target video generation model is a model obtained by fine-tuning an initial video generation model and a preset reward model to align the model;
[0180] The return module 802 is configured to return a video generation call response through a second application programming interface, wherein the response data carried in the video generation call response includes: a target video.
[0181] By adopting the embodiment of the present disclosure, a video generation call request can be obtained by calling the first application programming interface, and the video generation call request carries a target text for describing the video content to be generated. The target text is subjected to video generation processing using the target video generation model according to the obtained video generation call request, so that the target video can be obtained. Then, a video generation call response including the target video can be returned by calling the second application programming interface, thereby achieving the purpose of generating a target video that meets the user's expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans, thereby solving the technical problem in the related technology of training video generation models based on network data, resulting in the video generated by the trained video generation model having poor quality and not meeting user expectations.
[0182] It should be noted that the acquisition module 801 and the return module 802 correspond to steps S51 and S52 in Example 2. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of the device and can be run in the computer terminal 10 provided in Example 1.
[0183] According to an embodiment of the present disclosure, another embodiment of a device for implementing the above-mentioned video generation method is also provided. FIG9 is a structural diagram of another video generation device according to embodiment 4 of the present disclosure. As shown in FIG9 , the device includes:
[0184] The acquisition module 901 is configured to acquire a currently input video generation dialogue request, wherein the information carried in the video generation dialogue request includes: a target text, the target text being used to describe the video content to be generated;
[0185] The first response module 902 is configured to respond to the video generation dialog request and return a video generation dialog reply. The video generation dialog reply includes information such as a target video, which is obtained by performing video generation processing on the target text using a target video generation model. The target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model.
[0186] The display module 903 is configured to display the video-generated dialogue response in the graphical user interface.
[0187] According to the embodiment of the present disclosure, a video generation dialogue request carrying a target text for describing the content of the video to be generated is obtained. Then, in response to the video generation dialogue request, video generation processing is performed on the target text based on a target video generation model obtained by fine-tuning the initial video generation model and the preset reward model. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video, that is, a video generation dialogue response carrying the target video. Finally, the video generation dialogue response is displayed in a graphical user interface, thereby achieving the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and the content of the target text, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in poor quality of the video generated by the trained video generation model that does not meet user expectations.
[0188] It should be noted that the acquisition module 901, the first response module 902, and the display module 903 correspond to steps S61 to S63 in Example 3. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules can also be run as part of the device in the computer terminal 10 provided in Example 1.
[0189] It should be noted that the preferred implementation scheme involved in the above embodiments of the present disclosure is the same as the solution provided in Example 1, as well as the application scenario and implementation process, but is not limited to the solution provided in Example 1.
[0190] Example 5
[0191] The embodiment of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.
[0192] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0193] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the video generation method: obtaining a target text, wherein the target text is used to describe the video content to be generated; using a target video generation model to perform video generation processing on the target text to obtain a target video, wherein the target video generation model is a model obtained by aligning the initial video generation model with the preset reward model using a fine-tuning method.
[0194] Optionally, Figure 10 is a structural block diagram of a computer terminal according to Embodiment 5 of the present disclosure. As shown in Figure 10, the computer terminal A may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0195] Among them, the memory can be configured to store software programs and modules, such as the program instructions / modules corresponding to the video generation method and device in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the stored software programs and modules, that is, realizing the above-mentioned video generation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0196] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtaining the target text, wherein the target text is used to describe the video content to be generated; using the target video generation model to perform video generation processing on the target text to obtain the target video, wherein the target video generation model is a model obtained by fine-tuning the initial video generation model and the preset reward model.
[0197] Optionally, the processor may also execute the program code of the following steps: sampling the initial video generation model using training samples to generate a sampled video, wherein the training samples include: multiple video-text pairs, and the multiple video-text pairs each include: a training video and a training text, and the training text is used to describe the video content of the training video; performing reward calculation on the sampled video using a preset reward model to obtain a target reward result; adjusting the model parameters of the initial video generation model based on the target reward result to generate a target video generation model.
[0198] Optionally, the processor may further execute program codes of the following steps: performing noise processing on the training video to obtain a noisy video; and performing video resampling on the initial video generation model using the noisy video and the training text to generate a sampled video.
[0199] Optionally, the processor may also execute the program code of the following steps: obtaining the number of noise addition steps and the noise level corresponding to the training video, wherein the number of noise addition steps is used to determine the number of steps to be noised in the training video through a preset noise addition function, and the noise level is used to determine the degree of damage to the training video; performing noise addition processing on the training video based on the number of noise addition steps and the noise level to obtain a noisy video.
[0200] Optionally, the processor may also execute the program code of the following steps: performing reward calculation on the sampled video using a preset reward model to obtain an initial reward result; adjusting the initial reward weight corresponding to the initial reward result using a time-decay reward method to generate a target reward result, wherein the initial reward weight is a default reward weight corresponding to the video frame sequence contained in the sampled video.
[0201] Optionally, the processor may further execute the program code of the following steps: performing video segment sampling on the video frame sequence to obtain segment sampling results; and performing reward calculation on the segment sampling results using a preset reward model to obtain an initial reward result.
[0202] Optionally, the processor may also execute the program code of the following steps: obtaining a feature space representation of a video frame sequence; performing video segmentation sampling on the feature space representation to obtain a color space representation of the segmented video frame; and determining the segmentation sampling result based on the color space representation.
[0203] Optionally, the processor may further execute the program code of the following steps: differentially adjusting the initial reward weight corresponding to the initial reward result using a time-decay reward method to obtain a target reward weight, wherein the target reward weight is used to indicate that the current reward weight of the first video frame in the video frame sequence is higher than the current reward weight of the second video frame, the first video frame is located in the middle position of the video frame sequence, and the second video frame is located at the edge position of the video frame sequence; generating the target reward result based on the initial reward result and the target reward weight.
[0204] Optionally, the target video generation model is a video diffusion model, and the preset reward model is an image reward model, wherein the image reward model is used to perform preference learning on the video diffusion model.
[0205] According to the embodiment of the present disclosure, a target text for describing the content of the video to be generated is obtained, and then a video generation process is performed on the target text based on a target video generation model obtained by aligning the initial video generation model with the preset reward model by fine-tuning. That is, the pre-trained U-shaped network is aligned with the image reward model by fine-tuning to obtain a fine-tuned target video generation model, and the fine-tuned target video generation model is used to generate a video based on the target text, thereby obtaining a target video. This achieves the purpose of generating a target video that meets user expectations, thereby achieving the technical effect of making the generated target video more consistent with human aesthetic preferences and the target text content, improving the video quality of the generated target video, and making the generated target video more popular with humans. This solves the technical problem in the related art of training a video generation model based on network data, resulting in poor quality of the video generated by the trained video generation model, which does not meet user expectations.
[0206] Those skilled in the art will appreciate that the structure shown in FIG10 is merely illustrative, and the computer terminal A may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG10 does not limit the structure of the aforementioned electronic devices. For example, the computer terminal A may include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG10 , or may have a configuration different from that shown in FIG10 .
[0207] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0208] Example 6
[0209] The embodiment of the present disclosure further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the video generation method provided in the first embodiment.
[0210] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0211] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining a target text, wherein the target text is used to describe the video content to be generated; performing video generation processing on the target text using a target video generation model to obtain a target video, wherein the target video generation model is a model obtained by aligning the initial video generation model with the preset reward model using a fine-tuning method.
[0212] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: sampling the initial video generation model using training samples to generate a sampled video, wherein the training samples include: multiple video-text pairs, and the multiple video-text pairs each include: a training video and a training text, and the training text is used to describe the video content of the training video; using a preset reward model to calculate the reward for the sampled video to obtain a target reward result; adjusting the model parameters of the initial video generation model based on the target reward result to generate a target video generation model.
[0213] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing noise processing on the training video to obtain a noisy video; and performing video resampling on the initial video generation model using the noisy video and the training text to generate a sampled video.
[0214] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining the number of noise addition steps and the noise level corresponding to the training video, wherein the number of noise addition steps is used to determine the number of steps to be noised in the training video through a preset noise addition function, and the noise level is used to determine the degree of damage to the training video; performing noise addition processing on the training video based on the number of noise addition steps and the noise level to obtain a noisy video.
[0215] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing reward calculations on the sampled video using a preset reward model to obtain an initial reward result; adjusting an initial reward weight corresponding to the initial reward result using a time-decay reward method to generate a target reward result, wherein the initial reward weight is a default reward weight corresponding to a video frame sequence contained in the sampled video.
[0216] Optionally, in this embodiment, the computer-readable storage medium is configured to store program codes for executing the following steps: performing video segment sampling on the video frame sequence to obtain segment sampling results; and performing reward calculation on the segment sampling results using a preset reward model to obtain an initial reward result.
[0217] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining a feature space representation of a video frame sequence; performing video segmentation sampling on the feature space representation to obtain a color space representation of the segmented video frame; and determining the segmentation sampling result based on the color space representation.
[0218] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: using a time-decay reward method to differentially adjust the initial reward weight corresponding to the initial reward result to obtain a target reward weight, wherein the target reward weight is used to indicate that the current reward weight of the first video frame in the video frame sequence is higher than the current reward weight of the second video frame, the first video frame is located in the middle position of the video frame sequence, and the second video frame is located at the edge position of the video frame sequence; generating the target reward result based on the initial reward result and the target reward weight.
[0219] Optionally, the target video generation model is a video diffusion model, and the preset reward model is an image reward model, wherein the image reward model is used to perform preference learning on the video diffusion model.
[0220] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.
[0221] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0222] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0223] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0224] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0226] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.
Claims
1. A video generation method, comprising: Obtaining a target text, wherein the target text is used to describe the video content to be generated; The target text is processed by video generation using a target video generation model to obtain a target video, wherein the target video generation model is a model obtained by aligning an initial video generation model with a preset reward model using a fine-tuning method.
2. The video generation method according to claim 1, wherein: The video generation method further includes: The initial video generation model is sampled using training samples to generate sampled videos, wherein the training samples include: a plurality of video-text pairs, each of the plurality of video-text pairs includes: a training video and a training text, and the training text is used to describe the video content of the training video; Using the preset reward model to calculate the reward for the sampled video to obtain a target reward result; The model parameters of the initial video generation model are adjusted based on the target reward result to generate the target video generation model.
3. The video generation method according to claim 2, wherein: The initial video generation model is sampled using the training sample to generate the sampled video, comprising: Performing noise processing on the training video to obtain a noisy video; The noisy video and the training text are used to perform video resampling on the initial video generation model to generate the sampled video.
4. The video generation method according to claim 3, wherein: The training video is subjected to noise processing to obtain the noise-added video, comprising: Obtaining the number of noise adding steps and the noise level corresponding to the training video, wherein the number of noise adding steps is used to determine the number of steps to be noise added to the training video through a preset noise adding function, and the noise level is used to determine the degree of damage to the training video; The training video is subjected to noise addition processing based on the noise addition step number and the noise level to obtain the noisy video.
5. The video generation method according to claim 2, wherein: The preset reward model is used to calculate the reward for the sampled video to obtain the target reward result, including: Using the preset reward model to calculate the reward for the sampled video to obtain an initial reward result; An initial reward weight corresponding to the initial reward result is adjusted by using a time decay reward method to generate the target reward result, wherein the initial reward weight is a default reward weight corresponding to a video frame sequence included in the sampled video.
6. The video generation method according to claim 5, wherein: The preset reward model is used to calculate the reward for the sampled video, and the initial reward result is obtained, including: Performing video segment sampling on the video frame sequence to obtain segment sampling results; The preset reward model is used to calculate the reward for the segmented sampling result to obtain the initial reward result.
7. The video generation method according to claim 6, wherein: Performing video segment sampling on the video frame sequence to obtain the segment sampling result includes: Obtaining a feature space representation of the video frame sequence; Performing video segment sampling on the feature space representation to obtain a color space representation of the segmented video frame; The segment sampling result is determined based on the color space representation.
8. The video generation method according to claim 7, wherein: The time decay reward method is used to adjust the initial reward weight corresponding to the initial reward result to generate the target reward result, including: The time decay reward method is used to differentially adjust the initial reward weight corresponding to the initial reward result to obtain a target reward weight, wherein the target reward weight is used to indicate that the current reward weight of the first video frame in the video frame sequence is higher than the current reward weight of the second video frame, the first video frame is located in the middle of the video frame sequence, and the second video frame is located at the edge of the video frame sequence; The target reward result is generated based on the initial reward result and the target reward weight.
9. The video generation method according to claim 1, wherein: The target video generation model is a video diffusion model, and the preset reward model is an image reward model, wherein the image reward model is used to perform preference learning on the video diffusion model.
10. A video generation method, comprising: Obtaining a video generation call request through a first application programming interface, wherein the request data carried in the video generation call request includes: a target text, the target text is used to describe the video to be generated The video generation call request is used to request to call a target video generation model to perform video generation processing on the target text to obtain a target video, and the target video generation model is a model obtained by aligning the initial video generation model with the preset reward model in a fine-tuning manner; A video generation call response is returned through the second application programming interface, wherein the response data carried in the video generation call response includes: the target video.
11. A video generation method, comprising: Acquire a currently input video generation dialogue request, wherein the information carried in the video generation dialogue request includes: a target text, wherein the target text is used to describe the video content to be generated; In response to the video generation dialogue request, a video generation dialogue reply is returned, wherein the information carried in the video generation dialogue reply includes: a target video, the target video is obtained by performing video generation processing on the target text using a target video generation model, and the target video generation model is a model obtained by aligning the initial video generation model with the preset reward model in a fine-tuning manner; Presenting the video within a graphical user interface generates a conversation response.
12. An electronic device comprising: A memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 11 when running.
13. A computer-readable storage medium, the computer-readable storage medium comprising a stored executable program, wherein: When the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method, model and device for training text graph model, and electronic equipment
CN116894880A
Picture generation method and device, storage medium and electronic equipment
CN116958969A
Generative artificial intelligence-based novel tweet video generation method and system
CN117078782A
Text-image generation method, system and device and storage medium
CN117095083A
Video generation method, electronic equipment and computer readable storage medium
CN117668297A
Cited By
Text feature generation model training method and device, equipment and storage medium
CN120494017A
Model adjustment method and device, video generation method and related equipment
CN120935381A
Model training method, video generation method, electronic equipment and storage medium
CN120953453A
Video generation model training method and device and storage medium
CN121190645A
A video generation model training method and device, and a storage medium
CN121190645B