Training Method of Video Generation Model, Video Generation Method, Device, Electronic Device and Readable Storage Medium
By building a generative adversarial network and training a generative model, the problem of unreal and unsmooth video frames in action migration technology is solved, and the high-quality generation of target videos is achieved.
Patent Information
- Application Number
- CN202210581262.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-05-26
AI Technical Summary
When existing action migration techniques migrate actions in the source video to the target image, it is easy to cause problems such as unrealistic video frames, blurred picture and unsmooth video frames in the target video, resulting in poor generation results.
A generative adversarial network is built, including a generative model and a discriminative model. By obtaining multiple sample videos, inputting them into the generative model to generate predicted video frames, and inputting the predicted video frames and sample videos to the discriminative model for discrimination. The generative adversarial network is trained based on the discriminative results until the training stop condition is met, and a video generation model is obtained.
Improve the authenticity and coherence of target video generation, ensuring the authenticity of details in target video frames and the coherence of generated target videos.
Smart Images

Figure CN115063713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method for training a video generation model, a video generation method, a device, an electronic device, and a readable storage medium. Background Art
[0002] Currently, the action migration technology is to migrate the actions in the source video to the target image to generate a target video, and its effect is to make the object in the target image show the actions in the source video. It can be applied to various scenarios such as social entertainment and special effect synthesis.
[0003] Since the poses of the objects in the source video and the target image may vary greatly, there may be problems such as unrealistic single video frames, blurred images, and unsmoothness between video frames in the target video generated by the current action migration technology, that is, the effect of generating the target video is not good. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method for training a video generation model, a video generation method, a device, an electronic device, and a readable storage medium, which can improve the effect of generating a target video through the trained video generation model. The specific technical solutions are as follows:
[0005] In the first aspect of the present invention, first, a method for training a video generation model is provided, including:
[0006] Obtain a plurality of sample videos;
[0007] Construct a generative adversarial network, and the generative adversarial network includes a generative model and a discriminative model;
[0008] Input the sample video into the generative model to obtain a predicted video frame;
[0009] Input the predicted video frame and the sample video into the discriminative model to obtain a discrimination result; the discriminative model is used to discriminate whether the predicted video frame matches the sample video;
[0010] Train the generative adversarial network based on the discrimination results of each sample video until the training stop condition is met, and obtain a video generation model.
[0011] In the second aspect of the present invention, first, a video generation method is provided, including:
[0012] Obtain a video frame sequence, and the video frame sequence includes: video frames of the source video and a target image;
[0013] Input a video frame sequence into a video generation model, where the video generation model includes: an image generation model and an optical flow network model; extract the foreground features of the video frame sequence through the image generation model, and extract the optical flow features of the source video through the optical flow network model;
[0014] Perform feature fusion on the foreground features and the optical flow features to generate a target video frame;
[0015] Generate a target video based on the target video frame.
[0016] In the third aspect of the implementation of the present invention, there is also provided a training device for a video generation model, including:
[0017] A first acquisition module, configured to acquire a plurality of sample videos;
[0018] A construction module, configured to construct a generative adversarial network, where the generative adversarial network includes a generative model and a discriminative model;
[0019] A first input module, configured to input the sample video into the generative model to obtain a predicted video frame;
[0020] The first input module is further configured to input the predicted video frame and the sample video into the discriminative model to obtain a discrimination result; the discriminative model is used to discriminate whether the predicted video frame matches the sample video;
[0021] A training module, configured to train the generative adversarial network based on the discrimination results of each sample video until a training stop condition is met, to obtain a video generation model.
[0022] In the fourth aspect of the implementation of the present invention, there is also provided a video generation device, including:
[0023] A second acquisition module, configured to acquire a video frame sequence, where the video frame sequence includes: video frames of the source video and a target image;
[0024] A second input module, configured to input the video frame sequence into a video generation model, where the video generation model includes: an image generation model and an optical flow network model; extract the foreground features of the video frame sequence through the image generation model, and extract the optical flow features of the source video through the optical flow network model;
[0025] A fusion module, configured to perform feature fusion on the foreground features and the optical flow features to generate a target video frame;
[0026] A generation module, configured to generate a target video based on the target video frame.
[0027] In yet another aspect of the implementation of the present invention, there is also provided a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is caused to execute any one of the above methods.
[0028] In yet another aspect of the implementation of the present invention, there is also provided a computer program product including instructions, which, when running on a computer, causes the computer to execute any one of the above-mentioned methods.
[0029] In the embodiments of the present invention, by obtaining a plurality of sample videos and constructing a generative adversarial network including a generative model and a discriminative model, first, the sample videos are input into the generative model to obtain predicted video frames, and then the predicted video frames and the sample videos are input into the discriminative model for determining whether the predicted video frames match the sample videos. Here, the authenticity of the predicted video frames and the smooth coherence of the connection between the predicted video frames and the sample videos can be continuously improved. Finally, the generative adversarial network is trained based on the discrimination results of each sample video until the training stop condition is met, and a video generation model is obtained. Thus, the target video frames generated based on the trained video generation model can ensure the authenticity of the details in the target video frames and the coherence of the target video generated based on the target video frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0031] Figure 1 It is a schematic structural diagram of a video generation model provided by an embodiment of the present invention;
[0032] Figure 2 It is a flowchart of a training method of a video generation model provided by an embodiment of the present invention;
[0033] Figure 3 It is a flowchart of a video generation method provided by an embodiment of the present invention;
[0034] Figure 4 It is a schematic diagram of a multi-scale network structure provided by an embodiment of the present invention;
[0035] Figure 5 It is a schematic structural diagram of a training device of a video generation model provided by an embodiment of the present invention;
[0036] Figure 6 It is a schematic structural diagram of a video generation device provided by an embodiment of the present invention;
[0037] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following will describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention.
[0039] The video generation method provided by the embodiments of the present invention can be applied to at least the following application scenarios, which will be described below.
[0040] Currently, the action transfer technology is to transfer the actions in the source video to the target image to generate a target video, and its effect is to make the objects in the target image show the actions in the source video. It can be applied to various scenarios such as social entertainment and special effect synthesis. For example, transfer the motion information of professional dancers in the source video to the body of amateurs and render to generate a target video. In the target video generated by action transfer, amateurs can learn to dance different styles like professional dancers. Dance video generation is the combination of action transfer and video generation. How to ensure the stable transfer of actions and make the stability of the generated target video frames and the continuity of the generated target video is a problem that needs to be solved.
[0041] Based on the above application scenarios, the video generation method provided by the embodiments of the present invention will be described in detail below.
[0042] First, an overall description of the video generation model structure provided by the embodiments of the present invention will be given below.
[0043] Figure 1 It is a schematic diagram of a video generation model structure provided by the embodiments of the present invention. As Figure 1 shown, the video generation model includes: a generation model and a discriminant model. The generation model includes: an image generation model and an optical flow network model; the discriminant model includes: an image discriminant model and a video discriminant model.
[0044] During the training process, by obtaining a plurality of sample videos and constructing a generative adversarial network including a generation model and a discriminant model, first input the sample videos into the generation model to obtain predicted video frames. Among them, the sample videos include a first video frame and a plurality of second video frames adjacent to the first video frame; specifically, input the plurality of second video frames and the sample pose information extracted from the sample videos into the generation model, extract the foreground training features of the plurality of second video frames through the image generation model; and extract the optical flow training features of the plurality of second video frames through the optical flow network model. Finally, fuse the foreground training features output by the image generation model and the optical flow training features output by the optical flow network model to obtain the fused predicted video frames. Then input the predicted video frames and the sample videos into the discriminant model for determining whether the predicted video frames match the sample videos.
[0045] Here, it is possible to continuously improve the authenticity of the predicted video frames and the smooth coherence of the connection between the predicted video frames and the sample videos. Finally, a generative adversarial network is trained based on the discrimination results of each sample video until the training stop condition is met, and a video generation model is obtained. Thus, the target video frames generated based on the trained video generation model can ensure the authenticity of the details in the target video frames and the coherence of the target video generated based on the target video frames.
[0046] In the application process, by inputting a video frame sequence including the video frames of the source video and the target image into the video generation model, the foreground features of the video frame sequence are extracted through the image generation model. Specifically, the target image in the video frame sequence is used as input A (in the case where the target video frame has been generated, input A also includes the generated target video frame), and input into the image generation model to extract the foreground features of the video frame sequence through the image generation model; on the other hand, the pose information extracted from the video frames of the source video is used as input B, input into the optical flow network model, and the optical flow features of the source video are extracted through the optical flow network model.
[0047] Here, the foreground features with authenticity and stability can be extracted through the image generation model, and the optical flow features of the source video can be extracted through the optical flow network model, and the stable and continuous optical flow features can be extracted through the optical flow network model; by fusing the foreground features and the optical flow features, real and stable target video frames can be generated; thus, a real and stable target video can be generated based on the target video frames, improving the generation effect of the target video.
[0048] Next, the training structure of the video generation model provided by the embodiments of the present invention will be described.
[0049] Figure 2 is a flowchart of a training method for a video generation model provided by an embodiment of the present invention;
[0050] As Figure 2 shown, the training method of the video generation model may include step 210 - step 250, specifically as follows:
[0051] Step 210, obtain a plurality of sample videos.
[0052] Step 220, construct a generative adversarial network, and the generative adversarial network includes a generative model and a discriminative model.
[0053] Step 230, input the sample video into the generative model to obtain predicted video frames.
[0054] Step 240, input the predicted video frames and the sample video into the discriminative model to obtain a discrimination result; the discriminative model is used to discriminate whether the predicted video frames match the sample video.
[0055] Step 250: Train a generative adversarial network based on the discrimination results of each sample video until the training stop condition is met, and obtain a video generation model.
[0056] In summary, in the embodiments of the present invention, by obtaining multiple sample videos and constructing a generative adversarial network including a generation model and a discrimination model, first input the sample videos into the generation model to obtain predicted video frames, and then input the predicted video frames and the sample videos into the discrimination model for determining whether the predicted video frames match the sample videos. Here, the authenticity of the predicted video frames and the smooth coherence of the connection between the predicted video frames and the sample videos can be continuously improved. Finally, train the generative adversarial network based on the discrimination results of each sample video until the training stop condition is met, and obtain a video generation model. Thus, for the target video frames generated based on the trained video generation model, the authenticity of the details in the target video frames can be ensured, and the coherence of the target video generated based on the target video frames can also be ensured.
[0057] The specific implementation manners of the above steps are introduced below.
[0058] Regarding step 210.
[0059] Obtain multiple sample videos.
[0060] Exemplarily, three consecutive video frames x (which can be any number of frames) can be extracted frame by frame from the sample video T-2 , x T-1 , x T .
[0061] Regarding step 220.
[0062] Construct a generative adversarial network, where the generative adversarial network includes a generation model and a discrimination model.
[0063] A generative adversarial network (GAN, Generative Adversarial Networks) is a deep learning model. The model generates quite good outputs through the mutual game learning of at least two modules in the framework: a generative model (Generative Model) and a discrimination model (Discriminative Model).
[0064] During the training process, the goal of the generative model is to generate as realistic pictures as possible to deceive the discrimination model. And the goal of the discrimination model is to distinguish the pictures generated by the generative model from the real pictures. In this way, the generative model and the discrimination model constitute a dynamic "game process". Finally, the result of the game is that the generative model can generate pictures that are "indistinguishable from the real ones". For the discrimination model, it is difficult to determine whether the pictures generated by the generative model are real or not.
[0065] Relates to step 230.
[0066] Input the sample video into the generation model to obtain predicted video frames.
[0067] To ensure the authenticity of details in the generated predicted video frames, the generation model involved in the present invention consists of an image generation model and an optical flow network model. The image generation model is used to generate foreground features of motion, and the optical flow network model is used to generate optical flow features of stillness. Finally, the output results of the image generation model and the optical flow network model are fused as the finally generated predicted video frames. Additionally, considering that a video is a synthesis of the spatial domain and the temporal domain. Therefore, the discriminant model involved in the present invention consists of an image discriminant model and a video discriminant model. The image discriminant model is used to identify the authenticity of a single-frame predicted video frame. The video discriminant model is used to identify the coherence of consecutive predicted video frames.
[0068] Specifically, among them, the generation model includes: an image generation model and an optical flow network model; the discriminant model includes: an image discriminant model and a video discriminant model; the sample video includes a first video frame and a plurality of second video frames adjacent to the first video frame; specifically, the sample video includes a first video frame (x T ), and a plurality of second video frames (x T-2 , x T-1 ) adjacent to the first video frame.
[0069] Step 230 may specifically include the following steps:
[0070] Input a plurality of second video frames into the image generation model to output foreground training features;
[0071] Input the sample pose information extracted from the sample video into the optical flow network model to output optical flow training features;
[0072] Fuse the foreground training features and the optical flow training features to obtain predicted video frames.
[0073] Among them, inputting the sample pose information extracted from the sample video into the optical flow network model to output optical flow training features may specifically include the following steps:
[0074] Extract sample pose information from the sample video frames of the sample video respectively; input the sample pose information into the optical flow network model to extract optical flow training features.
[0075] Among the steps of fusing the foreground training features and the optical flow training features to obtain predicted video frames as mentioned above, it may specifically include the following steps:
[0076] Generate the first pixel point according to the foreground training feature; generate the second pixel point according to the optical flow training feature; synthesize the first pixel point and the second pixel point to obtain the predicted video frame.
[0077] Among them, the sample video includes the target object, and the target object can be a dynamic person, animal, mobile device, etc. Specifically, the first pixel point and the second pixel point can be synthesized according to the position of the target object in the video frame of the sample video to obtain the predicted video frame.
[0078] First, input multiple second video frames (x T-2 , x T-1 ) into the image generation model to output the foreground training feature; then, input the sample pose information (s T-2 , s T-1 , x T ) extracted from the sample video (x T-2 , s T-1 ,, s T ) into the optical flow network model to output the optical flow training feature; finally, fuse the foreground training feature output by the image generation model and the optical flow training feature output by the optical flow network model to obtain the fused predicted video frame (x T *).
[0079] Involve step 240.
[0080] Input the predicted video frame and the sample video into the discriminant model to obtain the discriminant result; the discriminant model is used to determine whether the predicted video frame matches the sample video.
[0081] Step 240 may specifically include the following steps:
[0082] Input the predicted video frame and the first video frame into the image discrimination model to obtain the first loss value;
[0083] Input the predicted video frame and the second video frame into the video discrimination model to obtain the second loss value;
[0084] Correspondingly, step 250 may specifically include the following steps:
[0085] Train the generative adversarial network according to the first loss value and the second loss value until the training stop condition is met to obtain the video generation model.
[0086] Input the predicted video frame (x T *) and the first video frame (x T ) into the image discrimination model, and calculate the first loss value of the image discrimination model. Input the predicted video frame (X T-2 *, X T-1 *, X T *) and the second video frame (X T-2,X T-1 ,X T ), input it into the video discrimination model, and calculate the second loss value of the video discrimination model. Determine the discrimination result according to the first loss value and the second loss value.
[0087] In addition, in the case of determining the predicted video frame of the first frame, X T-2 *,X T-1 * does not exist. In the process of inputting the predicted video frame and the second video frame into the video discrimination model, the corresponding predicted video frame (X T-2 *,X T-1 *) can be replaced by X T-2 and X T-1 .
[0088] Among them, the discrimination result determined according to the first loss value and the second loss value is used to train the generative adversarial network until the training stop condition is satisfied, and a video generation model is obtained. Specifically, it may include:
[0089] Determine the loss function value of the generative adversarial network according to the discrimination results determined by the first loss values and the second loss values of each sample video, train the generative adversarial network, and obtain the video generation model when the loss function value of the generative adversarial network obtained according to the discrimination results satisfies the loss threshold.
[0090] Before step 240, the following steps may also be included:
[0091] Calculate the third loss value of the image generation model according to the predicted video frame and the first video frame;
[0092] Calculate the fourth loss value of the optical flow network model according to the optical flow training feature and the optical flow ground truth, and the optical flow ground truth is extracted from the sample video through a preset optical flow extraction algorithm;
[0093] Determine the loss value of the generation model according to the third loss value and the fourth loss value;
[0094] Here, to calculate the loss value of the generator generation model, on the one hand, calculate the third loss value of the image generation model according to the predicted video frame and the first video frame, and the numerical value of this third loss value is the same as the first loss value involved above. On the other hand, the optical flow ground truth is extracted from the sample video based on the preset optical flow extraction algorithm, and the fourth loss value of the optical flow network model is calculated according to the optical flow training feature and the optical flow ground truth. Finally, determine the loss value of the generation model according to the third loss value of the image generation model and the fourth loss value of the flow network model.
[0095] Correspondingly, step 250 may specifically include the following steps:
[0096] Determine the loss value of the discrimination model according to the first loss value and the second loss value;
[0097] Based on the loss value of the generation model and the loss value of the discriminative model, perform backpropagation training on the generative adversarial network until the generative adversarial network meets the preset convergence condition, and obtain a trained video generation model.
[0098] Specifically, determine the sum of the third loss value and the fourth loss value as the loss value of the generation model; determine the sum of the first loss value and the second loss value as the loss value of the discriminative model; perform backpropagation training on the generative adversarial network according to the loss value of the generation model and the loss value of the discriminative model.
[0099] Among them, backpropagation is a process of iteratively optimizing the loss value using the gradient descent method to find the minimum value. The model algorithm optimization process is to use a suitable loss function to measure the output loss of the training samples, optimize the loss function to find the minimum value, that is, the process of iteratively optimizing the loss function of the deep neural network using the gradient descent method for a series of corresponding linear coefficient matrices to find the minimum value is the backpropagation algorithm, and various loss functions and activation functions can be used.
[0100] Involve step 250.
[0101] Train the generative adversarial network based on the discrimination results of each training sample until the training stop condition is met, and obtain a video generation model.
[0102] Here, for the trained generative adversarial network, the generation model can generate target video frames that are "realistic enough to pass for real". For the discriminative model, on the one hand, it is difficult to determine whether the target video frames generated by the generation model are actually real; on the other hand, it is also difficult to determine whether the target video frames generated by the generation model match the source video, that is, it is also difficult to determine whether the target video frames originally exist in the source video. Thus, the target video frames generated by the trained generative adversarial network can ensure the authenticity of the details in the target video frames and also ensure the coherence of the target video generated based on the target video frames.
[0103] In a possible embodiment, step 250 may specifically include the following steps:
[0104] Input the predicted video frames and the first video frames into the feature extraction network. The feature extraction network includes multiple scale layers, and each scale layer outputs sub-loss values of the predicted video frames and the first video frames respectively;
[0105] Determine the multi-scale loss value according to the multiple sub-loss values;
[0106] Train the generative adversarial network according to the multi-scale loss value, the first loss value, and the second loss value until the training stop condition is met, and obtain a video generation model.
[0107] The predicted video frame and the first video frame are respectively input into different scale layers of a feature extraction network (such as VGG), and each scale layer outputs the sub-loss values of the predicted video frame and the first video frame at that scale; according to multiple sub-loss values, a multi-scale loss value is determined; correspondingly, the generative adversarial network is trained according to the multi-scale loss value, the first loss value, and the second loss value until the training stop condition is met, and a video generation model is obtained.
[0108] Among them, the multi-scale loss value can be specifically calculated by the following formula:
[0109]
[0110] Among them, L featrure is the multi-scale loss value, where T represents the total number of scale layers, X is the predicted video frame, Y is the first video frame, and f i represents feature extraction, and f i (X) - f i (Y) is the sub-loss value.
[0111] In addition, as Figure 4 shown, in the multi-scale feature extraction network, the encoding network 1 and the decoding network 1 are low-scale networks. The network input input1 is 256 * 256, and the output output1 is also 256 * 256. The network corresponding to this scale layer is used to generate a global low-resolution video. The encoding 2 and decoding 2 are high-scale networks, and the scales of input2 and output2 are both 512 * 512. The high-scale network is used to generate a locally refined high-resolution video.
[0112] Thus, by designing a multi-scale network in the generation model, a high-resolution video can be generated, making the generated video more refined; in the discriminant model, by designing a multi-scale feature loss value, the details of the generated predicted video frame can be made more real.
[0113] In summary, for the video generation model in the embodiment of the present invention, by obtaining multiple sample videos and constructing a generative adversarial network including a generation model and a discriminant model, first, the sample videos are input into the generation model to obtain predicted video frames, and then the predicted video frames and the sample videos are input into the discriminant model for determining whether the predicted video frames match the sample videos. Here, the authenticity of the predicted video frames and the smooth coherence of the connection between the predicted video frames and the sample videos can be continuously improved. Finally, the generative adversarial network is trained based on the discrimination results of each sample video until the training stop condition is met, and a video generation model is obtained. Thus, for the target video frames generated based on the trained video generation model, the authenticity of the details in the target video frames can be guaranteed, and the coherence of the target video generated based on the target video frames can also be guaranteed.
[0114] The video generation method provided by the embodiments of the present invention will be described below.
[0115] Figure 3 It is a flowchart of a video generation method provided by an embodiment of the present invention.
[0116] As Figure 3 shown, the video generation method may include step 310-step 340. This method is applied to a video generation device and is specifically as follows:
[0117] Step 310, obtain a video frame sequence, where the video frame sequence includes: video frames of the source video and a target image.
[0118] Step 320, input the video frame sequence into a video generation model, where the video generation model includes: an image generation model and an optical flow network model; extract the foreground features of the video frame sequence through the image generation model, and extract the optical flow features of the source video through the optical flow network model.
[0119] Step 330, perform feature fusion on the foreground features and the optical flow features to generate a target video frame.
[0120] Step 340, generate a target video based on the target video frame.
[0121] In the embodiments of the present disclosure, by inputting a video frame sequence including video frames of the source video and a target image into the video generation model, the foreground features of the video frame sequence are extracted through the image generation model. Here, the foreground features with authenticity and stability can be extracted through the image generation model, and the optical flow features of the source video are extracted through the optical flow network model. Here, stable and continuous optical flow features can be extracted through the optical flow network model; perform feature fusion on the foreground features and the optical flow features to generate a real and stable target video frame; thus, a real and stable target video can be generated based on the target video frame, improving the generation effect of the target video.
[0122] The specific implementation manners of the above steps will be introduced below.
[0123] Regarding step 310.
[0124] Obtain a video frame sequence, where the video frame sequence includes: video frames of the source video and a target image.
[0125] Among them, the video frames of the source video may include: video frames y0-y T , and the target image may include: images z0, z1 of the target person. The video frame sequence is {z0, z1, y0-y T}.
[0126] Regarding step 320.
[0127] Input a video frame sequence into a video generation model, which includes an image generation model and an optical flow network model; extract the foreground features of the video frame sequence through the image generation model, and extract the optical flow features of the source video through the optical flow network model.
[0128] On the one hand, input (z0, z1) in the video frame sequence into the image generation model to extract the foreground features of the video frame sequence through the image generation model; on the other hand, extract the pose information (s0, s1, s2) from the video frames (y0 - y T ) of the source video in the video frame sequence, input the pose information into the optical flow network model, and extract the optical flow features of the source video through the optical flow network model.
[0129] Among them, in the step of extracting the foreground features of the video frame sequence through the image generation model, it may specifically include the following steps:
[0130] Extract the action features of the source video and the shape features of the target image through the image generation model;
[0131] Generate foreground features based on the action features of the source video and the shape features of the target image.
[0132] Thus, the action of the object in the source video can be migrated to the object in the target image to generate foreground features.
[0133] Here, by designing the structures of the image generation model and the optical flow network model in the video generation model, the optical flow network model is used to model the optical flow features of the background, and the image generation model is used to model the foreground features, and then the two model branches are fused. While ensuring the smoothness between frames of the generated predicted video frames and avoiding frame skipping, it also ensures the accuracy of migrating the action in the source video to the object in the target image.
[0134] Involve step 330.
[0135] Perform feature fusion on the foreground features and the optical flow features to generate the target video frame (y0*).
[0136] Each video frame in the video frame sequence includes a timing identifier; the video frame sequence includes a first sequence and a second sequence. The first sequence includes the target image and the generated target video frames; the second sequence includes the video frames of the source video. Extracting the foreground features of the video frame sequence through the image generation model and extracting the optical flow features of the source video through the optical flow network model includes:
[0137] According to the target timing identifier of the target video frame to be generated, obtain the third video frame corresponding to the timing identifier adjacent to the target timing identifier from the first sequence;
[0138] Obtain the fourth video frame corresponding to the timing identifier adjacent to the target timing identifier from the second sequence according to the target timing identifier;
[0139] Extract the foreground feature of the third video frame through an image generation model, and extract the optical flow feature of the fourth video frame through an optical flow network model.
[0140] To make the transition between video frames of the generated target video smoother, when a preset number of target video frames (y0*, …, y T *) are generated, the generated target video frames can be added to the video frame sequence.
[0141] For example, a preset number of first target video frames have been generated. The first target video frames are generated by extracting features from the video frames of the source video and the target image through a trained image generation model. The trained image generation model can ensure the smoothness between the first target video frames and the video frame sequence.
[0142] Then, input the video frame sequence including the first target video frames into the video generation model to obtain the predicted second target video frames. The trained image generation model can also ensure the smoothness between the second target video frames and the video frame sequence. Here, the video frame sequence includes the first target video frames, so the smoothness between the first target video frames and the second target video frames is ensured.
[0143] Since the final target video is composed of target video frames (the first target video frames, the second target video frames, ……), ensuring the inter-frame smoothness between the first target video frames and the second target video frames can improve the inter-frame smoothness and coherence of the target video.
[0144] Specifically, after feature fusion of the foreground feature and the optical flow feature to generate a target video frame, the following steps may further be included: updating the video frame sequence based on the target video frame to obtain an updated video frame sequence.
[0145] In this way, when generating the predicted video frame next time, according to the target timing identifier (t) of the target video frame (y t ) to be generated, obtain the third video frame (y t-2 *,y t-1 *) corresponding to the timing identifier adjacent to the target timing identifier from the first sequence (z0,z1, y0*, …, y t-2 *,y t-1 *);
[0146] According to the target timing identifier, obtain the fourth video frame (y T ) corresponding to the timing identifier adjacent to the target timing identifier from the second sequence (y0 - y t-2 ,y t-1,y t ), and extract the pose information (s from the fourth video frame t-2 ,s t-1 ,s t );
[0147] Among them, specifically, the fourth video frame can be input into a preset pose recognition network to obtain the pose information.
[0148] Extract the foreground features of the third video frame through an image generation model, and extract the pose information (s from the fourth video frame t-2 ,s t-1 ,s t ), and then extract the optical flow features of the pose information through an optical flow network model.
[0149] In a possible embodiment, in the step of performing feature fusion on the foreground features and the optical flow features to generate the target video frame, the following steps may specifically be included:
[0150] Obtain a weight image, where the weight image is used to represent the weight value corresponding to the pixel points in the target video frame;
[0151] Based on the weight image, perform weighted fusion processing on each pixel point in the foreground features and each pixel point in the optical flow features to generate the target video frame, where the foreground features and the optical flow features may be images, and the image sizes of the foreground features and the optical flow features match.
[0152] It involves step 340.
[0153] Generate a target video based on the target video frame.
[0154] After respectively outputting the target video frames, generate a target video based on the target video frames (y0*,...,y T *).
[0155] In summary, in the embodiments of the present invention, by inputting a video frame sequence including a source video and a target image into a video generation model, extracting the foreground features of the video frame sequence through an image generation model, where the foreground features with authenticity and stability can be extracted through the image generation model, and extracting the optical flow features of the source video through an optical flow network model, where stable and continuous optical flow features can be extracted through the optical flow network model; performing feature fusion on the foreground features and the optical flow features can generate real and stable target video frames; thus, a real and stable target video can be generated based on the target video frames, improving the generation effect of the target video.
[0156] Based on the above Figure 2 shown training method of the video generation model, the embodiments of the present invention further provide a training device for the video generation model, as Figure 5As shown in the figure, the training device 500 of the video generation model may include:
[0157] A first acquisition module 510, configured to acquire a plurality of sample videos.
[0158] A construction module 520, configured to construct a generative adversarial network, where the generative adversarial network includes a generative model and a discriminative model.
[0159] A first input module 530, configured to input the sample video into the generative model to obtain a predicted video frame.
[0160] The first input module 530 is further configured to input the predicted video frame and the sample video into the discriminative model to obtain a discrimination result; the discriminative model is used to determine whether the predicted video frame matches the sample video.
[0161] A training module 540, configured to train the generative adversarial network based on the discrimination results of each sample video until a training stop condition is met, to obtain a video generation model.
[0162] In a possible embodiment, the generative model includes: an image generation model and an optical flow network model; the discriminative model includes: an image discriminative model and a video discriminative model; the sample video includes a first video frame and a plurality of second video frames adjacent to the first video frame; the first input module 530 is specifically configured to:
[0163] Input a plurality of second video frames into the generative model, and extract foreground training features of the plurality of second video frames through the image generation model; and extract optical flow training features of the plurality of second video frames through the optical flow network model;
[0164] Fuse the foreground training features and the optical flow training features to obtain a predicted video frame.
[0165] In a possible embodiment, the first input module 530 is specifically configured to:
[0166] Input the predicted video frame and the first video frame into the image discrimination model to obtain a first loss value;
[0167] Input the predicted video frame and the second video frame into the video discrimination model to obtain a second loss value;
[0168] The training module 540 is specifically configured to: train the generative adversarial network according to the first loss value and the second loss value until a training stop condition is met, to obtain a video generation model.
[0169] In a possible embodiment, the training device 500 of the video generation model may further include:
[0170] A calculation module, configured to calculate a third loss value of the image generation model according to the predicted video frame and the first video frame.
[0171] A calculation module is further configured to calculate a fourth loss value of the optical flow network model according to the optical flow training features and the optical flow ground truth, where the optical flow ground truth is extracted from the sample video through a preset optical flow extraction algorithm.
[0172] A determination module is configured to determine the loss value of the generation model according to the third loss value and the fourth loss value.
[0173] The training module 540 is specifically configured to:
[0174] Determine the loss value of the discriminant model according to the first loss value and the second loss value;
[0175] Perform backpropagation training on the generative adversarial network according to the loss value of the generation model and the loss value of the discriminative model until the generative adversarial network meets the preset convergence condition, and obtain a trained video generation model.
[0176] In a possible embodiment, the training module 540 is specifically configured to:
[0177] Input the predicted video frame and the first video frame into the feature extraction network, where the feature extraction network includes multiple scale layers, and each scale layer outputs a sub-loss value of the predicted video frame and the first video frame respectively;
[0178] Determine the multi-scale loss value according to the multiple sub-loss values;
[0179] Train the generative adversarial network according to the multi-scale loss value, the first loss value and the second loss value until the training stop condition is met, and obtain a video generation model.
[0180] In summary, in the embodiments of the present invention, by obtaining multiple sample videos and constructing a generative adversarial network including a generation model and a discriminant model, first inputting the sample videos into the generation model to obtain predicted video frames, and then inputting the predicted video frames and the sample videos into the discriminant model for determining whether the predicted video frames match the sample videos. Here, the authenticity of the predicted video frames and the smooth coherence of the connection between the predicted video frames and the sample videos can be continuously improved. Finally, train the generative adversarial network based on the discrimination results of each sample video until the training stop condition is met, and obtain a video generation model. Thus, the target video frames generated based on the trained video generation model can ensure the authenticity of the details in the target video frames and the coherence of the target video generated based on the target video frames.
[0181] Based on the above Figure 3 shown video generation method, an embodiment of the present invention further provides a video generation device, as Figure 6 shown, the video generation device 600 may include:
[0182] The second acquisition module 610 is configured to acquire a video frame sequence, where the video frame sequence includes video frames of a source video and a target image.
[0183] The second input module 620 is configured to input the video frame sequence into a video generation model, where the video generation model includes an image generation model and an optical flow network model; extract foreground features of the video frame sequence through the image generation model, and extract optical flow features of the source video through the optical flow network model.
[0184] The fusion module 630 is configured to perform feature fusion on the foreground features and the optical flow features to generate a target video frame.
[0185] The generation module 640 is configured to generate a target video based on the target video frame.
[0186] In a possible embodiment, each video frame in the video frame sequence includes a timing identifier; the video frame sequence includes a first sequence and a second sequence, where the first sequence includes the target image and the generated target video frames; the second sequence includes video frames of the source video, and the second input module 620 is specifically configured to:
[0187] Obtain a third video frame corresponding to a timing identifier adjacent to the target timing identifier from the first sequence according to the target timing identifier of the target video frame to be generated;
[0188] Obtain a fourth video frame corresponding to a timing identifier adjacent to the target timing identifier from the second sequence according to the target timing identifier;
[0189] Extract foreground features of the third video frame through the image generation model, and extract optical flow features of the fourth video frame through the optical flow network model.
[0190] In summary, in the embodiment of the present invention, by inputting a video frame sequence including video frames of a source video and a target image into a video generation model, foreground features of the video frame sequence are extracted through an image generation model. Here, the image generation model can extract foreground features with authenticity and stability, and optical flow features of the source video are extracted through an optical flow network model. Here, the optical flow network model can extract stable and continuous optical flow features; feature fusion is performed on the foreground features and the optical flow features to generate a real and stable target video frame; thus, a real and stable target video can be generated based on the target video frame, improving the generation effect of the target video.
[0191] The embodiment of the present invention further provides an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete communication with each other through the communication bus 704.
[0192] A memory 703 for storing a computer program;
[0193] A processor 701, when executing the program stored on the memory 703, implements the following steps:
[0194] Obtain a plurality of sample videos; construct a generative adversarial network, the generative adversarial network including a generative model and a discriminative model; input the sample videos into the generative model to obtain predicted video frames; input the predicted video frames and the sample videos into the discriminative model to obtain a discrimination result; the discriminative model is used to determine whether the predicted video frames match the sample videos; train the generative adversarial network based on the discrimination results of each sample video until a training stop condition is satisfied to obtain the video generation model. Or,
[0195] Obtain a video frame sequence, the video frame sequence including: video frames of a source video and a target image;
[0196] Input the video frame sequence into a video generation model, the video generation model including: an image generation model and an optical flow network model; extract foreground features of the video frame sequence through the image generation model, and extract optical flow features of the source video through the optical flow network model;
[0197] Perform feature fusion on the foreground features and the optical flow features to generate a target video frame;
[0198] Generate a target video based on the target video frame.
[0199] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0200] The communication interface is used for communication between the above terminal and other devices.
[0201] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0202] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0203] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is caused to execute any of the methods described in the above embodiments.
[0204] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute any of the methods described in the above embodiments.
[0205] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0206] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0207] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0208] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A training method for a video generation model, characterized in that, The method includes: Obtaining a plurality of sample videos; the sample videos include a target object, a first video frame, and a plurality of second video frames adjacent to the first video frame; Constructing a generative adversarial network, the generative adversarial network including a generative model and a discriminative model; the generative model includes: an image generation model and an optical flow network model; the optical flow network model is used to model the optical flow characteristics of the background; the discriminative model includes: an image discriminative model and a video discriminative model; Inputting the sample videos into the generative model to obtain predicted video frames, including: inputting the plurality of second video frames into the generative model, extracting foreground training features of the plurality of second video frames through the image generation model, the foreground training features being generated based on the action features of the sample video and the shape features of the target object; and, extracting optical flow training features of the plurality of second video frames through the optical flow network model, the optical flow training features being obtained by inputting the sample pose information extracted from the sample video into the optical flow network model; fusing the foreground training features and the optical flow training features to obtain the predicted video frames; Inputting the predicted video frames and the sample videos into the discriminative model to obtain a discrimination result; the discriminative model is used to discriminate whether the predicted video frames match the sample videos, including discriminating the authenticity of a single-frame predicted video frame and the coherence of consecutive predicted video frames; Training the generative adversarial network based on the discrimination results of each sample video until a training stop condition is satisfied to obtain a video generation model.
2. The method according to claim 1, wherein the inputting the predicted video frames and the sample videos into the discriminative model to obtain a discrimination result, training the generative adversarial network based on the discrimination results of each training sample until a training stop condition is satisfied to obtain the video generation model includes: Inputting the predicted video frames and the first video frame into the image discriminative model to obtain a first loss value; Inputting the predicted video frames and the second video frames into the video discriminative model to obtain a second loss value; Training the generative adversarial network according to the first loss value and the second loss value until a training stop condition is satisfied to obtain the video generation model.
3. The method according to claim 2, before the inputting the predicted video frames and the sample videos into the discriminative model to obtain a discrimination result, the method further includes: Calculating a third loss value of the image generation model according to the predicted video frames and the first video frame; Calculating a fourth loss value of the optical flow network model according to the optical flow training features and the optical flow ground truth, the optical flow ground truth being extracted from the sample video through a preset optical flow extraction algorithm; Determining the loss value of the generative model according to the third loss value and the fourth loss value; The training the generative adversarial network based on the discrimination results of each training sample until a training stop condition is satisfied to obtain the video generation model includes: Determining the loss value of the discriminative model according to the first loss value and the second loss value; Based on the loss value of the generation model and the loss value of the discriminative model, perform backpropagation training on the generative adversarial network until the generative adversarial network meets a preset convergence condition, and obtain the trained video generation model.
4. The method according to claim 2 or 3, training the generative adversarial network according to the first loss value and the second loss value until a training stop condition is met to obtain the video generation model, including: Input the predicted video frame and the first video frame into a feature extraction network, the feature extraction network includes multiple scale layers, and each scale layer outputs a sub-loss value of the predicted video frame and the first video frame respectively; Determine a multi-scale loss value according to multiple sub-loss values; Train the generative adversarial network according to the multi-scale loss value, the first loss value and the second loss value until a training stop condition is met to obtain the video generation model.
5. A video generation method, characterized in that, The method includes: Obtain a video frame sequence, the video frame sequence includes: video frames of a source video and a target image; Input the video frame sequence into a video generation model, the video generation model includes: an image generation model and an optical flow network model; the optical flow network model is used to model the optical flow characteristics of the background; extract the foreground characteristics of the video frame sequence through the image generation model, and extract the optical flow characteristics of the source video through the optical flow network model; the foreground characteristics are generated based on the action characteristics of the source video and the appearance characteristics of the target image; the optical flow characteristics are obtained by inputting the pose information extracted from the video frames of the source video into the optical flow network model; Perform feature fusion on the foreground characteristics and the optical flow characteristics to generate a target video frame; Generate a target video based on the target video frame.
6. The method according to claim 5, each video frame in the video frame sequence includes a timing identifier; the video frame sequence includes a first sequence and a second sequence, the first sequence includes the target image and the generated target video frames; the second sequence includes the video frames of the source video, and the extracting the foreground characteristics of the video frame sequence through the image generation model and the extracting the optical flow characteristics of the source video through the optical flow network model includes: According to the target timing identifier of the target video frame to be generated, obtain a third video frame corresponding to the timing identifier adjacent to the target timing identifier from the first sequence; According to the target timing identifier, obtain a fourth video frame corresponding to the timing identifier adjacent to the target timing identifier from the second sequence; Extract the foreground characteristics of the third video frame through the image generation model, and extract the optical flow characteristics of the fourth video frame through the optical flow network model.
7. A training device for a video generation model, characterized in that The device includes: A first acquisition module, configured to acquire multiple sample videos; the sample videos include a target object, a first video frame, and multiple second video frames adjacent to the first video frame; A building module for building a generative adversarial network, the generative adversarial network including a generative model and a discriminative model; the generative model includes: an image generation model and an optical flow network model; the optical flow network model is used for modeling the optical flow features of the background; the discriminative model includes: an image discriminative model and a video discriminative model; A first input module for inputting the sample video into the generative model to obtain predicted video frames; specifically, for inputting the multiple second video frames into the generative model, extracting the foreground training features of the multiple second video frames through the image generation model, the foreground training features being generated based on the action features of the sample video and the shape features of the target object; and extracting the optical flow training features of the multiple second video frames through the optical flow network model, the optical flow training features being obtained by inputting the sample pose information extracted from the sample video into the optical flow network model; fusing the foreground training features and the optical flow training features to obtain the predicted video frames; The first input module is further configured to input the predicted video frames and the sample video into the discriminative model to obtain a discriminative result; the discriminative model is used for discriminating whether the predicted video frames match the sample video, including discriminating the authenticity of a single-frame predicted video frame and the coherence of consecutive predicted video frames; A training module for training the generative adversarial network based on the discriminative results of each sample video until a training stop condition is met, to obtain a video generation model.
8. A video generation device, characterized in that, The apparatus includes: A second acquisition module for acquiring a video frame sequence, the video frame sequence including: video frames of the source video and a target image; A second input module for inputting the video frame sequence into a video generation model, the video generation model including: an image generation model and an optical flow network model; the optical flow network model is used for modeling the optical flow features of the background; extracting the foreground features of the video frame sequence through the image generation model, and extracting the optical flow features of the source video through the optical flow network model; the foreground features are generated based on the action features of the source video and the shape features of the target image; the optical flow features are obtained by inputting the pose information extracted from the video frames of the source video into the optical flow network model; A fusion module for performing feature fusion on the foreground features and the optical flow features to generate target video frames; A generation module for generating a target video based on the target video frames.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
11. A computer program product, including a computer program, the computer program realizing the method according to any one of claims 1-6 when executed by a processor.
Citation Information
Patent Citations
Video generation method for action migration and neural network training method and device
CN110210386A
Camouflage image generation method based on multi-scale generative adversarial network
CN112288622A
Railway crossing abnormal event detection method and system based on generative adversarial network
CN112861762A