Video generation method and device, and electronic equipment
By introducing multi-scale source texture implant structures into the video generation method, extracting and fusing source textures and action features to generate high-quality animated videos, the problem of insufficient picture quality and stability in the prior art is solved.
Patent Information
- Application Number
- CN202411958038.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing neural network-based action sequence-driven character animation generation methods have shortcomings in image quality and stability, and it is difficult to effectively maintain the source texture of the output image.
By introducing a more accurate multi-scale source texture implantation structure, the multi-scale source texture reference features in the source reference image and the action features in the action guide sequence image are extracted, and the target latent vector with source texture information and action information is generated, and restored to an image to generate the target video.
The image quality and stability of the generated video are enhanced, and the ability to reconstruct the source role by the generator model and distinguish the front and back scenes of the reference image are improved.
Smart Images

Figure CN119996784A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and more specifically, to a video generation method, device, and electronic device. Background Art
[0002] Motion-controlled character animation inputs action sequences (such as human posture maps, depth maps, etc.) and source character images. The model generates realistic and vivid corresponding character animation videos that are consistent with the input actions. The source of the action sequence can be extracted from the character action video or generated by other models. Driving the animation of human, animal, cartoon and other human-like characters has attracted a lot of research in the industry and has many potential applications in online retail, entertainment videos, artistic creation and virtual characters.
[0003] From the perspective of the backbone generation model, the existing neural network-based action sequence driven character animation generation methods are mainly divided into two categories: GAN network-based and diffusion model-based methods. Although both have made great progress, there are still some limitations. Summary of the invention
[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a video generation method, device, and electronic device. By introducing a more accurate multi-scale source texture implantation structure, the source texture retention capability of the output image can be enhanced, thereby enhancing the image quality and stability of the generated video.
[0005] In a first aspect, an embodiment of the present application provides a video generation method, comprising the following steps:
[0006] extracting source texture reference features from a source reference image;
[0007] Extracting motion features from motion-guided sequence images;
[0008] Generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0009] The target latent vector is restored to an image to generate a target video.
[0010] Further, extracting source texture reference features from the source image includes:
[0011] Obtaining an intermediate latent vector according to the source reference image and the source posture skeleton graph; and
[0012] The intermediate latent vectors are concatenated, and multi-scale source texture reference features are extracted.
[0013] Furthermore, extracting action features from the action-guided sequence images includes:
[0014] generating a posture guidance feature, a head guidance feature and a hand guidance feature according to the posture skeleton map, the 3D head surface map and the 3D hand surface map of the action guidance sequence image; and
[0015] The posture guidance feature, the head guidance feature and the hand guidance feature are fused to extract the action feature.
[0016] Further, generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature includes:
[0017] The source texture reference feature and the motion feature are input into a diffusion model, and the diffusion model iteratively samples from random noise based on input conditions to output the target latent vector.
[0018] Further, before the source texture reference feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input condition, the method further comprises:
[0019] The diffusion model is trained.
[0020] Furthermore, the training of the diffusion model includes:
[0021] Obtain a training sample set;
[0022] Selecting a frame of image and its corresponding posture skeleton graph from the video of the training sample set as source reference input;
[0023] Selecting a frame of image from the remaining frames of the video of the training sample set as a target frame;
[0024] Selecting the action guidance map of the target frame and inputting it into the diffusion model, training the diffusion model so that the generated map is consistent with the image of the target frame; and
[0025] The above steps are iterated repeatedly until the diffusion model converges.
[0026] Furthermore, the obtaining of the training sample set includes:
[0027] Extract human body key points through human body key points and draw posture skeleton diagram;
[0028] Extracting 3D head information through the 3D head model and drawing a 3D head surface map; and
[0029] The 3D hand information is extracted through the 3D hand model and a 3D hand surface map is drawn.
[0030] In a second aspect, an embodiment of the present application further provides a video generating device, including:
[0031] A reference feature extraction module, used for extracting source texture reference features from a source reference image;
[0032] An action feature extraction module is used to extract action features from action-guided sequence images;
[0033] a latent vector generation module, configured to generate a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0034] The target video generation module is used to restore the target latent vector into an image to generate a target video.
[0035] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video generation method according to the first aspect when executing the program.
[0036] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is used to implement the video generation method according to the first aspect above.
[0037] The embodiments of the present application bring the following beneficial effects:
[0038] In the video generation method provided in the embodiment of the present application, firstly, multi-scale source texture reference features are extracted from the source reference image, and action features are extracted from the action guide sequence image, and then a target latent vector with source texture information and action information is generated according to the source texture reference features and the action features, and finally the target latent vector is restored to an image to generate a target video. The video generation method provided in the embodiment of the present application can enhance the source texture retention capability of the output image, improve the ability of the generated model to reconstruct the source character and distinguish the foreground and background of the reference image by introducing a more accurate multi-scale source texture implantation structure, thereby enhancing the image quality and stability of the generated video. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0040] Figure 1 A schematic diagram of a flow chart of a video generation method provided in an embodiment of the present application;
[0041] Figure 2 A schematic diagram of the structure of the overall algorithm used in the video generation method provided in the embodiment of the present application;
[0042] Figure 3 A schematic diagram of the structure of a reference feature encoder used in the video generation method provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of the structure of the motion feature encoder used in the video generation method provided in an embodiment of the present application;
[0044] Figure 5 A structural block diagram of a video generation device provided in an embodiment of the present application;
[0045] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0046] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments described in the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0048] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.
[0049] Figure 1 FIG. 1 is a flow chart of a video generation method according to an embodiment of the present application. Figure 1 As shown, the video generation method of the embodiment of the present application is used for gesture recognition training, including the following steps:
[0050] S101: extracting multi-scale source texture reference features from a source reference image;
[0051] The character images supported by the embodiment of the present application are not limited to real portrait images, but can also support images with human-like images such as animals, cartoons, and puppets.
[0052] S102: extracting action features from action-guided sequence images;
[0053] The action guidance sequence images involved in the embodiments of the present application may have various sources, for example, they may be extracted from character action videos (driving videos), or they may be generated using other character action generation models.
[0054] S103: generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0055] S104: Restoring the target latent vector into an image to generate a target video.
[0056] That is, refer to Figure 2 The main purpose of the embodiment of the present application is to drive the diffusion model to generate an animation video whose appearance is consistent with the source reference image and complies with the control action based on the source reference condition and the action condition. The diffusion model in the embodiment of the present application uses AnimateDiff, which already includes a VAE (Variational Auto-Encoding) module. Figure 2 As shown, first, the control conditions are generated step by step. The reference feature encoder extracts the source texture reference features, and the action feature encoder extracts the action features. Then, the source reference features and the action features are input as conditions into the diffusion model. The diffusion model generates a target latent vector with source texture information and action information based on the source texture reference features and the action features. Finally, the VAE decoder restores the latent vector to an RGB image to generate the target video.
[0057] Therefore, in the video generation method provided in the embodiment of the present application, firstly, multi-scale source texture reference features are extracted from the source reference image, and action features are extracted from the action guide sequence image, and then a target latent vector with source texture information and action information is generated according to the source texture reference features and the action features, and finally the target latent vector is restored to an image to generate a target video. The video generation method provided in the embodiment of the present application can enhance the source texture retention capability of the output image, improve the ability of the generated model to reconstruct the source character and distinguish the foreground and background of the reference image by introducing a more accurate multi-scale source texture implantation structure, thereby enhancing the image quality and stability of the generated video.
[0058] Furthermore, in some embodiments of the present application, extracting source texture reference features from the source image includes:
[0059] Obtaining an intermediate latent vector according to the source reference image and the source posture skeleton graph; and
[0060] The intermediate latent vectors are concatenated, and multi-scale source texture reference features are extracted.
[0061] Specifically, the embodiments of the present application support input of character images that are not only real portrait images, but also images with human-like images such as animals, cartoons, and puppets. Since different images have different body appearance features, and are limited by the limitations of training samples, it is impossible for training samples to cover all morphological and appearance combinations. Therefore, in order to enhance the robustness of the model, additional features are needed to indicate the semantics of each image part of the reference image. Therefore, a source posture skeleton graph that is consistent with the posture of the source reference image is introduced as the source reference input together with the source character image.
[0062] In order to ensure that the source reference features are accurately restored, the embodiment of the present application does not adopt the features compressed by the visual model commonly used in the industry (such as CLIP), but instead adopts the features compressed by the VAE encoder. Therefore, it can ensure the homology of the reference features and the features in the diffusion model, and the feature distribution is more similar.
[0063] The source reference feature extraction process is as follows: the source reference image and the source posture skeleton map are respectively input into the VAE encoder to obtain the corresponding intermediate latent vector, and then the two intermediate latent vectors are spliced according to the vector dimension, and then input into the reference feature encoder to extract multi-scale source texture features (the scale level is the same as the resolution level in the diffusion model), and finally the source texture features are respectively input into the corresponding layers of the diffusion model for feature implantation. Each level of resolution in the diffusion model has a cross-attention layer for fusing external conditions. The source texture features at all levels are implanted through the cross-attention layer of the corresponding resolution, and the source texture features at all levels have the same dimension as the corresponding cross-attention layer input features.
[0064] Reference Figure 3 The feature encoder consists of a convolutional layer and a fully connected layer. The convolution kernel of the convolution layer is 3x3, the step size is 1, and the number of channels is 64. The downsampling layer is used to reduce the feature map to the target resolution step by step. The number of fully connected layers is the same as the resolution level of the diffusion model, and the number of output feature channels is the same as the corresponding feature dimension of the diffusion model. Of course, the specific number of network layers and convolution kernel parameters of the reference feature encoder in the source texture implantation process vary.
[0065] The calculation logic of the cross attention layer in the diffusion model is expressed by formula (1):
[0066]
[0067] Among them, cross attention is the cross attention algorithm, Q i , K i , V i are three input vectors, Qi is the hidden vector from the i-th level in the diffusion model K i and V i All come from the i-th level reference feature L represents the corresponding linear layer in the cross-attention layer.
[0068] Furthermore, in some embodiments of the present application, extracting action features from the action guidance sequence images includes:
[0069] generating a posture guidance feature, a head guidance feature and a hand guidance feature according to the posture skeleton map, the 3D head surface map and the 3D hand surface map of the action guidance sequence image; and
[0070] The posture guidance feature, the head guidance feature and the hand guidance feature are fused to extract the action feature.
[0071] Specifically, in the embodiment of the present application, the action guidance sequence image can have multiple sources, such as being extracted from a character action video (driving video), or being generated using other character action generation models. The action guidance sequence image consists of three parts: a posture skeleton image, a 3D head surface image, and a 3D hand surface image. The action information contained in the three parts is the same, but the expression forms are different.
[0072] The three action guide graphs play different roles. The posture skeleton graph controls the overall movement posture. Since the face and hand control information in the skeleton graph is sparse or inaccurate, it is easy to cause distortion and defects in the face and hands in the generated image. By introducing precise and strict 3D face and hand information, it helps to generate detailed and realistic images.
[0073] The three action guidance images generate corresponding posture guidance features, head guidance features, and hand guidance features through the posture guide, head guide, and hand guide respectively. The three features are input into the guidance feature fusion device to finally obtain multi-level guidance features. The level of guidance features is the same as the resolution level of the diffusion model, and the feature dimension of each level is the same as the corresponding diffusion model feature; the guidance features are directly accumulated to the corresponding features of the diffusion model to realize the implantation of action guidance features.
[0074] Reference Figure 4 , showing the encoder structure of the action feature. The posture guide, head guide, and hand guide have the same structure, all consisting of convolution layers, with a convolution kernel of 3x3 and a step size of 1. The action feature fusion consists of a downsampling layer and a convolution layer, with two types of convolution kernels: 3x3 and 1x1, and a step size of 1. Of course, the three guides and feature fusion in the action feature implantation process vary in the specific number of network layers and convolution kernel parameters.
[0075] The embodiment of the present application generates high-quality animation videos using a diffusion model by specifying source character images and action sequences. The rich multi-level texture features ensure the accuracy of the reconstruction of the source texture. At the same time, clear semantic information will also enhance the generalization ability of the model. Due to the use of semantic indication information, although the model is only trained on real-life samples, it can also support non-real-life character images after training. It can guide the face and hands, which are areas sensitive to the human eye, to present more delicate details through more precise 3D head and hand movements, and the overall image quality is more realistic.
[0076] Further, in some embodiments of the present application, generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature includes:
[0077] The source texture reference feature and the motion feature are input into a diffusion model, and the diffusion model iteratively samples from random noise based on input conditions to output the target latent vector.
[0078] As described above, the main purpose of the embodiment of the present application is to drive the diffusion model to generate an animation video whose appearance is consistent with the source reference image and complies with the control action based on the source reference condition and the action condition. The diffusion model in the embodiment of the present application uses AnimateDiff, which already includes a VAE (Variational Autoencoding) module. Figure 2 As shown, first, the control conditions are generated step by step, the reference feature encoder extracts the source texture reference features, the action feature encoder extracts the action features, and then the source reference features and the action features are input as conditions to the diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input conditions. The video generation method provided in the embodiment of the present application can enhance the source texture retention ability of the output image by introducing a more accurate multi-scale source texture implantation structure, improve the ability of the generation model to reconstruct the source character and distinguish the foreground and background of the reference image, thereby enhancing the image quality and stability of the generated video.
[0079] Further, in some embodiments of the present application, before the source texture reference feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input condition, the method includes:
[0080] The diffusion model is trained.
[0081] That is, the expansion model needs to be fitted and trained on the training sample set before it can work properly.
[0082] Further, in some embodiments of the present application, the training of the diffusion model includes:
[0083] Obtain a training sample set;
[0084] Selecting a frame of image and its corresponding posture skeleton graph from the video of the training sample set as source reference input;
[0085] Selecting a frame of image from the remaining frames of the video of the training sample set as a target frame;
[0086] Selecting the action guidance map of the target frame and inputting it into the diffusion model, training the diffusion model so that the generated map is consistent with the image of the target frame; and
[0087] The above steps are iterated repeatedly until the diffusion model converges.
[0088] That is, when training the diffusion model, we first need to select a processed video from the training sample set, select a frame image and its corresponding posture skeleton map from the video as the source reference input, select a frame from the remaining frames of the video as the target frame, select the action guide map of the target frame and input it into the diffusion model, and train the diffusion model so that the generated map is consistent with the target frame. Figure 1 The above steps are repeated until the model converges. Training is performed in the latent space, usually using the v-predict training method.
[0089] Further, in some embodiments of the present application, obtaining a training sample set includes:
[0090] Extract human body key points through human body key points and draw posture skeleton diagram;
[0091] Extracting 3D head information through the 3D head model and drawing a 3D head surface map; and
[0092] The 3D hand information is extracted through the 3D hand model and a 3D hand surface map is drawn.
[0093] Specifically, obtaining a training sample set means collecting real human action videos. Human key points can be extracted frame by frame through a human key point extraction model such as DWPose and a color posture skeleton map can be drawn. 3D head information can be extracted through a 3D head model (such as FLAME) and a head surface map can be drawn. 3D hand information can be extracted through a 3D hand model (such as HaMeR) and a 3D hand surface map can be drawn.
[0094] Figure 5 2 is a structural block diagram of a video generation device 200 provided in an embodiment of the present application. Figure 5 As shown, the video generation device 200 of the embodiment of the present application includes: a reference feature extraction module 210, an action feature extraction module 220, a latent vector generation module 230 and a target video generation module 240, wherein:
[0095] A reference feature extraction module 210 is used to extract source texture reference features from a source reference image;
[0096] An action feature extraction module 220 is used to extract action features from action guide sequence images;
[0097] A latent vector generating module 230, configured to generate a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0098] The target video generation module 240 is used to restore the target latent vector into an image to generate a target video.
[0099] In the video generation device provided in the embodiment of the present application, firstly, multi-scale source texture reference features are extracted from the source reference image, and action features are extracted from the action guide sequence image, and then a target latent vector with source texture information and action information is generated according to the source texture reference features and the action features, and finally the target latent vector is restored to an image to generate a target video. The video generation method provided in the embodiment of the present application can enhance the source texture retention capability of the output image, improve the ability of the generated model to reconstruct the source character and distinguish the foreground and background of the reference image by introducing a more accurate multi-scale source texture implantation structure, thereby enhancing the image quality and stability of the generated video.
[0100] It should be noted that the specific implementation of the video generating device of the embodiment of the present application is similar to the specific implementation of the video generating method of the embodiment of the present application. Please refer to the description of the method part for details, and no further details will be given here.
[0101] Figure 6 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.
[0102] like Figure 6 As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage part 302 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0103] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.
[0104] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 309, and / or installed from a removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.
[0105] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0106] In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0107] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based electronic device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0108] The units or modules involved in the embodiments of the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor, and the processor is used to implement the video generation method when executing the program:
[0109] extracting source texture reference features from a source reference image;
[0110] Extracting motion features from motion-guided sequence images;
[0111] Generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0112] The target latent vector is restored to an image to generate a target video.
[0113] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the video generation method described in the present application:
[0114] extracting source texture reference features from a source reference image;
[0115] Extracting motion features from motion-guided sequence images;
[0116] Generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0117] The target latent vector is restored to an image to generate a target video.
[0118] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the video generation method described in the present application:
[0119] extracting source texture reference features from a source reference image;
[0120] Extracting motion features from motion-guided sequence images;
[0121] Generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and
[0122] The target latent vector is restored to an image to generate a target video.
[0123] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural changes made by using the contents of the present application specification and drawings under the application concept of the present application, or directly / indirectly used in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A video generation method, characterized in that: The following steps are involved: Extracting multi-scale source texture reference features from the source reference image; Extracting motion features from motion-guided sequence images; Generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and The target latent vector is restored to an image to generate a target video.
2. The video generation method according to claim 1, characterized in that: The step of extracting multi-scale source texture reference features from the source image comprises: Obtaining an intermediate latent vector according to the source reference image and the source posture skeleton graph; and The intermediate latent vectors are concatenated, and multi-scale source texture reference features are extracted.
3. The video generation method according to claim 1, characterized in that: The step of extracting action features from the action-guided sequence images comprises: generating a posture guidance feature, a head guidance feature and a hand guidance feature according to the posture skeleton map, the 3D head surface map and the 3D hand surface map of the action guidance sequence image; and The posture guidance feature, the head guidance feature and the hand guidance feature are fused to extract the action feature.
4. The video generation method according to claim 1, characterized in that: The step of generating a target latent vector with source texture information and action information according to the source texture reference feature and the action feature comprises: The source texture reference feature and the motion feature are input into a diffusion model, and the diffusion model iteratively samples from random noise based on input conditions to output the target latent vector.
5. The video generation method according to claim 4, characterized in that: Before the source texture reference feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on an input condition, the method includes: The diffusion model is trained.
6. The video generation method according to claim 5, characterized in that: The step of training the diffusion model comprises: Obtain a training sample set; Selecting a frame of image and its corresponding posture skeleton graph from the video of the training sample set as source reference input; Selecting a frame of image from the remaining frames of the video of the training sample set as a target frame; Selecting the action guidance map of the target frame and inputting it into the diffusion model, training the diffusion model so that the generated map is consistent with the image of the target frame; and The above steps are iterated repeatedly until the diffusion model converges.
7. The video generation method according to claim 6, characterized in that: The step of obtaining a training sample set includes: Extract human body key points through human body key points and draw posture skeleton diagram; Extracting 3D head information through the 3D head model and drawing a 3D head surface map; and The 3D hand information is extracted through the 3D hand model and a 3D hand surface map is drawn.
8. A video generating device, characterized in that: include: A reference feature extraction module, used for extracting source texture reference features from a source reference image; An action feature extraction module is used to extract action features from action-guided sequence images; A latent vector generation module, used to generate a target latent vector with source texture information and action information according to the source texture reference feature and the action feature; and The target video generation module is used to restore the target latent vector into an image to generate a target video.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video generating method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the video generation method according to any one of claims 1-7.
Citation Information
Patent Citations
Character action video generation method and system based on human skeleton sequence information and storage medium
CN112419455A
Human body video generation method based on time sequence consistent hidden space guide diffusion model
CN117994708A
Video generation method and device, electronic equipment and readable storage medium
CN118015509A
Video generation method and device, electronic equipment and computer storage medium
CN118741260A
Video generation method and device, equipment and medium
CN118799460A