Video generation method and device, and electronic equipment

By extracting reference image pictures, face images and action features, and combining diffusion models to generate portrait animation videos, the problem of insufficient information integrity, similarity and fidelity of video generation in the prior art is solved, and higher information integrity, similarity and fidelity are achieved.

CN119996785APending Publication Date: 2025-05-13WONDERSHARE TECH (HUNAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411963559.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing neural network-based action sequence-driven character animation generation method is insufficient when generating videos, and it is difficult to effectively combine face features and reference features.

Method used

By extracting reference features and face features from the reference image map, and combining action features, portrait animation video is generated from random noise, and the diffusion model is used to iterate the output target latent vector based on input conditions and decode iteratively.

Benefits of technology

Improves the information integrity of the generated video, making it more similar and fidelity with the input face image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996785A_ABST
    Figure CN119996785A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method. The video generation method comprises the following steps: extracting reference features from a reference image picture; extracting facial features from the face image; extracting action features from the action guide sequence image; and according to the reference feature, the face feature and the action feature, sampling from random noise to generate a portrait animation video. According to the video generation method provided by the embodiment of the invention, through decoupling of the face features and the reference features, the generated video has higher information integrity, and has higher similarity and fidelity with the input face image. The embodiment of the invention further provides a video generation device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and more specifically, to a video generation method, device, and electronic device. Background Art

[0002] Motion-controlled character animation inputs action sequences (such as human posture maps, depth maps, etc.) and source character images. The model generates realistic and vivid corresponding character animation videos that are consistent with the input actions. The source of the action sequence can be extracted from the character action video or generated by other models. Driving the animation of human, animal, cartoon and other human-like characters has attracted a lot of research in the industry and has many potential applications in online retail, entertainment videos, artistic creation and virtual characters.

[0003] From the perspective of the backbone generation model, the existing neural network-based action sequence driven character animation generation methods are mainly divided into two categories: GAN network-based and diffusion model-based methods. Although both have made great progress, there are still some limitations. Summary of the invention

[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a video generation method, device, and electronic device, which decouple facial features from reference features so that the generated video has higher information integrity and higher similarity and fidelity with the input facial image.

[0005] In a first aspect, an embodiment of the present application provides a video generation method, comprising the following steps:

[0006] Extracting reference features from the reference image;

[0007] Extract facial features from face images;

[0008] Extracting motion features from the motion-guided sequence images; and

[0009] A portrait animation video is generated by sampling random noise according to the reference features, the facial features and the action features.

[0010] Furthermore, extracting reference features from the reference image includes:

[0011] Selecting the reference image;

[0012] painting the facial area of ​​the reference image; and

[0013] Extracting reference features of the reference image.

[0014] Furthermore, extracting facial features from a facial image includes:

[0015] Selecting the face image;

[0016] Apply to areas other than the facial area; and

[0017] Facial features are extracted from the facial region.

[0018] Furthermore, the step of sampling random noise to generate a portrait animation video according to the reference feature, the facial feature and the action feature includes:

[0019] Inputting the reference feature, the facial feature, and the action feature into a diffusion model, wherein the diffusion model iteratively samples and outputs the target latent vector from random noise based on input conditions; and

[0020] The target latent vector is decoded to generate the portrait animation video.

[0021] Further, before the reference feature, the facial feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input condition, the method further comprises:

[0022] The diffusion model is trained.

[0023] Furthermore, the training of the diffusion model includes:

[0024] Obtain a training sample set;

[0025] Selecting a frame of image from the video of the training sample set as a reference image;

[0026] Extracting a face image from any one of the remaining frames of the video of the training sample set;

[0027] Selecting a set of sequence frames from the video of the training sample set as a target sequence;

[0028] inputting the smeared reference image, the face image and the action sequence image into the diffusion model; and

[0029] The diffusion model is trained so that the diffusion model can reconstruct the target sequence.

[0030] Furthermore, after training the diffusion model so that the diffusion model can reconstruct the target sequence, the method further includes:

[0031] The iterations are repeated until the diffusion model converges.

[0032] In a second aspect, the embodiment of the present application further provides a video generating device, including:

[0033] A reference feature extraction module, used for extracting reference features from a reference image;

[0034] A facial feature extraction module, used to extract facial features from a face image;

[0035] An action feature extraction module, used to extract action features from action-guided sequence images; and

[0036] The video generation module is used to generate a portrait animation video by sampling from random noise according to the reference features, the facial features and the action features.

[0037] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video generation method according to the first aspect when executing the program.

[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is used to implement the video generation method according to the first aspect above.

[0039] The embodiments of the present application bring the following beneficial effects:

[0040] In the video generation method provided in the embodiment of the present application, a portrait animation video is generated by combining the appearance of the portrait and the background in the reference image, the facial features in the face image, and the motion posture of the action sequence, so that the reference image in the generated video has the facial features of the input face, thereby making the generated video have higher information integrity and higher similarity and fidelity with the input face image. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0042] Figure 1 A schematic diagram of a flow chart of a video generation method provided in an embodiment of the present application;

[0043] Figure 2 A schematic diagram of the structure of the overall algorithm used in the video generation method provided in the embodiment of the present application;

[0044] Figure 3 A schematic diagram of the structure of a reference feature encoder used in the video generation method provided in an embodiment of the present application;

[0045] Figure 4 A schematic diagram of the structure of the motion feature encoder used in the video generation method provided in an embodiment of the present application;

[0046] Figure 5 A structural block diagram of a video generation device provided in an embodiment of the present application;

[0047] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0048] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments described in the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0050] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0051] Figure 1 FIG. 1 is a flow chart of a video generation method according to an embodiment of the present application. Figure 1 As shown, the video generation method of the embodiment of the present application is used for gesture recognition training, including the following steps:

[0052] S101: extracting reference features from the reference image;

[0053] The reference image image described here provides the generated portrait animation video with an image appearance other than the face, so as to make the image in the generated portrait animation video concrete.

[0054] S102: extracting facial features from the face image;

[0055] The reference image provides image appearance for the generated portrait animation video, while the face image provides facial features for the portrait animation video.

[0056] S103: extracting action features from the action-guided sequence images; and

[0057] The action guidance sequence images involved in the embodiments of the present application may have various sources, for example, they may be extracted from character action videos (driving videos), or they may be generated using other character action generation models.

[0058] S104: Generate a portrait animation video by sampling from random noise according to the reference features, the facial features and the action features.

[0059] The video generation method provided in the embodiment of the present application generates a portrait animation video by combining the appearance of the portrait (excluding the face) and the background in the reference image, the facial features in the face image, and the motion posture of the action sequence, so that the reference image in the generated video has the facial features of the input face and moves according to the action sequence posture.

[0060] It should be noted that, refer to Figure 2 The video generation method provided in the embodiment of the present application generates a portrait animation video based on a diffusion model. The overall algorithm structure adopted by the video generation method provided in the embodiment of the present application is composed of a reference feature and face feature extraction module, an action feature extraction module and a diffusion model.

[0061] Therefore, in the video generation method provided in the embodiment of the present application, a portrait animation video is generated by combining the appearance of the portrait and the background in the reference image, the facial features in the face image, and the motion posture of the action sequence, so that the reference image in the generated video has the facial features of the input face, thereby making the generated video have higher information integrity and higher similarity and fidelity with the input face image.

[0062] Further, in some embodiments of the present application, extracting reference features from the reference image includes:

[0063] Selecting the reference image;

[0064] painting the facial area of ​​the reference image; and

[0065] Extracting reference features of the reference image.

[0066] Specifically, the reference image and the face image each contain the face and other background information, such as hairstyle, accessories, some clothing, etc. To achieve the goal of face swapping, the model will logically obtain background and image information (such as background, body, hairstyle, etc.) from the image reference image and obtain facial information from the face image. Therefore, it is necessary to erase the face part in the image reference image and erase the information other than the face in the face image, so as to ensure that the facial information comes from the face image and other reference information comes from the image reference image.

[0067] Furthermore, extracting facial features from a facial image includes:

[0068] Selecting the face image;

[0069] Apply to areas other than the facial area; and

[0070] Facial features are extracted from the facial region.

[0071] Specifically, the reference image and the face image each contain the face and other background information, such as hairstyle, accessories, some clothing, etc. To achieve the goal of face swapping, the model will logically obtain background and image information (such as background, body, hairstyle, etc.) from the image reference image and obtain facial information from the face image. Therefore, it is necessary to erase the face part in the image reference image and erase the information other than the face in the face image, so as to ensure that the facial information comes from the face image and other reference information comes from the image reference image.

[0072] The diffusion model denoising process is performed in the VAE latent space. In order to ensure the homology of the reference features, facial features and features in the diffusion model, the feature extraction and implantation of the two are also performed in the latent space. In addition, since the VAE latent space can ensure the integrity of image information, the reference features and facial features extracted from it will also have richer detailed information, which can more accurately reconstruct the reference image and input face.

[0073] After the reference image is compressed into the latent space by VAE, it is then passed through the reference feature encoder to extract multi-level reference features. Similarly, after the face image is compressed into the latent space by the VAE encoder, it is then passed through the face feature encoder to extract multi-level face features. The above feature levels are the same as the diffusion model levels. After the reference features at each level and the corresponding face features are spliced, they will be implanted into the corresponding diffusion model level. The feature implantation operation is performed through the cross-attention layer.

[0074] The concatenation method of the i-th level reference feature and the i-th level face feature is shown in formula (1):

[0075]

[0076] According to the splicing method in formula (1), in general, the reference feature before splicing The data dimension is (B, N1, D), face features The data dimension is (B, N2, D), which is concatenated along the second dimension. The data dimension is (B, N1+N2, D).

[0077] The cross attention layer in the diffusion model is used to embed conditional information. In the present invention, reference features and face features are also embedded through the cross attention layer. The calculation logic of the i-th level cross attention layer is expressed as follows:

[0078]

[0079] In formula (2), cross attention is a known algorithm, Q i , K i , V i are three input vectors, Q i is the hidden vector of level i in the diffusion model, K i and V i All come from the i-th level splicing reference features L represents the corresponding linear layer in the cross-attention layer.

[0080] Reference Figure 3 The reference feature encoder and the face feature encoder are two modules with the same structure. Their specific structure mainly consists of convolutional layers and fully connected layers. The downsampling layer realizes the alignment of the feature vector with the spatial resolution of the diffusion model step by step. The number of fully connected layers is the same as the number of diffusion models, and the output dimension is the same as the feature dimension of the corresponding level of the diffusion model. The convolution kernel of the convolution layer is 3x3, the step size is 1, and the number of channels is 64.

[0081] Furthermore, the step of sampling random noise to generate a portrait animation video according to the reference feature, the facial feature and the action feature includes:

[0082] Inputting the reference feature, the facial feature, and the action feature into a diffusion model, wherein the diffusion model iteratively samples and outputs the target latent vector from random noise based on input conditions; and

[0083] The target latent vector is decoded to generate the portrait animation video.

[0084] Specifically, the diffusion model adopted in the embodiment of the present application uses AnimateDiff, which contains a VAE (Variational Autoencoding) module and a denoising module (Denosing UNet). The Denosing UNet in the diffusion model is a multi-level resolution structure. The resolution is first halved step by step, and then doubled step by step. The final output vector is the same size as the input vector, and its resolution level is the diffusion model level adopted in the embodiment of the present application. In order to improve the training efficiency of the diffusion model, the diffusion model uses a VAE encoder to compress the image into a latent space for training and sampling. VAE is divided into two parts, an encoder and a decoder. The VAE encoder compresses the image into a latent vector in the latent space (for example, in the AnimateDiff model, the image is compressed to the size of (H / 4, W / 4, 4)), and the VAE decoder can restore the latent vector to an image almost losslessly.

[0085] Reference Figure 4 , the human body posture skeleton sequence graph is used to represent the action sequence, where different colors are used to represent the semantics of different parts of the human body. The posture skeleton graph is input into the action feature encoder, and multi-level action features are extracted and then implanted into the diffusion model. The number of feature levels is the same as that of the diffusion model.

[0086] like Figure 4 As shown in the figure, the motion feature encoder is mainly composed of convolution layers. The convolution kernels are 3x3 and 1x1. The number of channels of the 3x3 convolution layer is 64. The output layer is a 1x1 convolution layer. The number of output layers is the same as the number of levels of the diffusion model, and the number of channels of the output layer is the same as the number of feature channels of the corresponding level of the diffusion model. The multi-level motion features are accumulated with the features of the corresponding level of the diffusion model to complete the implantation of the motion features.

[0087] Further, before the reference feature, the facial feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input condition, the method further comprises:

[0088] The diffusion model is trained.

[0089] That is, the expansion model needs to be fitted and trained on the training sample set before it can work properly.

[0090] Furthermore, the training of the diffusion model includes:

[0091] Obtain a training sample set;

[0092] Selecting a frame of image from the video of the training sample set as a reference image;

[0093] Extracting a face image from any one of the remaining frames of the video of the training sample set;

[0094] Selecting a set of sequence frames from the video of the training sample set as a target sequence;

[0095] inputting the smeared reference image, the face image and the action sequence image into the diffusion model; and

[0096] The diffusion model is trained so that the diffusion model can reconstruct the target sequence.

[0097] Specifically, when training the diffusion model, it is necessary to first collect a large number of single-person frontal action videos (such as standing dancing or other frontal sports videos), and then annotate the videos. The annotation method is: use the portrait face segmentation model to process the video frame by frame, detect the portrait face mask in the video, use the face detection model to detect the face frame by frame, frame the face area, use the human key point detection model to detect the human key points frame by frame, and then specify the color lines to connect the key points to draw the posture skeleton sequence diagram.

[0098] The training process is as follows: select a video from the sample set, select a frame from the video as the reference image, and select a frame from the remaining frames to extract the face image based on the annotated face frame; then select a set of sequence frames from the remaining frames as the target sequence, and select the posture skeleton map corresponding to the target sequence as the action sequence map. According to the corresponding portrait face mask map, erase the face of the reference image (for example, fill it with pure white), and erase the non-face area of ​​the face map (for example, fill it with pure white); finally, input the smeared reference image, face map and action sequence map into the diffusion model; train the diffusion model so that it can reconstruct the target sequence.

[0099] Furthermore, after training the diffusion model so that the diffusion model can reconstruct the target sequence, the method further includes:

[0100] The iterations are repeated until the diffusion model converges.

[0101] That is, samples are selected using the above steps and trained repeatedly until the model converges. The training method of the diffusion model is, for example, a v-predict training method.

[0102] Normal reasoning can only be performed after the diffusion model training is completed, that is, selecting an image reference image and smearing the face, selecting a face image and smearing the non-facial area, selecting an action sequence image, inputting it into the diffusion model, and sampling from the noise to generate the corresponding portrait animation video.

[0103] Figure 5 2 is a structural block diagram of a video generation device 200 provided in an embodiment of the present application. Figure 5As shown, the video generating device 200 of the embodiment of the present application includes: a reference feature extraction module 210, a facial feature extraction module 220, an action feature extraction module 230 and a video generating module 240, wherein:

[0104] A reference feature extraction module 210, for extracting reference features from a reference image;

[0105] A facial feature extraction module 220, for extracting facial features from a face image;

[0106] An action feature extraction module 230, used to extract action features from action guide sequence images; and

[0107] The video generation module 240 is used to generate a portrait animation video by sampling from random noise according to the reference features, the facial features and the action features.

[0108] In the video generation device provided in the embodiment of the present application, a portrait animation video is generated by combining the appearance of the portrait and the background in the reference image, the facial features in the face image, and the motion posture of the action sequence, so that the reference image in the generated video has the facial features of the input face, thereby making the generated video have higher information integrity and higher similarity and fidelity with the input face image.

[0109] It should be noted that the specific implementation of the video generating device of the embodiment of the present application is similar to the specific implementation of the video generating method of the embodiment of the present application. Please refer to the description of the method part for details, and no further details will be given here.

[0110] Figure 6 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.

[0111] like Figure 6 As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage part 302 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0112] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.

[0113] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 309, and / or installed from a removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.

[0114] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0115] In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0116] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based electronic device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0117] The units or modules involved in the embodiments of the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor, and the processor is used to implement the video generation method when executing the program:

[0118] Extracting reference features from the reference image;

[0119] Extract facial features from face images;

[0120] Extracting motion features from the motion-guided sequence images; and

[0121] A portrait animation video is generated by sampling random noise according to the reference features, the facial features and the action features.

[0122] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the video generation method described in the present application:

[0123] Extracting reference features from the reference image;

[0124] Extract facial features from face images;

[0125] Extracting motion features from the motion-guided sequence images; and

[0126] A portrait animation video is generated by sampling random noise according to the reference features, the facial features and the action features.

[0127] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the video generation method described in the present application:

[0128] Extracting reference features from the reference image;

[0129] Extract facial features from face images;

[0130] Extracting motion features from the motion-guided sequence images; and

[0131] A portrait animation video is generated by sampling random noise according to the reference features, the facial features and the action features.

[0132] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural changes made by using the contents of the present application specification and drawings under the application concept of the present application, or directly / indirectly used in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A video generation method, characterized in that: The following steps are involved: Extracting reference features from the reference image; Extract facial features from face images; Extracting motion features from motion-guided sequence images; and A portrait animation video is generated by sampling random noise according to the reference features, the facial features and the action features.

2. The video generation method according to claim 1, characterized in that: The step of extracting reference features from the reference image includes: Selecting the reference image; painting the facial area of ​​the reference image; and Extracting reference features of the reference image.

3. The video generation method according to claim 1, characterized in that: The extracting facial features from the facial image comprises: Selecting the face image; Apply to areas other than the facial area; and Facial features are extracted from the facial region.

4. The video generation method according to claim 1, characterized in that: The step of sampling random noise to generate a portrait animation video according to the reference feature, the facial feature and the action feature includes: Inputting the reference feature, the facial feature, and the action feature into a diffusion model, wherein the diffusion model iteratively samples and outputs the target latent vector from random noise based on input conditions; and The target latent vector is decoded to generate the portrait animation video.

5. The video generation method according to claim 4, characterized in that: Before the reference feature, the facial feature and the action feature are input into a diffusion model, and the diffusion model iteratively samples and outputs the target latent vector from random noise based on the input condition, the method includes: The diffusion model is trained.

6. The video generation method according to claim 5, characterized in that: The step of training the diffusion model comprises: Obtain a training sample set; Selecting a frame of image from the video of the training sample set as a reference image; Extracting a face image from any one of the remaining frames of the video of the training sample set; Selecting a set of sequence frames from the video of the training sample set as a target sequence; inputting the smeared reference image, the face image and the action sequence image into the diffusion model; and The diffusion model is trained so that the diffusion model can reconstruct the target sequence.

7. The video generation method according to claim 6, characterized in that: After training the diffusion model so that the diffusion model can reconstruct the target sequence, the method further includes: The iterations are repeated until the diffusion model converges.

8. A video generating device, characterized in that: include: A reference feature extraction module, used for extracting reference features from a reference image; A facial feature extraction module, used to extract facial features from a face image; An action feature extraction module is used to extract action features from action-guided sequence images; and The video generation module is used to generate a portrait animation video by sampling from random noise according to the reference features, the facial features and the action features.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the video generating method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the video generation method according to any one of claims 1-7.

Citation Information

Cited By

  • Action migration method and device, electronic equipment, computer readable storage medium and program product

    CN121354212A