Time-consistent human image animation production method
By combining appearance encoder, posture control network and video diffusion model in the computing system, the problems of time consistency and identity fidelity in human image animation production are solved, and high-quality and stable animation generation is achieved.
Patent Information
- Application Number
- CN202411684571.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-22
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to maintain time consistency and fidelity of reference identities in human image animation production, resulting in animation instability and loss of details.
A computing system is employed that includes a processing circuitry configured to implement a appearance encoder, a posture control network, and a video diffusion model. The appearance encoder encodes the reference image as appearance embedding, the posture control network extracts motion conditions, and the video diffusion model uses the temporal attention mechanism to generate a denoising animation sequence.
It achieves long-term consistency, robust appearance coding and high-fidelity animation effects, significantly improving video fidelity and single-frame quality.
Smart Images

Figure CN120047582A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 602,509, filed on November 24, 2023, which is hereby incorporated by reference in its entirety for all purposes. Background Art
[0003] The present disclosure relates to human image animation tasks, which aim to generate videos of a specific reference identity according to a specific motion sequence. Existing methods typically use frame warping techniques to animate reference images to achieve the target motion. However, due to the lack of temporal modeling and poor preservation of the reference identity, these methods face challenges in maintaining the temporal consistency of the entire animation. Summary of the Invention
[0004] To address the above problems, a computing system is provided. In one example, the computing system includes processing circuitry configured to implement an appearance encoder configured to encode a reference image into an appearance embedding. The processing circuitry is further configured to implement a pose control network configured to receive a target pose sequence as input and, in response, extract motion conditions from the target pose sequence. The processing circuitry is further configured to implement a video diffusion model including a temporal attention mechanism, wherein the video diffusion model is configured to receive the appearance embedding and the motion conditions as input and generate a denoised animation sequence.
[0005] The Summary of the Invention is provided to introduce in a simplified form some concepts that will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Moreover, the claimed subject matter is not limited to embodiments that solve any or all disadvantages noted in any part of this disclosure. Brief Description of the Drawings
[0006] Figure 1 Shows a comparison of animations produced by the computing system and method of the present disclosure with the prior art in response to a series of motion signals.
[0007] Figure 2 Shows a schematic view of a computing system including processing circuitry configured to implement an appearance encoder, a pose control network, and a trained video diffusion model.
[0008] Figure 3 Shows a qualitative comparison between the present disclosure and a baseline on the datasets of Dataset A and Dataset B.
[0009] Figure 4A and Figure 4B shows a visualization of the ablation study, where the errors are highlighted in white boxes, and each frame has a reference image superimposed on the lower left corner and a target pose superimposed on the lower right corner.
[0010] Figure 5A 、 Figure 5B and Figure 5C shows the animation generalization results of unseen domains, text-to-image combinations, and multi-person animations.
[0011] Figure 6A and Figure 6B shows the quantitative comparison results between the present disclosure and the baselines on two benchmark datasets.
[0012] Figures 7A to 7E shows the data collected during the ablation on the social network platform dataset.
[0013] Figure 8 shows a flowchart of a computerized method for generating a denoised animation sequence according to the present disclosure.
[0014] Figure 9 shows a schematic view of an example computing environment in which the computing systems and methods described herein can be implemented. Detailed Description
[0015] 1. Introduction
[0016] Given a series of motion signals (such as video, depth, or pose), the image animation task aims to bring static images to life. Animating humans, animals, cartoons, or other general objects has attracted the attention of researchers. Among them, human image animation has been the most widely explored because of its potential applications in various fields including social media, the film industry, and the entertainment industry. Compared with traditional graphics methods, rich data makes it possible to develop low-cost data-driven animation frameworks.
[0017] Based on the generative backbone models used, existing data-driven methods for human image animation can be divided into two main categories, namely, GAN-based frameworks and diffusion-based frameworks. GAN-based frameworks typically use warping functions to deform the reference image to the target pose and utilize GAN models to extrapolate missing or occluded body parts. In contrast, diffusion-based frameworks use appearance and pose conditions to generate the target image based on a pre-trained diffusion model. Although visually plausible animations are generated, these methods usually have several limitations: 1) GAN-based methods have limited motion transfer ability, resulting in unrealistic details in occluded areas and limited generalization ability for cross-identity scenarios, such asFigure 1 as depicted. 2) On the other hand, diffusion-based methods process long videos frame-by-frame and then stack the results along the temporal dimension. This method ignores temporal consistency, leading to unstable results. Additionally, these works typically rely on a technique called CLIP to encode the reference appearance, which is known to be less effective in retaining details, as Figure 1 highlighted in the white box in
[0018] To address this challenge, the present disclosure provides a computing system implementing a human image animation framework, which is a diffusion-based framework, herein referred to as the "animation program", which aims to enhance temporal consistency, faithfully retain the reference image, and improve animation fidelity. First, the computing system includes a video diffusion model for encoding temporal information. Second, to maintain appearance coherence between frames, the computing system includes a new appearance encoder for retaining the complex details of the reference image. Leveraging these two innovations, the computing system further adopts a video fusion technique to facilitate smooth transitions for long video animations. Empirical results demonstrate that on two benchmarks, the present system and method outperform the baseline methods. Notably, on the challenging Dataset A dance dataset, the present method is more than 38% higher in video fidelity than the strongest baseline.
[0019] As demonstrated by the result data discussed below, the present system and method can achieve long-term temporal consistency, robust appearance encoding, and high per-frame quality. To achieve these results, a video diffusion model is utilized to encode temporal information by incorporating temporal attention blocks into the diffusion network. Second, an innovative appearance encoder is used to retain the human identity and background information obtained from the reference image. Different from existing works that adopt CLIP-encoded visual features, this appearance encoder is capable of extracting dense visual features to guide the animation, thus better retaining identity, background, clothing, etc. To further improve the fidelity of each frame, an image-video joint training strategy is adopted, leveraging different single-frame image data for augmentation, thereby providing richer visual cues and improving the framework's ability to model details. Finally, a video fusion technique is used to achieve smooth-transition long video animations.
[0020] The following disclosure describes a novel diffusion-based method for human image animation that integrates temporal consistency modeling, precise appearance encoding, and temporal video fusion for synthesizing temporally consistent human animations of arbitrary length. State-of-the-art performance is achieved on two benchmarks. Notably, on the challenging Dataset A dance dataset, the method described herein is more than 38% higher in video quality than the strongest baseline. Further, the method herein demonstrates robust generalization ability, applicable to cross-identity animations and various downstream applications, including animations of unseen domains and multi-person animations.
[0021] 2. Related Work
[0022] 2.1 Data-Driven Animation Production
[0023] Previous efforts in image animation production have mainly focused on the human body or face, leveraging rich and diverse training data and domain-specific knowledge such as key points, semantic parsing, and statistical parametric models. Based on these motion signals, a series of works have emerged. These methods can be classified into two categories based on their animation production pipelines, namely, implicit animation production and explicit animation production. Implicit animation production methods transform the source image into the target motion signal by deforming the reference image in the sub-expression space or manipulating the latent space of the generative model. The generative backbone synthesizes the animation conditioned on the target motion signal. In contrast, explicit methods warp the source image to the target through 2D optical flow, 3D deformation fields, or directly swapping the face of the target image. In addition to deforming the source image or 3D mesh, recent research work has also explored explicit deformation of points in 3D neural representations for human body and face synthesis, demonstrating improved temporal and multi-view Figure 1 consistency.
[0024] 2.2 Diffusion Models for Animation Production
[0025] The remarkable progress of diffusion models has led to unprecedented success in text-to-image generation, spawning a large number of subsequent works such as controllable image generation and video generation. Recent works have applied diffusion models to human-centered video generation and animation production. In these works, a common approach is to develop a diffusion model for generating 2D optical flow and then use frame warping techniques to animate the reference image. Additionally, many diffusion-based animation production frameworks adopt StableDiffusion as their image generation backbone and utilize ControlNet to regulate the animation process on the OpenPose key point sequence. For reference image conditioning, the pre-trained image-language model CLIP is usually adopted to encode the image into the semantic-level text symbol space and guide the image generation process through cross-attention. Although these works produce visually reasonable results, most of them process each video frame independently and ignore the temporal information in the animated video, which inevitably leads to unstable animation results.
[0026] 3. Methods Disclosed in This Subject
[0027] Given a reference image I ref and a motion sequence p 1:N = [p 1 ,..., p N , where N is the number of frames, the goal is to synthesize an animation with the appearance of I ref while complying with the provided motion p1:N Continuous video I 1:N = [I 1 ,..., I N .
[0028] Existing diffusion - based frameworks process each frame independently, ignoring the temporal consistency between different frames, which leads to unstable animations. To address this issue, a video diffusion model for temporal modeling is provided where a temporal attention block is incorporated into the diffusion backbone (Section 3.1). Additionally, existing works use a CLIP encoder to encode reference images. The semantic - level features introduced by the CLIP encoder are considered too sparse and compact to capture complex details. Therefore, this paper provides a novel appearance encoder (Section 3.2) to encode I ref into appearance embeddings and regulate the model to achieve animations that preserve identity and background.
[0029] Figure 2 The overall pipeline of the system of the present disclosure (Section 3.3) is depicted in Figure 2 A schematic view of a computing system 10 including a processing circuitry 12 is shown. The processing circuitry is configured to execute an animation program 18 to implement an appearance encoder 20, a pose control network 22, and a trained video diffusion model 24 (text - to - image diffusion model). For example, the computing system 10 may include a cloud server platform comprising multiple server devices, and the processing circuitry 12 may be one processor with a single server device or multiple processors with multiple server devices. The computer system 10 may also include one or more client devices communicating with the server devices, and the processing circuitry 12 may be located in such client devices.
[0030] The figure shows the processing pipeline employed by the computing system and method of the present disclosure. Given a reference image 16 and a target DensePose motion sequence 34, the computing system and method described herein employ a video diffusion model 24 and an appearance encoder 20 for temporal modeling and identity preservation, respectively (left figure). To support long - video animations, a video fusion strategy that produces smooth video transitions during inference is used (right figure).
[0031] First, at training time, the processing circuitry 12 is configured to implement the appearance encoder 20, which is configured to encode the reference image 16 into an appearance embedding 32, where the reference image 16 is embedded into the appearance embedding using the appearance encoder 20 For example, the reference image 16 may include a human image. The processing circuitry 12 is further configured to implement a pose control network 22, which is configured to receive a target pose sequence 34 as input and, in response, extract motion conditions 36 from the target pose sequence 34, where the target pose sequence 34 (i.e., DensePose) is passed to the Pose ControlNet to extract the motion conditions Conditioned on these two signals, the video diffusion model 24 is trained to animate the reference human identity to follow the given motion. In practice, due to memory limitations, the entire video is processed in a segment-by-segment manner.
[0032] At inference time, the processing circuitry 12 is configured to implement an appearance encoder 20, which is configured to encode the reference image 16 into an appearance embedding 32. The processing circuitry 12 is further configured to implement a pose control network 22, which is configured to receive a target pose sequence 34 as input and, in response, extract motion conditions 36 from the target pose sequence 34. The processing circuitry 12 is further configured to implement a trained video diffusion model 24 that includes a temporal attention mechanism, where the video diffusion model 24 is configured to receive the appearance embedding 32 and the motion conditions 36 as input and generate a denoised animation sequence 50.
[0033] Due to temporal modeling and robust appearance encoding, the animation program 18 can largely maintain temporal and appearance consistency between segments. Nevertheless, there are still slight discontinuities between segments. To alleviate this, video fusion methods are utilized to improve transition smoothness. Specifically, as Figure 2 depicted, the video animation is generated through multiple overlapping segments, where the entire video is decomposed into multiple overlapping segments and the predictions of the overlapping frames are averaged. Finally, joint training 38 is performed using an image dataset and a video dataset at training time, where an image-video joint training strategy is used to further enhance reference retention ability and single-frame fidelity (Section 3.4).
[0034] 3.1 Temporal Consistency Modeling
[0035] To ensure temporal consistency between video frames, the image diffusion model is extended to the video domain. Specifically, the original 2D UNet is inflated into a 3D temporal UNet by inserting temporal attention layers. The temporal UNet is denoted as where θ T are trainable parameters. The architecture of the inflated UNet block is as Figure 2 shown. First, randomly initialized latent noise is generated Among them, K is the length of the video frame. Then, K consecutive poses are stacked into a DensePose sequence p 1:K , for motion guidance. Next, the process reshapes the input features from into and inputs into the video diffusion backbone In the temporal module, the features are reshaped into to calculate cross-frame information along the temporal dimension.
[0036] The transferable motion prior is learned from video data. To learn the motion prior, the video diffusion model 24 includes a pre-trained motion module, and the pre-trained motion module includes a transformer that has been trained on consecutive video frames to predict the motion of image features in consecutive frames. The processing circuitry is further configured to perform a temporal video fusion operation using the pre-trained motion module to generate a denoised animation sequence. The pre-trained motion module is trained by adding sinusoidal frame-by-frame position encoding to consecutive frames of the training video to encode the position of each frame within the training video. The pre-trained motion module is further trained by training the motion module to predict the motion of visual features within the images of the consecutive video frames using a temporal attention mechanism, and the temporal attention mechanism uses the frame-by-frame position encoding to calculate the attention of each element of the image over the elements in the consecutive video frames.
[0037] Sinusoidal position encoding is added to enable the model to understand the position of each frame in the video. In this way, an attention operation is used to calculate the temporal attention, where Q = W Q z t , K = W K z t , V = W V z t . Through this attention mechanism, the animation program 18 aggregates the temporal information from adjacent frames and synthesizes K frames with improved temporal consistency, where the denoised animation sequence exhibits temporal consistency.
[0038] 3.2 Appearance Encoder
[0039] The goal of human image animation is to generate results under the guidance of a reference image I ref . The core goal of the appearance encoder of this subject is to represent I ref with detailed identity and background-related features, and these features can be injected into the video diffusion model for redirection under the guidance of motion signals. A novel appearance encoder is utilized, which can improve identity and background retention to enhance single-frame fidelity and temporal coherence. Specifically, the appearance encoder 20 creates a base UNet Another trainable copy, and calculate the reference image I for each denoising step t ref The conditional features of. This process is mathematically expressed as follows:
[0040]
[0041] is a set of normalized attention hidden states of the middle block and the upsampling block. Different from ControlNet that adds conditions in a residual manner, these features are passed to the spatial self-attention layer in the UNet block by connecting each feature in with the original UNet self-attention hidden state, thereby injecting appearance information. The appearance conditioning process is mathematically expressed as follows:
[0042]
[0043] where [·] represents the concatenation operation. Through this operation, the motion condition is concatenated with the pose control network hidden state and passed to the spatial self-attention layer of the trained video diffusion model 24, thereby passing part of the reference image 16 to the corresponding part of the target motion sequence. Through this operation, the spatial self-attention mechanism in the video diffusion model can be adjusted to a hybrid mechanism. This hybrid attention mechanism can not only maintain the semantic layout of the synthesized image, such as the pose and position of humans in the image, but also query the content of the reference image 16 during the denoising process to retain details, including identity, clothing, accessories, and background. This improved retention ability benefits the disclosed framework in two aspects: (1) The method can faithfully transfer the reference image 16 to the target motion; and (2) Due to retaining the same identity, background, and other details throughout the video, strong appearance conditioning helps with temporal consistency.
[0044] 3.3 Animation production pipeline
[0045] By incorporating temporal consistency modeling and the appearance encoder 20, these elements are combined with pose conditioning (i.e., ControlNet) to transform the reference image 16 into the target pose.
[0046] Motion transfer. ControlNet for OpenPose key points is usually used to animate the reference human image. One challenge of this method is that the main body key points are sparse and not robust enough for some motions (such as rotation). Therefore, DensePose is selected as the motion signal p i to obtain dense and robust pose conditions. The pose ControlNet is adopted where the pose condition for frame i is calculated as follows
[0047]
[0048] Among them, is a set of conditional residuals of the residuals of the intermediate blocks and upsampling blocks added to the diffusion model. In the pipeline, the motion features of each pose are concatenated in a DensePose sequence into
[0049] the denoising process. Under the appearance condition and the motion condition , the animation program animates the reference image 16 according to the DensePose sequence. During the denoising process, the noise estimation function is mathematically expressed by the following formula:
[0050]
[0051] where θ is the set of all trainable parameters (i.e., θ T , θ a and θ p ).
[0052] Long video animation. Using temporal consistency modeling and appearance encoders, temporally consistent human image animation results of arbitrary length can be generated via per-segment processing. However, because the temporal attention block cannot model the long-term consistency between different segments, unnatural transitions and inconsistent details may occur between segments.
[0053] To address this challenge, a sliding window method is adopted in the inference stage to improve the transition smoothness. In other words, the sliding window technique is applied during inference to smooth the transitions between segments, where the denoised animation sequence is a video animation generated from multiple segments. As Figure 2 shown, the long motion sequence is divided into multiple segments with temporal overlap, where the length of each segment is K. First, for the entire video with N frames, the noise is sampled and the noise is also divided into overlapping noise segments where and s is the overlap stride, where s < K. If (N - K) mod (K - s) ≠ 0, i.e., the size of the last segment is less than K, for simplicity, only the first few frames are used for padding to construct a K-frame segment. In addition, it is empirically found that sharing the same initial noise for all segments can improve the video quality. For each denoising time step t, the predicted noise is obtained for each segment and is then merged into by averaging the overlapping frames. When t = 0, the final animated video I 1:N is obtained.
[0054] 3.4 Training
[0055] Learning objective. The animation program 18 adopts a multi-stage training strategy. In the first stage, the temporal attention layer is temporarily omitted, and the appearance encoder 20 is trained together with the pose control network 22, where all attention layers are temporarily omitted to train the appearance encoder 20 and the pose control network 22 together. The loss term for this stage is calculated as follows
[0056]
[0057] p i is the DensePose of the target image I i The learnable modules are and In the second stage, only the temporal attention layer in is optimized, and the learning objective is formulated as
[0058]
[0059] Image-video joint training. Compared with image datasets, the scale of human video datasets is much smaller, and the diversity in terms of identity, background, and pose is lower. This limits the effective learning of the reference condition ability of the proposed animation framework. To alleviate this problem, an image-video joint training strategy is adopted.
[0060] In the first stage, when pre-training the appearance encoder 20 and the pose control network 22, a probability threshold τ 0 is set for sampling human images from the large-scale image dataset. A random number r ∼ U(0, 1) is taken, where U(·, ·) represents the uniform distribution. If r ≤ τ 0 , the sampled image is used for training. In this case, the adjusted pose p ref is estimated according to I i , and the learning objective of the framework becomes reconstruction.
[0061] Although introducing temporal attention in the second stage helps improve temporal modeling, it is worth noting that this leads to a decrease in the quality of each frame. To improve temporal coherence and maintain single-frame image fidelity simultaneously, joint training is also adopted in this stage. Specifically, two probability thresholds τ 1 and τ 2 are empirically selected and compared with r ∼ U(0, 1). When r ≤ τ 1 , training data is sampled from the image dataset, otherwise data is sampled from the video dataset. Based on different training data, the denoising process in the training stage is formulated as
[0062]
[0063] 4. Experiments
[0064] The performance of the animation program was evaluated using two datasets, namely, Dataset A and Dataset B. Dataset A consists of 350 dancing videos, while Dataset B consists of 1203 video clips extracted from an online video sharing platform. To ensure a fair comparison with state-of-the-art methods, Dataset A was evaluated using the same test set as DisCo, and Dataset B was divided into training / test sets. All datasets underwent the same preprocessing pipeline.
[0065] 4.1. Comparisons
[0066] Baselines. The animation program was comprehensively compared with several state-of-the-art human image animation methods as follows: (1) State-of-the-art GAN-based animation method 1, which estimates optical flow from the driving sequence to warp the source image and then uses a GAN model to repair the occluded regions. (2) State-of-the-art diffusion-based animation method 2, which integrates pose, human, and background disentanglement conditional modules into a pre-trained diffusion model to perform human image animation. (3) An additional baseline, labeled IPA+CtrlN, was constructed by combining a state-of-the-art image conditional method 3 with a pose control network. To make a fair comparison, a temporal attention block was added to this framework, and a video version baseline labeled IPA+CtrlNV was constructed. Additionally, method 1 uses the ground truth video as the driving signal. To ensure a fair comparison, an alternative version of method 1 was trained using the same driving signal (DensePose) as the animation program.
[0067] Evaluation metrics. When making these comparisons, the established evaluation metrics adopted in previous studies were followed. For Dataset A, the single-frame image quality and video fidelity were evaluated. The metrics for single-frame quality include L1 error, SSIM, LPIPS, PSNR, and FID. The video fidelity was evaluated by FID-FVD and FVD. On Dataset B, the following L1 error, average keypoint distance (AKD), missing keypoint rate (MKR), and average Euclidean distance (AED) for method 1 were reported. However, these evaluation metrics are designed for single-frame evaluation and lack a perceptual measurement of the animation results. Therefore, FID, FID-VID, and FVD were also calculated on Dataset B to measure the perceptual quality of the images and videos.
[0068] Figure 6A and Figure 6B Shows the quantitative comparison results of the present disclosure and the baselines on two benchmark datasets (Dataset A and Dataset B). Figure 6A and Figure 6BIncluding a quantitative comparison with the baseline, where the best results are presented in bold and the second-best results are underlined. The symbol * in the figure indicates that the original Method 1 directly uses the ground truth video frames for animation, and the presented results are for reference only. As Figure 6A shown, in terms of the reconstruction metrics (i.e., L1, PSNR, SSIM, and LPIPS) of Dataset A, this method outperforms all baselines. Notably, for SSIM and LPIPS, this disclosure improves by 6.9% and 18.2% respectively compared to the strongest baseline (DisCo). Additionally, this disclosure achieves state-of-the-art video fidelity, demonstrating a significant 63.7% improvement in the performance of FID-VID and a significant 38.8% improvement in the performance of FVD compared to DisCo. As Figure 6B shown, this method also exhibits excellent video fidelity on Dataset B, achieving the best FID-VID of 19.00 and an FVD of 131.51. This performance is particularly significant compared to the second-best method (Method 1), where the FVD is improved by 28.1%. Additionally, this disclosure demonstrates state-of-the-art single-frame fidelity, obtaining the best FID score of 22.78. Compared to the diffusion-based baseline method DisCo, this disclosure demonstrates a significant improvement of 17.2%. However, it is important to note that this disclosure has a higher L1 error compared to the baseline. This may be due to the lack of background information in the DensePose control signal. Therefore, this disclosure is unable to learn a consistent dynamic background as presented in the TED talk dataset, resulting in an increase in the L1 error. Nevertheless, this disclosure achieves an L1 error comparable to the strongest baseline (Method 1) in the foreground human region, demonstrating its effectiveness in human animation. Furthermore, this method achieves the best performance in AKD, MKR, and AED, demonstrating its excellent identity protection ability and animation accuracy.
[0069] Figure 3Shows a qualitative comparison between the present disclosure and the baselines on the datasets of Dataset A and Dataset B. The target pose is superimposed on the lower left corner of the synthetic frame, and the artifacts generated by the strongest baseline (DisCo) are highlighted in the white box. Notably, the dance videos in Dataset A exhibit significant pose variations, which pose challenges to GAN-based methods (such as Method 1) as they struggle to produce reasonable results when there are significant pose differences between the reference image 16 and the driving signal. In contrast, diffusion-based baselines (IPA+CtrlN, IPA+CtrlN-V, and DisCo) show better single-frame quality. However, since IPA+CtrlN and DisCo generate each frame independently, their temporal consistency is not satisfactory, as evidenced by the color variations of the clothes and inconsistent backgrounds in the occluded regions. The video diffusion baseline IPA+CtrlN-V shows more consistent content but has poor single-frame quality due to weak reference conditioning. Instead, the present disclosure generates temporally consistent animations and high-fidelity details for the background, clothes, face, and hands.
[0070] Different from Dataset A, Dataset B includes speech videos recorded under dim lighting conditions. The motions in the Dataset B dataset mainly involve gestures and are less challenging than dance videos. Therefore, Method 1 produces visually more reasonable results, although the motions are inaccurate. In contrast, IPA+CtrlN, IPA+CtrlN-V, DisCo, and the present disclosure demonstrate more precise body pose control capabilities because these methods extract appearance conditions from the reference image to guide the animation rather than directly warping the source image. As Figure 3 shown, among all these methods, the present disclosure exhibits excellent identity and background retention capabilities because the appearance encoder 20 of the present system extracts detailed information from the reference image.
[0071] Cross-identity animation. In addition to animating each identity with its corresponding motion sequence, the cross-identity animation capabilities of the present disclosure and the state-of-the-art baselines (i.e., DisCo and Method 1) were investigated. Specifically, two DensePose motion sequences were sampled from the test set of Dataset A, and these sequences were used to animate reference images from other videos. Figure 1 Shows a comparison of the animations produced by the computing system and method of the present disclosure with the prior art in response to a series of motion signals. The systems and methods described herein generate temporally consistent animations for the reference identity images, while the state-of-the-art methods fail to generalize or retain the reference appearance, as highlighted by the white boxes in the figure. In the figure, the motion sequences are superimposed at the corners. Note that Method 1 directly uses the video frames as the driving signal. Figure 1It is shown that Method 1 fails to generalize driving videos containing a large number of pose differences, while DisCo has difficulty in retaining details in the reference image, resulting in artifacts in the background and clothing. In contrast, given the target motion, the proposed method faithfully animates the reference image, thus demonstrating its robustness.
[0072] 4.2. Ablation Study
[0073] Figures 7A to 7E Data collected during the ablation of the present disclosure on Dataset A is shown, with the best results in bold. The architecture design and training strategy are different to study their effectiveness. Specifically, Figure 7A The effect of modeling temporal information is shown. Figure 7B The effect of the appearance encoder 20 is shown. Figure 7C The effect of image-video joint training is shown. Figure 7D The effect of temporal video fusion in the inference stage is shown. Figure 7E The effect of all video segments sharing the same initial noise is shown. To simplify the numerical values, Figures 7A to 7E L1×10 is reported -4 . To verify the effectiveness of the design choices in the present disclosure, ablation experiments are conducted on Dataset A, which is characterized by significant pose variations, a wide range of identities, and diverse backgrounds.
[0074] Temporal modeling. To evaluate the impact of the proposed temporal attention layer, a version of the present disclosure is trained without the temporal attention layer for comparison. Figure 7A The results presented in show that when the temporal attention layer is discarded, both the single-frame quality and video fidelity evaluation metrics decrease, highlighting the effectiveness of the temporal modeling of the proposed method. This is further supported by Figure 4A the qualitative ablation results presented in, where models without explicit temporal modeling cannot maintain the temporal coherence of humans and the background.
[0075] Appearance encoder. To evaluate the enhancement brought by the proposed appearance encoding strategy, the appearance encoder in the present disclosure is replaced with CLIP and IP-Adapter to establish baselines. Figure 7B summarizes the ablation results. Apparently, the proposed method is significantly superior to these two baselines in terms of reference image retention, with significant improvements in both single-frame and video fidelity.
[0076] Inference stage video fusion. The present disclosure utilizes video fusion technology to enhance the transition smoothness of long-term animations. Figure 7D and Figure 7EDemonstrated the effectiveness of the design choices in the present system and method. Generally speaking, skipping video fusion or using different initial random noises for different video segments will reduce the animation performance, which can be seen from the degradation in appearance and video quality.
[0077] Image-video joint training. An image-video joint training strategy is introduced to improve the animation quality. As Figure 7C shown, applying image-video joint training in the appearance encoding and temporal modeling stages consistently improves the animation quality. This improvement can also be observed in Figure 4B . Without the joint training strategy, the model has difficulty modeling complex details and tends to produce incorrect clothing and accessories, as Figure 4B shown.
[0078] 4.3. Applications
[0079] Although trained only on real human data, this disclosure demonstrates the ability to generalize to various application scenarios, including animating data from unseen domains, integration with text-to-image diffusion models, and multi-person animation. Figure 5A 、 Figure 5B and Figure 5C show the animation generalization results for unseen domains, text-to-image combinations, and multi-person animation. Figure 5A shows the animation results for unseen domains. Figure 5B shows the combination of this disclosure with generative AI technologies. Figure 5C shows multi-person animation. The motion signals are superimposed at the corners of each frame in Figure 5A and Figure 5B .
[0080] Animation of unseen domains. This disclosure demonstrates the generalization ability for unseen image styles and motion sequences. As Figure 5A shown, the present method can animate oil paintings and movie images to perform actions such as running and yoga while maintaining a stable background and repairing occluded areas with temporally consistent results.
[0081] Combined text-to-image generation. Due to its strong generalization ability, this disclosure can be used to animate images generated by text-to-image (T2I) models (e.g., generative AI technologies). As Figure 5B shown, first, generative AI technologies are used to synthesize reference images through various prompts. Then, our model can animate these reference images to perform various actions.
[0082] Multi-person animation. This disclosure also shows strong generalization ability for multi-person animation. As Figure 5CAs shown, given a reference frame and a motion sequence including two dancing individuals, animations can be generated for multiple individuals.
[0083] Figure 8 FIG. 1 shows a flowchart of a computerized method 100 for generating a denoised animation sequence according to the present disclosure. Method 100 may be implemented by Figure 2 the computing system 10 shown. At 102, method 100 may include encoding a reference image into an appearance embedding. At 104, method 100 may further include receiving a target pose sequence as input and, in response, extracting motion conditions from the target pose sequence. At 106, method 100 may further include receiving the appearance embedding and the motion conditions as input via a trained video diffusion model including a temporal attention mechanism, and generating a denoised animation sequence.
[0084] The present disclosure describes a computing system and method, referred to as an animation program, which is a novel diffusion-based framework designed for human avatar animation, with a particular focus on temporal consistency. By effectively modeling temporal information, the present disclosure enhances the overall temporal coherence of the animation results. The proposed appearance encoder not only improves the single-frame quality but also contributes to improving temporal consistency. Additionally, the integration of video frame fusion techniques enables seamless transitions on the animated video. The present disclosure demonstrates state-of-the-art performance in terms of both single-frame and video quality. Moreover, its strong generalization ability makes it applicable to unseen domains and multi-person animation scenarios.
[0085] In some embodiments, the methods and processes described herein may be related to the computing systems of one or more computing devices. Specifically, such methods and processes may be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.
[0086] Figure 9 FIG. 2 schematically shows a non-limiting embodiment of a computing system 800 that may implement one or more of the above methods and processes. Computing system 800 is shown in a simplified form. Computing system 800 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (e.g., smartphones), and / or other computing devices, as well as wearable computing devices (such as smartwatches and head-mounted augmented reality devices).
[0087] Computing system 200 includes a logical processor 202, volatile memory 204, and a non-volatile storage device 206. Computing system 200 may optionally include a display subsystem 202, an input subsystem 210, a communication subsystem 212, and / or Figure 9 other components not shown in FIG. 3.
[0088] The logical processor 202 includes one or more physical devices configured to execute instructions. For example, a logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform a task, implement a data type, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.
[0089] The logical processor may include one or more physical processors (hardware) configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 202 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logical processor may be distributed across two or more separate devices, which may be remotely located and / or configured to coordinate processing. The various aspects of the logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. It will be understood that, in such cases, these virtualized aspects run on different physical logical processors of various different machines.
[0090] The non-volatile storage device 206 includes one or more physical devices configured to store instructions executable by the logical processor to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 206 may change, for example to store different data.
[0091] The non-volatile storage device 206 may include removable and / or built-in physical devices. The non-volatile storage device 206 may include optical memory (e.g., CD, DVD, HD-DVD, Blu-ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and / or magnetic memory (e.g., hard disk drive, floppy disk drive, tape drive, MRAM, etc.) or other mass storage device technologies. The non-volatile storage device 206 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that the non-volatile storage device 206 is configured to store instructions even when the power to the non-volatile storage device 206 is turned off.
[0092] The volatile memory 204 may include a physical device that includes random access memory. The volatile memory 204 is typically used by the logic processor 202 to temporarily store information during the processing of software instructions. It should be understood that when the power supply to the volatile memory 204 is cut off, the volatile memory 204 generally does not continue to store instructions.
[0093] Aspects of the logic processor 202, the volatile memory 204, and the non-volatile storage device 206 may be integrated together into one or more hardware logic components. Such hardware logic components may include field programmable gate arrays (FPGAs), programmed and application specific integrated circuits (PASIC / ASICs), programmed and application specific standard products (PSSP / ASSPs), systems on a chip (SOCs), and complex programmable logic devices (CPLDs), among others.
[0094] The terms "module", "program", and "engine" may be used to describe an aspect of the computing system 200 that is typically implemented in software by a processor to perform a specific function using portions of the volatile memory, the function involving transformative processing for specifically configuring the processor to perform the function. Thus, a module, program, or engine may be instantiated by executing instructions saved by the non-volatile storage device 206 using portions of the volatile memory 204 via the logic processor 202. It should be understood that different modules, programs, and / or engines may be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may encompass single or multiple sets of executable files, data files, libraries, drivers, scripts, database records, etc.
[0095] The display subsystem 208 (when included) may be used to present a visual representation of data saved by the non-volatile storage device 206. The visual representation may take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data saved by the non-volatile storage device and thus transform the state of the non-volatile storage device, the state of the display subsystem 208 may also be transformed accordingly to visually represent the changes in the underlying data. The display subsystem 208 may include one or more display devices that utilize almost any type of technology. Such display devices may be combined with the logic processor 202, the volatile memory 204, and / or the non-volatile storage device 206 in a shared enclosure, or such display devices may be peripheral display devices.
[0096] The input subsystem 210 (when included) may include one or more user input devices, such as a keyboard, mouse, touch screen, or game controller, or interface with such user input devices. In some embodiments, the input subsystem may include or interface with selected natural user input (NUI) components. These components may be integrated or peripheral components, and the conversion and / or processing of input actions may be performed on-board or off-board. Example NUI components may include a microphone for voice and / or speech recognition; infrared, color, stereo, and / or depth cameras for machine vision and / or gesture recognition; a head tracker, eye tracker, accelerometer, and / or gyroscope for motion detection and / or intent recognition; and / or any other suitable sensors.
[0097] The communication subsystem 212 (when included) may be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 212 may include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem may be configured to communicate via a wireless telephone network, or a wired or wireless local or wide area network. In some embodiments, the communication subsystem may allow the computing system 200 to send messages to and / or receive messages from other devices via a network such as the Internet.
[0098] The following paragraphs provide additional description of the subject matter of the present disclosure. In one aspect, a computer system is provided that includes processing circuitry configured to implement an appearance encoder configured to encode a reference image into an appearance embedding. The processing circuitry is further configured to implement a pose control network configured to receive a target pose sequence as input and, in response, extract motion conditions from the target pose sequence. The processing circuitry is further configured to implement a trained video diffusion model including a temporal attention mechanism, the video diffusion model being configured to receive the appearance embedding and the motion conditions as input and generate a denoised animation sequence.
[0099] In this regard, the video diffusion model includes a pre-trained motion module that includes a transformer that has been trained on consecutive video frames to predict the motion of image features in consecutive frames, and the processing circuitry is further configured to use the pre-trained motion module to perform a temporal video fusion operation to generate the denoised animation sequence.
[0100] In this regard, the pre-trained motion module is trained by adding sinusoidal frame-by-frame positional encoding to consecutive frames of the training video to encode the position of each frame within the training video. The pre-trained motion module is further trained by training the motion module to predict the motion of visual features within the images of consecutive video frames by using a temporal attention mechanism that uses the frame-by-frame positional encoding to compute the attention of each element of the image over the elements in consecutive video frames.
[0101] In this regard, the denoised animation sequence is a video animation generated by multiple segments, and a sliding window technique is applied during inference to smooth the transition between segments.
[0102] In this regard, the video animation is generated by multiple overlapping segments, and the predictions of the overlapping frames are averaged.
[0103] In this regard, the reference image includes a human image.
[0104] In this regard, the motion conditions are connected to the pose control network hidden state and are passed to the spatial self-attention layer of the trained video diffusion model, thereby passing parts of the reference image to the corresponding parts of the target motion sequence.
[0105] In this regard, the denoised animation sequence exhibits temporal consistency.
[0106] In this regard, at training time, all attention layers are temporarily omitted to train the appearance encoder together with the pose control network.
[0107] In this regard, at training time, joint training is performed using an image dataset and a video dataset.
[0108] On the other hand, a computerized method is provided that includes encoding a reference image into an appearance embedding. The method further includes receiving a target pose sequence as input and, in response, extracting motion conditions from the target pose sequence. The method further includes receiving the appearance embedding and the motion conditions as input via a trained video diffusion model that includes a temporal attention mechanism, and generating a denoised animation sequence.
[0109] In this regard, the video diffusion model includes a pre-trained motion module that includes a transformer that has been trained on consecutive video frames to predict the motion of image features in consecutive frames, and the computerized method further includes using the pre-trained motion module to perform a temporal video fusion operation to generate the denoised animation sequence.
[0110] In this regard, the pre-trained motion module is trained by adding sinusoidal per-frame positional encoding to consecutive frames of the training video to encode the position of each frame within the training video. The pre-trained motion module is further trained by using a temporal attention mechanism to train the motion module to predict the motion of visual features within the images of consecutive video frames, where the temporal attention mechanism uses the per-frame positional encoding to calculate the attention of each element of the image over the elements in the consecutive video frames.
[0111] In this regard, the denoised animation sequence is a video animation generated from multiple segments, and a sliding window technique is applied during inference to smooth the transitions between segments.
[0112] In this regard, the video animation is generated from multiple overlapping segments, and the predictions of the overlapping frames are averaged.
[0113] In this regard, the reference image includes a human image.
[0114] In this regard, the motion conditions are connected to the pose control network hidden state and are passed to the spatial self-attention layer of the trained video diffusion model, thereby passing parts of the reference image to corresponding parts of the target motion sequence.
[0115] In this regard, the denoised animation sequence exhibits temporal consistency.
[0116] In this regard, at training time, all attention layers are temporarily omitted to train the appearance encoder together with the pose control network.
[0117] On the other hand, a non-transitory computer-readable storage medium storing computer-executable instructions is provided. When executed by a processing circuitry, the computer-executable instructions cause the processing circuitry to be configured to encode a reference image into an appearance embedding. The processing circuitry is further configured to receive a target pose sequence as input and, in response, extract motion conditions from the target pose sequence. The processing circuitry is further configured to receive the appearance embedding and the motion conditions as input via a trained video diffusion model including a temporal attention mechanism and generate a denoised animation sequence.
[0118] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered restrictive as there may be various variations. The specific routines or methods described herein may represent one or more of any number of processing strategies. Accordingly, the various actions illustrated and / or described may be performed in the order illustrated and / or described, in other orders, in parallel, or omitted. Similarly, the order of the above processes may also be changed.
[0119] The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems, and configurations, as well as other features, functions, acts, and / or properties disclosed herein, and any and all equivalents thereof.
Claims
1. A computing system comprising: processing circuitry configured to implement: an appearance encoder configured to encode a reference image into an appearance embedding; a gesture control network configured to receive as input a target gesture sequence and, in response, extract motion conditions from the target gesture sequence; as well as A trained video diffusion model, the video diffusion model comprising a temporal attention mechanism, the video diffusion model configured to receive the appearance embedding and the motion condition as input and generate a denoised animation sequence.
2. The computing system of claim 1, wherein: The video diffusion model includes a pre-trained motion module, the pre-trained motion module includes a transformer that has been trained on consecutive video frames to predict the motion of image features in consecutive frames, and The processing circuitry is further configured to perform a temporal video fusion operation using the pre-trained motion module to generate the denoised animation sequence.
3. The computing system of claim 2, wherein: The pre-trained motion module is trained in the following way: Adding sinusoidal frame-by-frame position coding to consecutive frames of a training video to encode the position of each frame in the training video; as well as The motion module is trained to predict the motion of visual features within the images of the consecutive video frames using a temporal attention mechanism, wherein the temporal attention mechanism uses the frame-by-frame position encoding to compute the attention of each element of the image over the elements in the consecutive video frames.
4. The computing system of claim 1, wherein: A denoised animation sequence is a video animation generated from multiple clips, and A sliding window technique is applied during inference to smooth the transition between segments.
5. The computing system of claim 4, wherein: The video animation is generated by a plurality of overlapping segments, and The predictions of overlapping frames are averaged.
6. The computing system of claim 1, wherein: The reference image includes an image of a human being.
7. The computing system of claim 1, wherein: The motion conditions are concatenated with the posture control network hidden states and passed to the spatial self-attention layer of the trained video diffusion model to transfer parts of the reference image to corresponding parts of the target motion sequence.
8. The computing system of claim 1, wherein the denoised animation sequence exhibits temporal consistency.
9. The computing system of claim 1, wherein at training time, all attention layers are temporarily omitted to train the appearance encoder together with the posture control network.
10. The computing system of claim 1, wherein at training time, joint training is performed using an image dataset and a video dataset.
11. A computerized method comprising: Encode the reference image as an appearance embedding; receiving as input a target posture sequence, and in response, extracting motion conditions from the target posture sequence; as well as The appearance embedding and the motion condition are received as input via a trained video diffusion model including a temporal attention mechanism, and a denoised animation sequence is generated.
12. The computerized method of claim 11, wherein: The video diffusion model includes a pre-trained motion module, the pre-trained motion module includes a transformer that has been trained on consecutive video frames to predict the motion of image features in consecutive frames, and The computerized method further includes performing a temporal video fusion operation using the pre-trained motion module to generate the denoised animation sequence.
13. The computerized method of claim 12, wherein: The pre-trained motion module is trained in the following way: Adding sinusoidal frame-by-frame position coding to consecutive frames of a training video to encode the position of each frame in the training video; as well as The motion module is trained to predict the motion of visual features within the images of the consecutive video frames using a temporal attention mechanism, wherein the temporal attention mechanism uses the frame-by-frame position encoding to compute the attention of each element of the image over the elements in the consecutive video frames.
14. The computerized method of claim 11, wherein: A denoised animation sequence is a video animation generated from multiple clips, and A sliding window technique is applied during inference to smooth the transition between segments.
15. The computerized method of claim 11, wherein: The video animation is generated by a plurality of overlapping segments, and The predictions of overlapping frames are averaged.
16. The computerized method of claim 11, wherein the reference image comprises an image of a human being.
17. A computerized method according to claim 11, wherein motion conditions are connected to the posture control network hidden state and passed to the spatial self-attention layer of the trained video diffusion model to transfer portions of the reference image to corresponding portions of the target motion sequence.
18. The computerized method of claim 11, wherein the denoised animation sequence exhibits temporal consistency.
19. The computerized method of claim 11, wherein at training time, all attention layers are temporarily omitted to train the appearance encoder together with the posture control network.
20. A non-transitory computer-readable storage medium storing computer-executable instructions, wherein: When executed by processing circuitry, the computer executable instructions cause the processing circuitry to be configured to: Encode the reference image as an appearance embedding; receiving as input a target posture sequence, and in response, extracting motion conditions from the target posture sequence; and The appearance embedding and the motion condition are received as input via a trained video diffusion model including a temporal attention mechanism, and a denoised animation sequence is generated.