Generating images using machine learning models

By training across identities and using ControlNet technology, the machine learning model effectively solves the problems of appearance leakage and identity drift in cross-identity image generation, generating high-fidelity, vivid animations that showcase subtle differences in head movements and facial expressions.

CN120997348APending Publication Date: 2025-11-21FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510650445.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-29
Filing Date
2025-05-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing machine learning models are prone to appearance leakage and identity drift when generating cross-identity images, making it difficult to effectively preserve the appearance features of the source object and achieve high-quality head pose and facial animation.

Method used

By employing a cross-identity training scheme and assisted ControlNet technology, identity features are decoupled through the second sub-model of the machine learning model. Combined with local motion control, high-fidelity videos are generated, reducing appearance leakage and enhancing the derivation of subtle facial features.

Benefits of technology

It effectively preserves the appearance features of the source object in cross-identity image generation, generating vivid and expressive animations that showcase subtle differences in head movements and facial expressions, thus enhancing the realism and robustness of the animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997348A_ABST
    Figure CN120997348A_ABST
Patent Text Reader

Abstract

This disclosure describes techniques for generating images using a machine learning model. A source image and a drive image are received. The source image includes an image of the first object. The drive image includes a second object and the drive image depicts a pose or appearance. Appearance features of a first object are extracted from a source image through a first sub-model of a machine learning model. A mask image is generated based on the driving image. The mask image comprises a mouth region and / or an eye region in the driving image. A pose or appearance is derived based on the drive image and the mask image by a second sub-model of the machine learning model. The image is generated through a machine learning model. The image preserves an appearance feature of the first object and follows a pose or appearance depicted in the drive image.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 649,734, filed May 20, 2024, and U.S. Patent Application No. 19 / 193,811, filed April 29, 2025, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Machine learning models are increasingly being used across various industries to perform a wide range of tasks. These tasks can include content generation. Improvements in techniques for leveraging machine learning models for content generation are expected. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for generating an image using a machine learning model is provided. The method includes: receiving a source image and a driving image via the machine learning model, wherein the source image includes an image of a first object, the driving image includes a second object different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model; extracting appearance features of the first object from the source image via the first sub-model; generating a mask image based on the driving image, wherein the mask image includes at least one of a mouth region or an eye region in the driving image; deriving the pose or appearance based on the driving image and the mask image via the second sub-model; and generating an image via the machine learning model, wherein the generated image retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

[0004] In a second aspect of this disclosure, a system for generating images using a machine learning model is provided. The system includes at least one processor; and at least one memory communicatively coupled to the at least one processor and including computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations. The operations include: receiving a source image and a driving image via a machine learning model, wherein the source image includes an image of a first object, the driving image includes a second object different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model; extracting appearance features of the first object from the source image via the first sub-model; generating a mask image based on the driving image, wherein the mask image includes at least one of a mouth region or an eye region in the driving image; deriving a pose or appearance based on the driving image and the mask image via the second sub-model; and generating an image via the machine learning model, wherein the generated image retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

[0005] In a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer-readable instructions that, when executed by a processor, cause the processor to perform operations. The operations include: receiving a source image and a driving image via a machine learning model, wherein the source image includes an image of a first object, the driving image includes a second object different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model; extracting appearance features of the first object from the source image via the first sub-model; generating a mask image based on the driving image, wherein the mask image includes at least one of a mouth region or an eye region in the driving image; deriving a pose or appearance based on the driving image and the mask image via the second sub-model; and generating an image via the machine learning model, wherein the generated image retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

[0006] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0007] The following detailed description will be better understood when read in conjunction with the accompanying drawings. Exemplary embodiments of various aspects of this disclosure are shown in the drawings for illustrative purposes; however, the invention is not limited to the specific methods and means disclosed.

[0008] Figure 1 An example system for generating images and / or videos using a machine learning model, according to this disclosure, is shown.

[0009] Figure 2 An example system for generating images and / or videos using a machine learning model, according to this disclosure, is shown.

[0010] Figure 3 An example system for generating videos using a machine learning model, according to this disclosure, is shown.

[0011] Figure 4 An example system for training machine learning models according to this disclosure is shown.

[0012] Figure 5 An example process for generating images using a machine learning model, according to this disclosure, is shown.

[0013] Figure 6 An example process for training a second sub-model of a machine learning model according to this disclosure is shown.

[0014] Figure 7 An example process for generating a control image according to this disclosure is shown.

[0015] Figure 8 An example process for training a second sub-model of a machine learning model according to this disclosure is shown.

[0016] Figure 9 An example process for training a machine learning model according to this disclosure is shown.

[0017] Figure 10 An example process for generating video using a machine learning model, according to this disclosure, is shown.

[0018] Figure 11 An example computing device is shown that can be used to perform any of the techniques disclosed herein. Detailed Implementation

[0019] Machine learning models can be used to generate avatar animations. Specifically, machine learning models can be used to animate static avatar images using motion information (such as head pose and / or facial aspects / appearance) derived from driving images or videos, where the driving images in the video typically represent a different subject than the static avatar image. Avatar animation has already become significant in a variety of downstream applications, such as video conferencing, visual effects, and digital agents.

[0020] This paper describes an improved technique for generating avatar animation. This improved technique can be used to generate high-fidelity videos of natural avatars in various styles, exhibiting highly dynamic head poses and expressive facial features. The machine learning model can leverage image diffusion priors for high-expressive avatar animation and a pose control scheme to mitigate expressiveness loss and appearance leakage. To fully preserve the driving head poses and facial features, motion is interpreted directly from the original driving images without resorting to any intermediate motion representations. A motion transfer network is employed to generate cross-identity training image pairs for training the machine learning model. The cross-identity driven training scheme simultaneously mitigates appearance leakage, enabling direct avatar animation during inference without any preprocessing. To further enhance the derivation of subtle facial features at minute scales, an auxiliary control network (ControlNet) is employed to direct conditioned motion attention to local facial movements.

[0021] Figure 1 An example system 100 according to this disclosure is shown. System 100 can be used for image or video generation using machine learning model 103. For example, system 100 can generate an output image or video using a single image and multiple driving frames from a driving video.

[0022] Source image 101 and driving image / video 102 can be input into machine learning model 103. Source image 101 may include an image of an object. Source image 101 may include an image of the object's face. Driving image / video 102 may depict a pose (e.g., head pose) or appearance (e.g., facial features). In embodiments, driving image / video 102 may depict the same object in a specific pose or with a specific appearance. For example, driving image / video 102 and source image 101 (e.g., driving image / video 102 and source image 101 may be different frames from the same video) may be extracted from the same video. In other embodiments, driving image / video 102 may depict different objects in a specific pose or with a specific appearance.

[0023] Machine learning model 103 can be trained to generate output image / video 122 based on transmitting head pose and / or facial aspects associated with driving image / video 102 to the object depicted in source image 101. For example, if the object depicted in source image 101 has a first appearance or pose (e.g., smiling), and driving image / video 102 depicts an object with a different appearance or pose (e.g., not smiling) (e.g., different object), then machine learning model 103 can generate output image / video 122 depicting the object in the source image with a different appearance or pose (e.g., not smiling).

[0024] Figure 2 An example system 200 according to this disclosure is shown. System 200 can be used for image or video generation using machine learning model 103. Machine learning model 103 may include a first sub-model 204, a second sub-model 212, and a third sub-model 215.

[0025] Machine learning model 103 may receive source image 201 and driving image / video 202. Source image 201 may include an image of a first object, while driving image / video 202 may include at least one image of a second object different from the first object. Driving image / video 202 may depict pose (e.g., head pose) or appearance (e.g., facial features). Source image 201 may be input into a first sub-model 204. First sub-model 204 may extract identity features 206 (appearance features, such as facial features) of the first object from source image 201. At least one mask image 210 may be generated based on driving image / video 202. Multiple mask images 210 may include at least one mouth region or eye region from driving image / video 202. Driving image / video 202 and multiple mask images 210 may be input into a second sub-model 212. Second sub-model 212 may derive motion information 214, such as information indicating pose or appearance, based on driving image / video 202 and multiple mask images 210. The third sub-model 215 can be trained to achieve temporal smoothness. The machine learning model 103 can generate an output image / video 122 with temporal smoothness based on identity features 206 and motion information 214. The output image / video 122 can retain the identity features of the first object, and the output image / video 122 can follow the pose or appearance depicted in the driving image / video 202. In some embodiments, the machine learning model 103 can utilize a frozen pre-trained latent diffusion model as a rendering backbone and merge the three sub-models 204, 212, and 215 for decoupled control of appearance, motion, and temporal smoothness.

[0026] Figure 3 An example system 300 according to this disclosure is shown. System 300 can be used for video generation using machine learning model 103. Machine learning model 103 may include a first sub-model 204, a second sub-model 212, and a third sub-model 215. Given one or more static images I s (Such as source image 301), system 300 can generate head animation sequence {I s →D i}, such as the head animation sequence depicted in output video 322, which has a length of q, to ​​drive the video. For example, drive video 308, where i = 0, ..., q represents the frame index.

[0027] Machine learning model 103 may receive source image 301 and driving video 308. Source image 301 may include an image of a first object (e.g., an identity including appearance features of the first object, such as facial features), while driving video 308 may include at least one image of a second object different from the first object. Driving video 308 may depict pose (e.g., head pose) or appearance (e.g., facial features). Source image 301 may be input into a first sub-model 204. First sub-model 204 may extract identity features 306 of the first object from source image 301. At least one mask image 310 may be generated based on driving video 308. For example, at least one mask image 310 may be generated for each frame of driving video 308.

[0028] Multiple mask images 310 may include at least one of the mouth or eye regions of the second object in the driving video 308. The driving video 308 and the multiple mask images 310 may be input into a second sub-model 212. The second sub-model 212 may derive motion information 314, such as information indicating pose or appearance, based on the driving video 308 and the multiple mask images 310. The machine learning model 103 may generate an output video 322 based on identity features 306 and motion information 314. The output video 322 may retain the identity features and background content of the first object depicted in the source image 301, and the output video 322 may follow the pose or appearance depicted in the driving video 308. To generate the output video 322, the machine learning model 103 may utilize one or more latent diffusion models with decoupled control over appearance, motion, and temporal smoothing. The multiple latent diffusion models may include a generative model designed to denoise Gaussian noise z through T denoising steps. T ~N(0,1) synthesize the desired data samples. The latent diffusion model can operate in the latent space determined by the pre-trained autoencoder.

[0029] To control facial appearance and head pose based on an image diffusion model, traditional methods typically employ ControlNet, trained to generate images based on facial keypoints. The control module can be trained to reconstruct images from the target image. D The extracted key points are conditional inputs and are defined by I. s The ID serves as input to the appearance reference module R. During training, the I representing the same object... s and I DIt can be two random video frames. While effective at coarse scales, such a control scheme introduces several problems, especially when zooming in on faces. First, the accuracy of the drive signal heavily depends on the accuracy of the third-party detector. This dependence introduces jitter control, motion blur, and can lead to poor animation when detection fails (e.g., due to facial occlusion). Second, conveying strong emotions or subtle expressions often involves detailed facial movements, such as those of the teeth, eyeballs, eyebrows, and brow (ajna). Coarse keypoint representation can significantly hinder animation performance, failing to capture the nuances required for accurate facial animation. Finally, the drive keypoints are compared with the target image I. D Facial structure alignment, target image I D Characterization and I S The objects in the traditional approach are identical to the objects in the self-driven training scheme. Therefore, as a simple approach, the traditional scheme tends to replicate the driving structure coupled with identity features (such as facial shape and proportions). As a result, during cross-identity animation inference, the traditional approach leads to undesirable identity drift towards the driving object.

[0030] To address the aforementioned issues, machine learning model 103 includes a second sub-model 212 (e.g., a control sub-model C). The second sub-model 212 may include novel conditional motion control completely decoupled from source identity features, while minimizing the loss of motion information at all scales (such as facial expressions and head pose). Representation and I S The original driving RGB image of the different objects represented D This can be used as a conditional input to the second sub-model 212. This allows the source image to be directly reproduced onto the driving video of different identities (different objects). However, it is difficult to obtain such image pairs with different identities but aligned motions for training.

[0031] Figure 4 An example system 400 according to this disclosure is shown. System 400 can be used to train a machine learning model 103 including a second sub-model 212. The second sub-model 212 can be trained by applying a cross-identity training scheme. The cross-identity training scheme can be configured to instruct the second sub-model 212 to derive an identity-decoupled pose or appearance.

[0032] Applying a cross-identity training scheme may include generating cross-identity image pairs. Each cross-identity image pair may include images of different objects. The generation of cross-identity image pairs can be supported by a pre-trained image reconstruction network F. Two randomly selected video frames representing the same object (e.g., appearance reference image 406 and reconstructed target image 404) may be selected. The pre-trained image reconstruction network F can generate an RGB control image 408 as a conditional input to the second sub-model 212, instead of relying on facial keypoints from the reconstructed target image 404. The control image 408 can be generated based on the cross-identity source image 402 and the reconstructed target image 404, where the cross-identity source image 402 is a frame randomly selected from videos with different identities. The cross-identity source image 402 may depict an object different from the objects in the appearance reference image 406 and the reconstructed target image 404. The control image 408 may depict the same object as the cross-identity source image 402. The control image 408 may share motion information with the reconstructed target image 404. The second sub-model 212 can be trained on the cross-identity image pairs to mitigate appearance leakage originating from driving signals. In the example, each cross-identity image pair may include an appearance reference image 406, a reconstructed target image 404, and a control image 408.

[0033] This cross-identity training scheme effectively instructs the second sub-model 212 to implicitly derive identity-decoupled motion from the control image 408. This mitigates appearance leakage from the driving signals, allowing direct application of the driving video for inference without third-party dependencies. The pre-trained image reconstruction network F provides reconstructed control images 408 with reasonable quality and motion accuracy for a wide range of conversational scenarios. Even with limited perceptual quality, the control image 408 contains richer motion information than keypoints, sufficient to enable the second sub-model 212 to effectively decipher the embedded motion structure, allowing it to adapt and correlate more refined expressions and poses when provided with ground truth motion for supervision. Thus, the second sub-model 212 is able to establish an implicit structural mapping between the control image 408 and the reconstructed target image 404, generalizing well to unseen expressions and head movements.

[0034] The trained second sub-model 212 provides significant improvements in capturing coarse keypoints in head transformations and low-frequency facial expressions. The trained second sub-model 212 can extract structural features from the control image 408 and can be integrated into the U-Net via skip connections during the denoising process. However, this additional conditional attention operates in the global image space, treating motion in each pixel with equal weight.

[0035] To guide the second sub-model 212 with enhanced local attention to key facial regions, aiming for better animation realism and finer control granularity, an auxiliary ControlNet is introduced. Motion control at subtle scales can be achieved using an auxiliary ControlNet conditioned on a local control image 410 that reveals only small patches around the eyes and mouth from control image 408. Specifically, keypoints in control image 408 can be detected for the eyes and mouth, and the centers of these keypoints can be used to crop a 128×128 patch as the local control image 410. This control branch effectively provides enhanced guidance to UNet denoising, focusing only on the local structures extracted from those cropped facial regions. The enhanced generator helps capture subtle motions in the hierarchically conditioned inputs (control image 408 and local control image 410), benefiting subsequent training of both control modules.

[0036] The first sub-model 204 (e.g., appearance reference module R) ensures the preservation of source identity features. The first sub-model 204 can derive appearance features from the appearance reference image 406, which can then be concatenated into the UNet transform block. Meanwhile, the cross-identity training scheme with the reconstructed control image 408 substantially mitigates appearance leakage originating from the driving signal. However, inherited from its self-supervised training, the pre-trained image reconstruction generator F is not entirely free of appearance entanglement. Therefore, facial attributes of the control image 408, particularly in terms of facial shape and eye / mouth size, can be impaired by the reconstructed target image 404, leading to slight identity drift, especially when there are significant differences in facial appearance between the source and driving images.

[0037] To mitigate these slight identity drifts, control image 408 and local control image 410 can be adjusted (e.g., scaled) during training using random heterogeneous scaling. This can induce slight facial distortion and structural misalignment between control image 408 / local control image 410 and the reconstructed target image 404, causing the network to rely on appearance reference image 406 for identity features. The scaling operation only affects head shape and does not modify driving facial expressions or head pose. While the induced excessive misalignment may hinder the learning of the control module, a random scaling factor within the range [0.9, 1.1] strikes a balance between identity preservation and motion performance. Furthermore, during cross-identity driving inference, facial shape differences can be minimized by applying affine transformations (translation and scaling) across the entire driving sequence to align the head bounding box of the source frame with the selected driving frame.

[0038] With a single appearance reference image 406, only a portion of the facial appearance is visible, and when the head pose or camera view changes, the network must rely on the general generative prior of the latent diffusion model (LDM) to fix unobservable facial regions. However, when more reference images are available, such as in videos, a more comprehensive appearance context can be incorporated without any network modifications. Due to the decoupled control described in this paper, the framework described can seamlessly fuse multiple extracted appearance features with the first sub-model 204 (e.g., appearance reference module R) into the UNet by simply concatenating them, generating animations with better-preserved identity attributes.

[0039] Figure 5 An example process 500 for generating images using a machine learning model is shown. Although in Figure 5 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0040] At point 502, a source image (e.g., source image 201) and a driving image (e.g., driving image / video 202) can be received. The source image and driving image can be received by a machine learning model (e.g., machine learning model 103). The machine learning model may include a first sub-model (e.g., first sub-model 204) and a second sub-model (e.g., second sub-model 212). The source image may include an image of a first object. The driving image may include a second object different from the first object. The driving image may depict a pose or appearance.

[0041] A source image can be input into a first sub-model. At 504, the identity features (appearance features, such as facial features) of the first object can be extracted from the source image through the first sub-model. At 506, at least one mask image (e.g., mask image 210) can be generated based on the driving image. The mask images(s) may include at least one of the mouth region or eye region in the driving image. The driving image and the mask images(s) can be input into a second sub-model. At 508, motion information (such as information indicating pose or appearance) can be derived from the driving image and the mask images(s) through the second sub-model. At 510, an image (e.g., output image / video 122) can be generated by a machine learning model. The generated image can retain the identity features of the first object and follow the pose or appearance depicted in the driving image.

[0042] Figure 6 An example process 600 is shown for training a second sub-model (e.g., second sub-model 212) of a machine learning model (e.g., machine learning model 103) according to this disclosure. Although in Figure 6The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0043] At 602, cross-identity image pairs can be generated. Each cross-identity image pair can include images of different objects. The generation of cross-identity image pairs can be supported by a pre-trained image reproduction network F. At 604, a second sub-model of the machine learning model can be trained on the cross-identity image pairs by applying a cross-identity training scheme. The cross-identity training scheme can be configured to instruct the second sub-model to derive identity-decoupled motion information. The second sub-model can be trained on the cross-identity image pairs to mitigate appearance leakage originating from driving signals.

[0044] Figure 7 An example process 700 for generating a control image according to this disclosure is shown. Although in Figure 7 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0045] Cross-identity image pairs can be generated. Each cross-identity image pair may include images of different objects. The generation of cross-identity image pairs can be supported by a pre-trained image reconstruction network F. At 702, two video frames representing the same object (e.g., appearance reference image 406 and reconstructed target image 404) can be selected. At 704, an RGB control image (e.g., control image 408) can be generated by the pre-trained image reconstruction network. The control image can be generated based on the cross-identity source image (e.g., cross-identity source image 402) and the reconstructed target image. The control image can represent an object different from the object in the appearance reference image and the reconstructed target image. The control image can share motion information with the reconstructed target image.

[0046] Figure 8 An example process 800 for training a second sub-model (e.g., second sub-model 212) according to this disclosure is shown. Although in Figure 8 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0047] Cross-identity image pairs can be generated. Each cross-identity image pair may include images of different objects. The generation of cross-identity image pairs can be supported by a pre-trained image representation network F. RGB control images (e.g., control image 408) can be generated by the pre-trained image representation network. At 802, local control images can be generated. Local control images can be generated based on control images in the cross-identity image pairs. Each local control image may include at least one of a mouth region or multiple eye regions. A second sub-model (e.g., second sub-model 212) of a machine learning model (e.g., machine learning model 103) can be trained on the cross-identity image pairs by applying a cross-identity training scheme. At 804, the second sub-model of the machine learning model can be guided to enhance attention to local facial movements using local control images. For example, the second sub-model can be instructed to derive identity-decoupled motion information. The second sub-model can be trained on the cross-identity image pairs to mitigate appearance leakage originating from driving signals.

[0048] Figure 9 An example process 900 for training a machine learning model (e.g., machine learning model 103) according to this disclosure is shown. Although in Figure 9 The operations are depicted as a sequence of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder, or modify the depicted operations.

[0049] Two video frames representing the same object (e.g., appearance reference image 406 and reconstructed target image 404) can be selected. At 902, an RGB control image (e.g., control image 408) can be generated by a pre-trained image reconstruction network. The control image can be generated based on a cross-identity source image (e.g., cross-identity source image 402) and the reconstructed target image. The control image can represent an object different from the object in the appearance reference image and the reconstructed target image. The control image can share motion information with the reconstructed target image.

[0050] At 904, local control images can be generated. Local control images can be generated based on control images across a pair of identity images. Each local control image may include at least one of a mouth region or one of several eye regions. At 906, a randomized heterogeneous scaling operation can be performed on the control images and local control images. The randomized heterogeneous scaling operation can be performed on the control images and local control images during training to enable a machine learning model to derive identity features from appearance reference images. In an embodiment, the random scaling factor of the randomized heterogeneous scaling operation can be greater than or equal to 0.9 and less than or equal to 1.1.

[0051] Figure 10An example process 1000 for generating video using a machine learning model (e.g., machine learning model 103) according to this disclosure is shown. Although in Figure 10 The operations are described as a series of operations, but those skilled in the art will understand that various embodiments may add, remove, reorder or modify the operations described.

[0052] At 1002, a source image (e.g., source image 201) and a driving video (e.g., driving video 202) can be received. The source image and driving video can be received by a machine learning model (e.g., machine learning model 103). The source image may include an image of a first object. The driving video may include a second object different from the first object. The driving video may include a sequence of frames. The driving video may characterize the second object with motions associated with a head or face. The first object is different from the second object. At 1004, video (e.g., output video 122) can be generated by the machine learning model. The generated video may retain the identity features of the first object and may follow the motions depicted in the driving video.

[0053] During inference, a cue word walk strategy is employed to enhance temporal smoothness. Using a Stable Diffusion UNet (SD UNet), machine learning model 103 demonstrates inherent compatibility with the latent consistency model. Notably, instead of denoising from random Gaussian noise, a forward diffusion process is applied to the source image as initial noise. This generated noise adds a subtle level of structural guidance in the early denoising step, resulting in improved consistency along with reduced flicker artifacts. The pre-trained image reconstruction network F is not utilized during inference.

[0054] Machine learning model 103 can create vivid and highly expressive animations, demonstrating a wide range of head movements (with rotations exceeding 150 degrees) and facial expressions (frowning, squinting, pouting, etc.) across realistic, humanoid, and stylized characters. Machine learning model 103 employs reference modules to effectively establish local spatial correspondences between input and output across query source appearance features. Once trained, machine learning model 103 is able to generalize to out-of-domain appearances, such as those embodied in stylized characters, through its learned latent space. Simultaneously, it maintains high identity similarity to the given source image throughout the generated video.

[0055] In summary, Machine Learning Model 103 ensures deterministic transfer of driving facial expressions and head pose. Model 103 excels at incorporating cross-identity driven inputs during training, supporting a balance between motion expressiveness, identity preservation, and animation robustness. The local control module emphasizes attention to detailed facial expressions, capturing subtle yet crucial elements for emotional delivery. The excellent performance of Model 103 on generalized source images and driving motion validates its effectiveness.

[0056] Figure 11 The diagram illustrates computing devices that can be used in various aspects, such as Figures 1 to 4 The services, networks, modules, and / or devices described in any of them. About Figures 1 to 4 Any or all components can be free. Figure 11 One or more instances of the computing device 1100 are implemented. Figure 11 The computer architecture shown illustrates a conventional server computer, workstation, desktop computer, laptop computer, tablet computer, network device, personal digital assistant (PDA), e-reader, digital cellular phone, or other computing node, and can be used to perform any aspect of the computer described herein, such as to implement the methods described herein.

[0057] The computing device 1100 may include a substrate or “motherboard,” which is a printed circuit board to which multiple components or devices may be connected via a system bus or other electrical communication path. One or more central processing units (CPUs) 1104 may operate in conjunction with a chipset 1106. The CPU 1104 may be a standard programmable processor that performs the arithmetic and logic operations necessary for the operation of the computing device 1100.

[0058] The CPU 1104 can perform necessary operations by manipulating switching elements that distinguish and change these states to transition from one discrete physical state to the next. Switching elements typically include electronic circuitry that maintains one of two binary states, such as flip-flops, and electronic circuitry that provides an output state based on a logical combination of the states of one or more other switching elements, such as logic gates. These basic switching elements can be combined to create more complex logic circuits, including registers, adder-subtractor units, arithmetic logic units, floating-point units, and so on.

[0059] The (multiple) CPUs 1104 can be enhanced or replaced by other processing units (e.g., GPUs 1105). The (multiple) GPUs 1105 may include processing units dedicated to, but not limited to, highly parallel computing, such as graphics and other visualization-related processing.

[0060] Chipset 1106 provides an interface between CPU(s) 1104 and the remaining components and devices on the substrate. Chipset 1106 may also provide an interface to random access memory (RAM) 1108, which serves as the main memory in computing device 1100. Chipset 1106 may also provide an interface to computer-readable storage media, such as read-only memory (ROM) 1120 or non-volatile RAM (NVRAM) (not shown), for storing basic routines that help boot computing device 1100 and transfer information between various components and devices. ROM 1120 or NVRAM may also store other software components necessary for the operation of computing device 1100 according to the aspects described herein.

[0061] Computing device 1100 can operate in a networked environment using a logical connection via a local area network (LAN) to remote computing nodes and computer systems. Chipset 1106 may include functionality for providing network connectivity via a network interface controller (NIC) 1122 (such as a Gigabit Ethernet adapter). NIC 1122 enables computing device 1100 to connect to other computing nodes via network 1118. It should be understood that multiple NICs 1122 may exist in computing device 1100, connecting the computing device to other types of networks and remote computer systems.

[0062] Computing device 1100 can be connected to mass storage device 1128, which provides non-volatile storage for the computer. Mass storage device 1128 can store system programs, application programs, other program modules, and data, as described in more detail herein. Mass storage device 1128 can be connected to computing device 1100 via storage controller 1124, which is connected to chipset 1106. Mass storage device 1128 can consist of one or more physical storage units. Mass storage device 1128 may include management components. Storage controller 1124 can interface with physical storage units via a serial attached SCSI (SAS) interface, a serial advanced technology accessory (SATA) interface, a Fibre Channel (FC) interface, or other types of interfaces used for physical connection and data transfer between the computer and physical storage units.

[0063] The computing device 1100 can store data on the mass storage device 1128 by changing the physical state of the physical storage units to reflect the stored information. The specific changes in the physical state may depend on various factors and different embodiments of this specification. Examples of these factors may include, but are not limited to, the techniques used to implement the physical storage units and the techniques used to characterize the mass storage device 1128 as primary or secondary storage.

[0064] For example, computing device 1100 can store information in mass storage device 1128 by issuing instructions via storage controller 1124 to change the magnetic properties of a specific location within a disk drive unit, the reflection or refraction properties of a specific location in an optical storage unit, or the electrical properties of a specific capacitor, transistor, or other discrete component in a solid-state storage unit. Other transformations of the physical medium are possible without departing from the scope and spirit of this specification, and the examples provided for the foregoing are for illustrative purposes only. Computing device 1100 can also read information from mass storage device 1128 by detecting the physical state or characteristics of one or more specific locations within a physical storage unit.

[0065] In addition to the aforementioned high-capacity storage device 1128, the computing device 1100 may also access other computer-readable storage media to store and retrieve information, such as program modules, data structures, or other data. Those skilled in the art will understand that a computer-readable storage medium can be any available medium that provides storage for non-transitory data and can be accessed by the computing device 1100.

[0066] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, transient computer-readable storage media and non-transitory computer-readable storage media implemented in any method or technology, as well as removable and non-removable media. Computer-readable storage media include, but are not limited to, RAM, ROM, erasable programmable ROM (“EPROM”), electrically erasable programmable ROM (“EEPROM”), flash memory or other solid-state memory technologies, optical disc ROM (“CD-ROM”), digital versatile disc (“DVD”), high-definition DVD (“HD-DVD”), Blu-ray disc (BLU-RAY) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, other magnetic storage devices, or any other medium that may be used to store desired information in a non-transitory manner.

[0067] Massive storage devices (e.g.) Figure 11 The mass storage device 1128 depicted herein may store an operating system used to control the operation of the computing device 1100. The operating system may include a version of the LINUX operating system. The operating system may include a version of the WINDOWS SERVER operating system from Microsoft Corporation. According to another aspect, the operating system may include a version of the UNIX operating system. Various mobile phone operating systems, such as iOS and Android, may also be used. It should be understood that other operating systems may also be used. The mass storage device 1128 may store other systems, applications, and data utilized by the computing device 1100.

[0068] Mass storage device 1128 or other computer-readable storage medium may also be encoded with computer-executable instructions that, when loaded into computing device 1100, transform the computing device from a general-purpose computing system into a special-purpose computer capable of implementing the aspects described herein. These computer-executable instructions transform computing device 1100 by specifying how CPU(s)(s)1104 transition between states, as described above. Computing device 1100 can access a computer-readable storage medium storing computer-executable instructions that, when executed by computing device 1100, can perform the methods described herein.

[0069] Such as Figure 11 The computing device 1100 depicted may also include an input / output controller 1132 for receiving and processing input from multiple input devices, such as a keyboard, mouse, touchpad, touchscreen, electronic stylus, or other types of input devices. Similarly, the input / output controller 1132 may provide output to a display, such as a computer monitor, flat panel display, digital projector, printer, plotter, or other types of output device. It should be understood that the computing device 1100 may not include... Figure 11 All components shown may include Figure 11 Other components not explicitly shown in the document, or those that can be utilized with Figure 11 The architecture shown is completely different from the one shown.

[0070] As described in this article, a computing device can be a physical computing device, such as... Figure 11 The computing device 1100. A computing node may also include virtual machine host processes and one or more virtual machine instances. Computer-executable instructions may be executed indirectly by the physical hardware of the computing device through the interpretation and / or execution of instructions stored and executed in the context of a virtual machine.

[0071] It should be understood that the methods and systems are not limited to a particular method, a particular component, or a particular implementation. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0072] As used in the specification and appended claims, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include plural referents. A range may be expressed herein as “about” a particular value, and / or “about” another particular value. When such a range is expressed, another embodiment includes from one particular value and / or another particular value. Similarly, when a value is expressed as an approximation using the antecedent “about,” it will be understood that the particular value forms another embodiment. It should also be understood that the endpoints of each range are significant relative to and independent of the other endpoint.

[0073] "Optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes instances where the event or situation occurs as well as instances where it does not occur.

[0074] Throughout the description and claims of this specification, the word “comprising” and variations thereof, such as “comprising” and “including,” means “including, but not limited to,” and is not intended to exclude, for example, other components, integers, or steps. “Exemplary” means “example” and is not intended to convey indications of preferred or ideal embodiments. “Like” is used not in a limiting sense but for purposes of explanation.

[0075] Components used to perform the described methods and systems can be described. When describing combinations, subsets, interactions, groups, etc., of these components, it should be understood that while specific references to each of the various individual and collective combinations and substitutions of these components may not be explicitly described, they are particularly contemplated and described herein for all methods and systems. This applies to all aspects of this application, including but not limited to operations in the described methods. Therefore, if various additional operations are available, it should be understood that each of these additional operations can be performed using any particular embodiment or combination of embodiments of the described methods.

[0076] The method and system can be more readily understood by referring to the preferred embodiments and the examples included therein, as well as the following detailed description of the accompanying drawings and their descriptions.

[0077] As those skilled in the art will understand, the methods and systems may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the methods and systems may take the form of a computer program product on a computer-readable storage medium having computer-readable program instructions (e.g., computer software) implemented on the storage medium. More specifically, the methods and systems may take the form of network-implemented computer software. Any suitable computer-readable storage medium, including hard disks, CD-ROMs, optical storage devices, or magnetic storage devices, may be used.

[0078] Embodiments of the methods and systems are described below with reference to block diagrams and flowcharts illustrating the methods, systems, apparatus, and computer program products. It should be understood that each block in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented by computer program instructions. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute on the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart blocks.

[0079] These computer program instructions may also be stored in a computer-readable storage medium that can instruct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of writing including computer-readable instructions for implementing the functions specified in the flowchart block or block. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in the flowchart block or block.

[0080] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Furthermore, in some embodiments, certain method or process blocks may be omitted. The methods and processes described herein are not limited to any particular sequence, and the associated blocks or states may be executed in other suitable sequences. For example, described blocks or states may be executed in a different order than specifically described, or multiple blocks or states may be combined in a single block or state. Example blocks or states may be executed serially, in parallel, or in some other manner. Blocks or states may be added to or removed from the described example embodiments. The example systems and components described herein may be configured differently from those described. For example, elements may be added, removed, or rearranged compared to the described example embodiments.

[0081] It should also be understood that various items are shown as being stored in memory or, when in use, in a storage device, and these items, or portions thereof, may be transferred between memory and other storage devices for memory management and data integrity purposes. Alternatively, in other embodiments, some or all of the software modules and / or systems may be executed in memory on another device and communicate with the illustrated computing system via inter-computer communication. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways, such as at least in part in firmware and / or hardware, including, but not limited to, one or more application-specific integrated circuits (“ASICs”), standard integrated circuits, controllers (e.g., by executing appropriate instructions, and including microcontrollers and / or embedded controllers), field-programmable gate arrays (“FPGAs”), complex programmable logic devices (“CPLDs”), etc. Some or all of the modules, systems, and data structures may also be stored (e.g., as software instructions or structured data) on computer-readable media, such as hard disks, memory, networks, or portable media articles that will be read by appropriate devices or via appropriate connections. Systems, modules, and data structures can also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) over various computer-readable transmission media, including wireless and wired / cable-based media, and can take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). In other embodiments, such computer program products can also take other forms. Therefore, the invention can be practiced with other computer system configurations.

[0082] While methods and systems have been described in conjunction with preferred embodiments and specific examples, they are not intended to limit the scope to the particular embodiments illustrated, as the embodiments herein are intended in all respects to be illustrative rather than restrictive.

[0083] Unless otherwise expressly stated, any method described herein should never be construed as requiring its operations to be performed in a particular order. Therefore, no inference of order is intended in any respect where a method claim does not actually describe the order of its operations or where the claims or description do not specifically state that the operations are limited to a particular order. This preserves any possible unexpressed basis for interpretation, including: logical questions regarding the arrangement of steps or operational flows; general meanings derived from grammatical organization or punctuation; and the number or type of embodiments described in the specification.

[0084] It will be apparent to those skilled in the art that various modifications and variations can be made without departing from the scope or spirit of this disclosure. Other embodiments will be apparent to those skilled in the art upon consideration of the description and practice described herein. This description and the example drawings are to be considered exemplary only, and the actual scope and spirit are indicated by the appended claims.

Claims

1. A method for generating images using a machine learning model, comprising: The machine learning model receives a source image and a driving image, wherein the source image includes an image of a first object, the driving image includes a second object different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model. The appearance features of the first object are extracted from the source image using the first sub-model; A mask image is generated based on the driving image, wherein the mask image includes at least one of the mouth region or eye region in the driving image; The pose or appearance is derived from the driving image and the mask image using the second sub-model. as well as The machine learning model generates an image that retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

2. The method of claim 1, wherein the second sub-model is trained by applying a cross-identity training scheme, and the cross-identity training scheme is configured to instruct the second sub-model to derive an identity-decoupled pose or appearance.

3. The method according to claim 2, wherein the application cross-identity training scheme comprises: Generate cross-identity image pairs, each of which includes a different object; as well as The second sub-model is trained on the cross-identity image pair to mitigate appearance leakage originating from the driving signal.

4. The method of claim 3, wherein generating each cross-identity image pair comprises: Select a reference image and a target image to represent the appearance of the same object; as well as A control image is generated by a pre-trained image reconstruction generator, wherein the control image represents an object that is different from the object in the appearance reference image and the reconstructed target image, and wherein the control image shares motion information with the reconstructed target image.

5. The method according to claim 4, further comprising: A local control image is generated based on the control image in the cross-identity image pair, wherein each local control image includes at least one of a mouth region or an eye region; as well as The second sub-model is guided by the local control image to enhance attention to local facial movements.

6. The method according to claim 4, further comprising: During training, random heterogeneous scaling is performed on the control image and local control images to enable the machine learning model to derive appearance features from the appearance reference image.

7. The method of claim 6, wherein the random scaling factor of the random heterogeneous scaling operation is greater than or equal to 0.9 and less than or equal to 1.

1.

8. The method of claim 1, wherein the pose includes a head pose and the appearance includes a facial appearance.

9. The method according to claim 1, further comprising: The source image and driving video are received through the machine learning model, wherein the driving video comprises a sequence of frames and the driving video uses motion associated with a head or face to characterize the second object; as well as The machine learning model generates a video that retains the appearance features of the first object and follows the motion depicted in the driving video.

10. A system for generating images using a machine learning model, comprising: At least one processor; as well as At least one memory, communicatively coupled to the at least one processor, and including computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: The machine learning model receives a source image and a driving image, wherein the source image includes an image of a first object, the driving image includes a second object different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model. The appearance features of the first object are extracted from the source image using the first sub-model; A mask image is generated based on the driving image, wherein the mask image includes at least one of the mouth region or eye region in the driving image; The pose or appearance is derived from the driving image and the mask image using the second sub-model; and The machine learning model generates an image that retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

11. The system of claim 10, wherein the second sub-model is trained by applying a cross-identity training scheme, wherein the cross-identity training scheme is configured to instruct the second sub-model to derive an identity-decoupled pose or appearance, and wherein applying the cross-identity training scheme comprises: Generate cross-identity image pairs, each of which includes a different object; as well as The second sub-model is trained on the cross-identity image pair to mitigate appearance leakage originating from the driving signal.

12. The system of claim 11, wherein generating each cross-identity image pair comprises: Select a reference image and a target image to represent the appearance of the same object; as well as A control image is generated by a pre-trained image reconstruction generator, wherein the control image represents an object that is different from the object in the appearance reference image and the reconstructed target image, and wherein the control image shares motion information with the reconstructed target image.

13. The system according to claim 12, further comprising: A local control image is generated based on the control image in the cross-identity image pair, wherein each local control image includes at least one of a mouth region or an eye region; as well as The second sub-model is guided by the local control image to enhance attention to local facial movements.

14. The system of claim 12, further comprising: During training, random heterogeneous scaling is performed on the control image and local control images to enable the machine learning model to derive appearance features from the appearance reference image.

15. The system according to claim 10, further comprising: The source image and driving video are received through the machine learning model, wherein the driving video comprises a sequence of frames and the driving video uses motion associated with a head or face to characterize the second object; as well as The machine learning model generates a video that retains the appearance features of the first object and follows the motion depicted in the driving video.

16. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a processor, cause the processor to perform operations, the operations including: A source image and a driving image are received through a machine learning model, wherein the source image includes an image of a first object, the driving image includes a second object that is different from the first object, the driving image depicts a pose or appearance, and the machine learning model includes a first sub-model and a second sub-model. The appearance features of the first object are extracted from the source image using the first sub-model; A mask image is generated based on the driving image, wherein the mask image includes at least one of the mouth region or eye region in the driving image; The pose or appearance is derived from the driving image and the mask image using the second sub-model. as well as The machine learning model generates an image that retains the appearance features of the first object and follows the pose or appearance depicted in the driving image.

17. The non-transitory computer-readable storage medium of claim 16, wherein the second sub-model is trained by applying a cross-identity training scheme, wherein the cross-identity training scheme is configured to instruct the second sub-model to derive an identity-decoupled pose or appearance, and wherein applying the cross-identity training scheme comprises: Generate cross-identity image pairs, each of which includes a different object; as well as The second sub-model is trained on the cross-identity image pair to mitigate appearance leakage originating from the driving signal.

18. The non-transitory computer-readable storage medium of claim 17, wherein generating each cross-identity image pair comprises: Select a reference image and a target image to represent the appearance of the same object; as well as A control image is generated by a pre-trained image reconstruction generator, wherein the control image represents an object that is different from the object in the appearance reference image and the reconstructed target image, and wherein the control image shares motion information with the reconstructed target image.

19. The non-transitory computer-readable storage medium of claim 18, further comprising: A local control image is generated based on the control image in the cross-identity image pair, wherein each local control image includes at least one of a mouth region or an eye region; as well as The second sub-model is guided by the local control image to enhance attention to local facial movements.

20. The non-transitory computer-readable storage medium of claim 18, further comprising: During training, random heterogeneous scaling is performed on the control image and local control images to enable the machine learning model to derive appearance features from the appearance reference image.