Diffusion model-based realistic makeup migration method for three-dimensional animatable head portrait

Through the three-dimensional animated avatar realistic makeup migration method based on the diffusion model, the problems of inconsistent makeup effects and identity characteristics maintenance during the three-dimensional avatar makeup process are solved, and high-quality makeup effects and real-time rendering are improved.

CN120339472AActive Publication Date: 2025-07-18SHANDONG UNIV OF SCI & TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510382426.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing three-dimensional avatar technology is difficult to achieve high-quality makeup effects during the makeup process. The makeup effects are inconsistent with dynamic expressions and multi-view angles, making it difficult to maintain the identity characteristics of the avatar, and the real-time rendering efficiency is low.

Method used

A three-dimensional animated avatar realistic makeup migration method based on diffusion model is adopted to generate makeup images through a stable makeup model, optimize global UV maps, and optimize makeup with refinement mechanism and loss function to generate high-quality makeup effects.

Benefits of technology

It achieves consistency of makeup effects in dynamic expressions and multi-view angles, maintains the identity characteristics of the avatar, improves the naturalness and aesthetics of the makeup effects, and improves real-time rendering efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339472A_ABST
    Figure CN120339472A_ABST
Patent Text Reader

Abstract

The invention provides a realistic makeup migration method of a three-dimensional animatable head portrait based on a diffusion model. The realistic makeup migration method specifically comprises the following steps: acquiring a reconstructed three-dimensional drivable target figure head portrait and a single-view reference picture with a target makeup effect as input of a stable makeup model; generating a series of makeup images by adopting a stable makeup model, optimizing a global UV chartlet by utilizing the makeup images, and sampling from the global UV chartlet to obtain a makeup image as supervision information; adding a refining mechanism to the stable makeup model, and further optimizing makeup details of the head portrait of the target person through the stable makeup model; and optimizing the makeup of the head portrait of the target person in combination with the loss function, and finally generating the head portrait of the target person with a high-quality makeup effect. According to the technical scheme, the problems that in the prior art, the high-quality make-up effect cannot be achieved for the three-dimensional head portrait, and the identity characteristics of the head portrait cannot be kept in the make-up process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional avatar construction, and particularly to a realistic makeup transfer method for three-dimensional animatable avatars based on a diffusion model. Background Art

[0002] Three-dimensional avatars are increasingly widely used in fields such as virtual reality, gaming, film and television production, and social media. Three-dimensional avatars can not only provide more realistic visual effects but also enhance the user's immersion through dynamic expressions and movements. However, to meet the user's needs for personalization and aesthetics, the customization of three-dimensional avatars has become an important research direction. Among them, the realization of makeup effects is one of the key factors in enhancing the visual appeal of three-dimensional avatars.

[0003] Although significant progress has been made in makeup techniques for two-dimensional images, such as makeup transfer methods achieved through generative adversarial networks (GANs) and stable makeup models, the application of these techniques to three-dimensional avatars still faces many challenges: First is the dynamic consistency problem: Existing three-dimensional makeup techniques often result in inconsistent makeup effects when dealing with dynamic expressions and multiple perspectives. For example, when the avatar makes exaggerated expressions or is observed from different angles, makeup details may be distorted, blurred, or lost. This inconsistency seriously affects the realism and user experience of three-dimensional avatars. Then is the lack of detail control: In the process of applying makeup to three-dimensional avatars, precise control of makeup details is a key issue. Existing methods are difficult to achieve fine adjustment of makeup details (such as eyeshadow, lip color, blush, etc.) while maintaining the identity characteristics of the avatar. This leads to the generated makeup effects often lacking naturalness and artistic sense. Next is the identity preservation problem: The purpose of makeup transfer is to change the appearance of the avatar while maintaining its original identity characteristics. However, existing three-dimensional makeup techniques perform poorly in this regard. For example, some methods may change the facial contour, skin color, or facial features of the avatar, resulting in a deviation in identity between the generated avatar and the original avatar. Finally is the real-time and efficiency problem: Most existing three-dimensional makeup techniques rely on complex optimization processes and are difficult to achieve real-time rendering. This limits their use in real-time interactive applications (such as virtual reality and gaming). For example, some methods based on neural radiance fields (NeRF) can generate high-quality three-dimensional effects but require a large amount of computational resources and time during the optimization process.

[0004] Therefore, it is necessary to develop a realistic makeup transfer method for three-dimensional animatable avatars that can achieve high-quality makeup effects and maintain the identity characteristics of the avatar during the makeup process. Summary of the Invention

[0005] The main object of the present invention is to provide a method for realistic makeup transfer of 3D animatable avatars based on diffusion models, so as to solve the problems in the prior art that high-quality makeup effects cannot be achieved for 3D avatars and the identity characteristics of avatars cannot be maintained during the makeup process.

[0006] To achieve the above object, the present invention provides a method for realistic makeup transfer of 3D animatable avatars based on diffusion models, specifically including the following steps:

[0007] S1. Obtain a reconstructed 3D drivable target person's head and a single-view reference photo with the target makeup effect as the input of the stable makeup model.

[0008] S2. Use the stable makeup model to generate a series of makeup images, optimize the global UV map using the makeup images, and then sample from the global UV map to obtain the makeup pictures as supervision information, thereby optimizing the makeup of the target person's head.

[0009] S3. Add a refinement mechanism to the stable makeup model, and further optimize the makeup details of the target person's head through the stable makeup model on the basis of the rough makeup in step S2.

[0010] S4. Optimize the makeup of the target person's head in combination with the loss function, and finally generate the target person's head with high-quality makeup effects.

[0011] Further, step S1 specifically includes the following steps:

[0012] S1.1. Use existing 3D head reconstruction technologies, such as Gaussian character models, to generate a 3D drivable target person's head.

[0013] S1.2. The user provides a single reference picture R with the target makeup effect.

[0014] Further, step S2 specifically includes the following steps:

[0015] S2.1. Given k camera views {π1, π2,..., π k}, render the pictures {x1, x2,..., x k} corresponding to the target person's head from the corresponding views.

[0016] S2.2. Add noise to the rendered image x i for t time steps to obtain Then, through the stable makeup model θ, perform a denoising process on the noisy picture to generate the makeup transfer images k under k camera views {π1, π2,..., π

[0017]

[0018] S2.3. Render the mapping image of each point in the makeup transfer image to the UV coordinates through the mesh renderer. Optimize the global UV map by accumulating the makeup transfer images from multiple perspectives. The optimization formula for the UV map is as follows:

[0019]

[0020] where (H, W) and (h, w) are pixel coordinates, is the pixel value at (H, W), and UVmap(h, w) is the pixel value of the UV map.

[0021] S2.4. Using the global UV map, render the makeup images of the target person's head at any expression and perspective through the mesh renderer. Use step S4 and take the makeup image as the supervision image to optimize the makeup of the target person's head.

[0022] Furthermore, step S3 specifically includes the following steps:

[0023] S3.1. Given the target person's head obtained in step S2, render a series of images {y1, y2,..., y n}.

[0024] S3.2. When adding noise to {y1, y2,..., y n}, select the time step t ′ ∈[20, 350] to obtain

[0025] S3.3. Input into the stable makeup model for denoising. Similar to step S2.2, obtain the denoised image

[0026] S3.4. Use as the supervision image to optimize the target person's head, and use step S4 to obtain the optimized makeup of the target person's head.

[0027] Furthermore, step S4 specifically includes the following steps:

[0028] S4.1. For the areas that need makeup transfer in the supervision images of both step S2 and step S3, use the L1 loss and L LPIPS loss for optimization:

[0029] L makeup = L1(M⊙I G , M⊙I r ) + LLPIPS (M⊙I G ,M⊙I r );

[0030] Among them, M represents the area where makeup transfer is required, and I G represents the supervision picture, and I r represents the rendering picture, and ⊙ is the element multiplication.

[0031] S4.2. For the areas where makeup transfer is not required, use the L1 loss to maintain the same texture features as the original character:

[0032] L Res = L1((1 - M)⊙I ID ,(1 - M)⊙I r );

[0033] Among them, I ID represents the rendering picture of the initially reconstructed target person's head portrait.

[0034] The present invention has the following beneficial effects:

[0035] By combining a stable makeup model and an optimization technique, the present invention achieves a high-quality 3D head makeup effect. The makeup effect is maintained consistently under dynamic expressions and multiple perspectives, solving the problems in the prior art. By adding a refinement mechanism to the stable makeup model, precise control of makeup details is achieved, improving the naturalness and aesthetics of the makeup effect. The present invention maintains the identity characteristics of the head during the makeup process, avoiding damage to the original features of the head by the makeup effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0037] Figure 1 Shows a flowchart of a realistic makeup transfer method for a 3D animatable head based on a diffusion model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The following will clearly and completely describe the technical solutions of the present invention with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0039] As Figure 1 shown, a realistic makeup transfer method for 3D animatable avatars based on diffusion models specifically includes the following steps:

[0040] S1. Obtain a reconstructed 3D drivable target human head and a single-view reference photo with the target makeup effect as the input for the stable makeup model.

[0041] S2. Coarse makeup stage: Use the stable makeup model to generate a series of makeup images, optimize the global UV map using the makeup images, and then sample from the global UV map to obtain the makeup pictures as supervision information, thereby optimizing the makeup of the target human head.

[0042] S3. Fine makeup stage: Add a refinement mechanism to the stable makeup model. Based on the coarse makeup in step S2, further optimize the makeup details of the target human head in the stable makeup model. Further enhance the details of the makeup effect while maintaining consistency with the coarse makeup stage.

[0043] S4. Optimize the makeup of the target human head in combination with the loss function, and finally generate a target human head with a high-quality makeup effect. By optimizing the features and opacity attributes of the 3D Gaussian distribution, and combining the loss function (including L1 loss, L LPIPS loss and constraint loss) to optimize the head, finally generate a 3D head with a high-quality makeup effect.

[0044] Specifically, step S1 specifically includes the following steps:

[0045] S1.1. Use existing 3D head reconstruction technologies, such as Gaussian character models, to generate a 3D drivable target human head. "Drivable" means capable of performing expression transformations.

[0046] S1.2. The user provides a single reference picture R with the target makeup effect.

[0047] Specifically, step S2 specifically includes the following steps:

[0048] S2.1. Given k camera viewpoints {π1, π2, …, π k}, render pictures {x1, x2, …, x k} of the target human head from the corresponding viewpoints. Select 16 camera viewpoints {π1, π2,..., π 16} to cover the entire face, and use a Gaussian Renderer to render pictures {x1, x2,..., x 16} of the target human head from the corresponding viewpoints.

[0049] S2.2, Perform the diffusion process. Add noise to the rendered image x i for t time steps to obtain Then, through the stable makeup model θ, perform the denoising process on the noisy image to generate the makeup transfer images under k camera views {π1, π2,..., π l}.

[0050]

[0051] Add noise to the rendered image x i for y = 1000 time steps to obtain Then, through the pre-trained stable makeup model θ (Stable-Makeup) based on the diffusion model, perform the denoising process on the noisy image. The denoising process uses the denoising diffusion implicit model DDIM for iterative denoising. Finally, generate the makeup transfer images under 16 camera views

[0052] S2.3, Render the mapping image of each point in the makeup transfer image to the UV coordinates through the mesh renderer Optimize the global UV map by accumulating the makeup transfer images from multiple views. The optimization formula for the UV map is as follows:

[0053]

[0054] where (H, W) and (h, w) are pixel coordinates, is the pixel value at (H, W), and UVmap(h, w) is the pixel value of the UV map.

[0055] Optimize the global UV map by accumulating the makeup images from 16 views to ensure its consistency under different views and expressions. First, through the mesh renderer, specifically Nvdiffrast, render the mapping image of pixels to UV coordinates.

[0056] S2.4, Using the global UV map, render the makeup images of the target person's head at any expression and view through the mesh renderer. Use the images in step S4 and take the makeup images as the supervision images to optimize the makeup of the target person's head. The images rendered when optimizing the makeup of the target person's head are used as the rendered images in this step.

[0057] Specifically, step S3 specifically includes the following steps:

[0058] S3.1, Given the target person's head obtained after step S2, render a series of images {y1, y2,..., y n}, as the rendered image for this step. At this time, the target person's head has a rough but highly consistent makeup in terms of various angles and expressions. A series of 1000 images with different perspectives and expressions are rendered for the target person's head.

[0059] S3.2, To further improve the makeup quality, a refinement mechanism is added to the stable makeup model, that is, for {y1, y2,..., y n} when adding noise, the time step t ′ ∈ [20, 350], to obtain

[0060] S3.3, Input into the stable makeup model for denoising. Similar to step S2.2, the denoised image

[0061] S3.4, Use as the supervision image to optimize the target person's head, and use step S4 to obtain a high-texture-quality makeup.

[0062] Specifically, step S4 includes the following steps:

[0063] S4.1, For the supervision images in the two stages of step S2 and step S3, the regions that need makeup transfer are optimized using both L1 loss and L lPIPS loss:

[0064] L makeup = L1(M ⊙ I G , M ⊙ I r ) + L LPIPS (M ⊙ I G , M ⊙ I r );

[0065] Among them, M is the mask obtained from the face-parsing model, representing the regions that need makeup transfer, I G represents the supervision image, I r represents the rendered image, and ⊙ is element-wise multiplication.

[0066] S4.2, For the regions that do not need makeup transfer, the L1 loss is used to maintain the same texture features as the original character:

[0067] L Res = L1((1 - M) ⊙ I ID , (1 - M) ⊙ I r );

[0068] Among them, I ID represents the rendered image of the initially reconstructed target person's head.

[0069] The present invention is implemented under the PyTorch framework, and the Adam algorithm is used to optimize the model, with a maximum number of iterations of 12,000. Among them, the rough makeup stage has 10,000 steps, and the refinement stage has 2,000 steps.

[0070] To verify the feasibility and beneficial effects of the present invention, the following comparative experiments were carried out. The numerical experiments were carried out under the dataset LADN.

[0071] The makeup transfer effect of the current CLIPFace model for drivable 3D avatar editing was selected for comparison, and the comparison results are shown in Tables 1 and 2. The CLIPFace model selects the mesh as the 3D representation and uses the CLIP loss for makeup transfer.

[0072] The present invention selects two evaluation indicators, CLIP similarity and DINO similarity, to evaluate the trained model. The CLIP similarity is used to evaluate the semantic similarity between the finally obtained makeup image and the reference image, and the larger the value, the higher the model accuracy. The DINO similarity is used to evaluate the similarity between the finally obtained makeup image and the reference image. To evaluate the makeup transfer quality of the rendered images at different angles and the quality of makeup retention when driving the human face, two groups of experiments were carried out. The first group tested the makeup transfer quality at different perspectives under a single expression. Images were rendered at three horizontal angles: 0°, 45° - 45°, and the CLIP similarity and DINO similarity of each group were calculated according to the corresponding angles respectively. Finally, the three values were averaged, and the comparison results are shown in Table 1. For the second group of experimental tests, the human face was driven to make different expressions at the same angle and the images were rendered, and multi-angle pictures were added to test the CLIP similarity and DINO similarity of these images, and the comparison results are shown in Table 2.

[0073] Table 1 Comparison of static multi-perspective generation results

[0074]

[0075]

[0076] Table 2 Comparison of dynamic multi-perspective generation results

[0077] model CLIP similarity DINO similarity AvatarMakeup 0.632 0.627 CLIPFace 0.664 0.337

[0078] As can be seen from Table 1 and Table 2, the CLIP similarity obtained using the CLIPFace model is higher numerically. This is because the CLIPFace model itself adopts the CLIP loss, aiming to improve the CLIP similarity, so the numerical value will be higher. However, the numerical value of the DINO similarity is much weaker than that of the present invention. The AvatarMakeup model of the present invention has relatively high numerical values in both CLIP similarity and DINO similarity, demonstrating the superiority of the present invention.

[0079] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A realistic makeup transfer method for 3D animatable avatars based on diffusion models, characterized in that, Specifically, it includes the following steps: S1. Obtain a reconstructed three-dimensional drivable target human head and a single-view reference photo with a target makeup effect as the input of the stable makeup model; S2. Use the stable makeup model to generate a series of makeup images, optimize the global UV map with the makeup images, and then sample the makeup images from the global UV map as supervision information to optimize the makeup of the target human head; S3. Add a refinement mechanism to the stable makeup model, and further optimize the makeup details of the target human head on the basis of the rough makeup in step S2; S4. Optimize the makeup of the target human head in combination with the loss function, and finally generate a target human head with a high-quality makeup effect.

2. The realistic makeup transfer method for a 3D animatable avatar based on a diffusion model according to claim 1, wherein, Step S1 specifically includes the following steps: S1.

1. Use existing three-dimensional head reconstruction technologies, such as Gaussian character models, to generate a three-dimensional drivable target human head; S1.

2. The user provides a single reference picture R with a target makeup effect.

3. A realistic makeup transfer method for three-dimensional animatable avatars based on diffusion models according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.1, given k camera perspectives {π1,π2,...,π k }, and render the target person's head portrait into a picture {x1,x2,...,x k }; S2.2, for the rendered image x i add noise for t time steps to obtain Then, through the stable makeup model θ, perform a denoising process on the noisy image to generate makeup transfer images under k camera views {π1, π2,..., π k} S2.

3. Render the mapping image of each point in the makeup transfer image to the UV coordinates through the mesh renderer Optimize the global UV texture map by accumulating the makeup transfer images from multiple perspectives. The optimization formula for the UV texture map is as follows: Among them, (H, W) and (h, w) are pixel coordinates, is the pixel value at (H, W), and UVmap(h, w) is the pixel value of the UV map; S2.

4. Use the global UV map to render the makeup images of the target human head at any expression and perspective through a mesh renderer, and use step S4 and take the makeup images as supervision pictures to optimize the makeup of the target human head.

4. A realistic makeup transfer method for three-dimensional animatable avatars based on a diffusion model according to claim 1, characterized in that, Step S3 specifically includes the following steps: S3.

1. Given the target person's head portrait obtained in step S2, a series of pictures {y1, y2,..., y n} are rendered for the target person's head portrait; S3.2, for the time step t n selected when adding noise to {y1, y2,..., y ′} and t ∈ [20, 350], we get S3.3, input into the stable makeup model for denoising. Similar to step S2.2, obtain the denoised image S3.4, Use as the supervision image to optimize the target person's head portrait, and use step S4 to obtain the optimized makeup of the target person's head portrait.

5. A realistic makeup transfer method for 3D animatable avatars based on a diffusion model according to claim 1, characterized in that, Step S4 specifically includes the following steps: S4.

1. For the supervised images in both the S2 and S3 stages, the regions that need to undergo makeup transfer are optimized using both the L1 loss and the L LPIOS loss: L makeup = L1(M⊙I G , M⊙I r ) + L LPIPS (M⊙I G , M⊙I r ); Among them, M represents the area where makeup transfer needs to be performed, and I G represents the supervision picture, I r represents the rendering picture, and ⊙ is the element multiplication; S4.

2. For areas where makeup transfer is not required, use the L1 loss to maintain the same texture features as the original character: L Res = L1((1 - M) ⊙ I ID , (1 - M) ⊙ I r ); Among them, I ID represents the rendered image of the initially reconstructed target person's head portrait.

Citation Information

Patent Citations

  • Makeup migration method based on generative adversarial network

    CN115496650A

  • Video makeup migration method and system

    CN115689869A

  • Chinese opera makeup migration method based on UV space mapping

    CN117196935A

  • Three-dimensional image generation method, three-dimensional figure image generation method and computing device

    CN118115706A

  • Face make-up real-time generation method based on two-stage diffusion model

    CN119295590A