SSS compatible face reflectivity estimation method based on diffusion model

Through the diffusion model-based method, UV maps are generated using DSN-Net and HM-Net algorithms, which solves the problems of strong lighting dependence and insufficient SSS compatibility in the existing facial reflectivity estimation method, real skin rendering and efficient calculation under different lighting conditions are achieved.

CN120279147AInactive Publication Date: 2025-07-08SHANGHAI XUESHEN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371587.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing facial reflectivity estimation method cannot simulate subsurface scattering, resulting in the skin looking stiff and unnatural, strong dependence on the lighting environment, high computational complexity, lack of SSS compatibility, cannot estimate hemoglobin and melanin, weak generalization ability, and cannot adapt to different skin tones, ages and lighting conditions.

Method used

Using a diffusion model-based method, latent variables are extracted by freezing encoder, single-image-multireflection mode and multi-image-single-diffuse reflection mode are used, and denoising optimization is combined with DSN-Net and HM-Net algorithms to generate UV maps to ensure light consistency and biological characteristics, and hemoglobin and melanin maps are generated.

Benefits of technology

The authenticity and stability of skin rendering under different lighting conditions are achieved, the computing efficiency and generalization ability are improved, the generated UV maps meet the characteristics of real skin, adapt to different lighting conditions, and enhance SSS compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279147A_ABST
    Figure CN120279147A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of three-dimensional face animation generation, in particular to an SSS compatible face reflectivity estimation method based on a diffusion model, and the method comprises the following steps: S1, taking a face image as input, carrying out the standardization processing, enabling the face image to meet the model input requirement, extracting a latent variable of the input image through a frozen encoder, and carrying out the recognition of the latent variable; according to the method, high-dimensional image data are converted into low-dimensional features, so that the calculation complexity is reduced, the training efficiency is improved, and the problems that according to an existing estimation method, subsurface scattering cannot be simulated, and consequently the skin looks stiff and unnatural; the dependence on the illumination environment is strong, and the rendering effect is unstable under different illumination conditions; the calculation complexity is high, and the performance is limited in high-quality 3D rendering; data dependence is high, high-quality data collected by Light Stage is needed, and training data is insufficient; the problems of lack of SSS compatibility, incapability of estimating hemoglobin and melanin, influence on sense of reality, weak generalization ability, incapability of adapting to different skin colors, ages and illumination conditions and the like are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional face animation generation, and specifically to a method for estimating SSS-compatible facial reflectance based on a diffusion model. Background Art

[0002] In the fields of computer graphics and computer vision, high-quality facial reflectance estimation is an important step in achieving realistic 3D face rendering. Existing facial reflectance estimation methods mainly rely on high-precision data collected from light fields. However, these methods are limited by high data acquisition costs, limited applicability, and difficulty in generalizing to more diverse individuals. In recent years, deep learning-based methods have been used for face reflectance estimation, and methods such as StyleGAN, Relightify, and ID2Reflectance have attempted to use data-driven techniques to improve the accuracy of reflectance prediction.

[0003] However, the existing estimation methods still have deficiencies, specifically: the existing estimation methods cannot simulate subsurface scattering, resulting in a rigid and unnatural appearance of the skin, strong dependence on the lighting environment, unstable rendering effects under different lighting conditions, high computational complexity, limited performance in high-quality 3D rendering, strong data dependence, requiring high-quality data collected by a Light Stage, insufficient training data; lack of SSS compatibility, unable to estimate hemoglobin and melanin, affecting the realism, weak generalization ability, and unable to adapt to different skin colors, ages, and lighting conditions.

[0004] Therefore, a method for estimating SSS-compatible facial reflectance based on a diffusion model is needed to solve the problems raised in the above background art. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for estimating SSS-compatible facial reflectance based on a diffusion model to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A method for estimating SSS-compatible facial reflectance based on a diffusion model, comprising the following steps:

[0008] S1. Using a face image as input, after standardization processing to make it meet the model input requirements, a frozen encoder is used to extract the latent variables of the input image, converting high-dimensional image data into low-dimensional features to reduce computational complexity and improve training efficiency. The single-image - multiple reflection mode is adopted. Under the same input image, diffuse reflection, specular reflection, and normal are predicted simultaneously to ensure that the lighting calculations of different reflectance information on the same face are consistent. The multi-image - single diffuse reflection mode is adopted to ensure that the diffuse reflection map remains consistent under different lighting conditions and avoid the influence of lighting on skin color;

[0009] S2. The joint reflection attention mechanism is used to optimize the map, making the lighting calculations of different maps at the same part consistent. The DSN-Net algorithm is used for denoising optimization in the latent space and trained using a loss function. After training, the DSN-Net algorithm generates three UV maps;

[0010] S3. Using the diffuse reflection map generated by the DSN-Net as input to remove the influence of lighting, making the subsequent prediction of hemoglobin and melanin more stable. The HM-Net algorithm is optimized for denoising through a diffusion model, gradually generating the hemoglobin map and the melanin map. The HM-Net algorithm is optimized for denoising in the latent space, gradually restoring the true distribution of hemoglobin and melanin to conform to the physical characteristics of subsurface scattering. A loss function is used to train the HM-Net algorithm to ensure that the generated UVh and Uvm maps conform to the biological characteristics of real skin. After training, the HM-Net algorithm generates two UV maps.

[0011] As a preferred solution of the present invention, the input face image data in S1 includes: an RGB color image with a size of 768×768, lighting change image data of the same face under different lighting conditions, and a high-resolution UV map for training the generation of diffuse reflection, specular reflection, and normal maps.

[0012] As a preferred solution of the present invention, the specific method of adopting the single-image - multiple reflection mode in S1 is to use a switch to control the currently predicted reflectance type (diffuse reflection, specular reflection, normal), and optimize through a multi-task loss function to ensure the geometric consistency of the three maps during training.

[0013] As a preferred solution of the present invention, the specific method of adopting the multi-image - single diffuse reflection mode in S1 is to use images of the same face under different lighting conditions as input, ensuring that the model learns to ignore the influence of lighting, only retain the inherent skin color, and limit the predicted values of UVd under different lighting conditions to be close through a consistency loss to improve the generalization ability.

[0014] As a preferred solution of the present invention, the loss function for training the DSN-Net algorithm in S2 is where \(i\in\{d, s, n\}\) and \(r\) i are the reflectivity domain and the domain indicator respectively, \(K\) is the index of the color image at different luminances, where \(\epsilon\) t i is Gaussian noise independently sampled from a multi-scale noise set.

[0015] As a preferred solution of the present invention, the three UV maps generated by the DSN-Net algorithm in S2 include a diffuse map, a specular map, and a normal map. The diffuse map is used to represent the base color of the skin and remove the influence of lighting. The specular map is used to represent the gloss information of the skin, such as grease reflection. The normal map is used to represent the surface structure of the skin, such as wrinkles and pores.

[0016] As a preferred solution of the present invention, the loss function for training the HM-Net algorithm in S3 is where \(j\in\{h, m\}\) and \(r\) d are the domain indicators of melanin and hemoglobin respectively, where \(\epsilon\) t j is Gaussian noise independently sampled from a multi-scale noise set.

[0017] As a preferred solution of the present invention, the two UV maps generated by the HM-Net algorithm in S3 are a hemoglobin map and a melanin map.

[0018] Compared with the prior art, the beneficial effects of the present invention are:

[0019] 1. In the present invention, a face image is used as the input. After being standardized to meet the model input requirements, a frozen encoder is employed to extract the latent variables of the input image, converting high-dimensional image data into low-dimensional features to reduce computational complexity and improve training efficiency. The single-image - multiple reflection mode is adopted. Under the same input image, diffuse reflection, specular reflection, and normal are predicted simultaneously to ensure that the lighting calculations of different reflectance information on the same face are consistent. The multi-image - single diffuse reflection mode is used to ensure that the diffuse reflection map remains consistent under different lighting conditions, avoiding the influence of lighting on skin color. The joint reflection attention mechanism is used to optimize the map, making the lighting calculations of different maps consistent at the same location. The DSN-Net algorithm is used to perform denoising optimization in the latent space and is trained using a loss function. After training, the DSN-Net algorithm generates three UV maps. Using the diffuse reflection map generated by DSN-Net as the input to remove the influence of lighting makes the subsequent prediction of hemoglobin and melanin more stable. The HM-Net algorithm performs denoising optimization through a diffusion model, gradually generating the hemoglobin map and the melanin map. The HM-Net algorithm performs denoising optimization in the latent space, gradually restoring the true distribution of hemoglobin and melanin to conform to the physical characteristics of subsurface scattering. A loss function is used to train the HM-Net algorithm to ensure that the generated UVh and UVm maps conform to the biological characteristics of real skin. After training, the HM-Net algorithm generates two UV maps. By adopting SSS-compatible modeling, the present invention makes skin rendering more realistic, estimates hemoglobin and melanin from outdoor face images, allows light to scatter under the skin, and generates a more natural skin lighting effect. The present invention uses Stable Diffusion as the prior model, pre-trains with large-scale Internet data, and can generate high-quality facial reflectance maps after fine-tuning on a small dataset. High-quality reflectance images can be generated from ordinary pictures. Using the two-stage algorithm structure of DSN-Net and HM-Net and the Latent Diffusion model to perform denoising inference in the latent space improves computational efficiency, reduces training and inference time, and further optimizes SSS-related information after generating the basic reflectance, making the skin texture clearer. At the same time, the multi-image - single diffuse reflection mode and the joint reflection attention mechanism are adopted to ensure stable lighting estimation and improve adaptability under different lighting conditions. Description of the Drawings

[0020] Figure 1 It is a flowchart of the DSN-Net algorithm of the present invention;

[0021] Figure 2 It is a flowchart of the HM-Net algorithm of the present invention;

[0022] Figure 3 It is a schematic diagram of the first specific embodiment of the present invention;

[0023] Figure 4 This is a schematic diagram of Specific Embodiment 2 of the present invention. Specific Embodiment

[0024] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0025] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0027] For the embodiments, please refer to Figures 1-4 , the present invention provides a technical solution:

[0028] An SSS-compatible facial reflectance estimation method based on a diffusion model, comprising the following steps:

[0029] S1. Using a face image as input, after standardization processing to make it meet the model input requirements, a frozen encoder is used to extract the latent variables of the input image, converting high-dimensional image data into low-dimensional features to reduce the computational complexity and improve the training efficiency. The single-image - multiple reflection mode is adopted to simultaneously predict the diffuse reflection, specular reflection, and normal on the same input image, ensuring that the lighting calculations of different reflectance information on the same face are consistent. The multi-image - single diffuse reflection mode is adopted to ensure that the diffuse reflection map remains consistent under different lighting conditions and avoid the influence of lighting on skin color;

[0030] S2. The joint reflection attention mechanism is used to optimize the map, making the lighting calculations of different maps consistent at the same part. The DSN-Net algorithm is used for denoising optimization in the latent space and trained with a loss function. After training, the DSN-Net algorithm generates three UV maps;

[0031] S3 uses the diffuse texture map generated by DSN-Net as input to remove the influence of lighting, making the subsequent prediction of hemoglobin and melanin more stable. The HM-Net algorithm performs denoising optimization through a diffusion model to gradually generate the hemoglobin texture map and the melanin texture map. The HM-Net algorithm performs denoising optimization in the latent space to gradually restore the true distributions of hemoglobin and melanin, making them conform to the physical characteristics of subsurface scattering. The loss function is used to train the HM-Net algorithm to ensure that the generated UVh and Uvm texture maps conform to the biological characteristics of real skin. After training, the HM-Net algorithm generates two UV texture maps.

[0032] Furthermore, the input face image data in S1 includes: an RGB color image with a size of 768×768, the lighting change image data of the same face under different lighting conditions, and the high-resolution UV texture map used for training the generation of diffuse, specular, and normal texture maps.

[0033] Furthermore, the specific method of using the single-image - multiple reflection mode in S1 is to use a switch to control the currently predicted reflectance type (diffuse, specular, normal), and optimize through a multi-task loss function to ensure the geometric consistency of the three texture maps during training.

[0034] Furthermore, the specific method of using the multi-image - single diffuse mode in S1 is to use the images of the same face under different lighting conditions as input, ensuring that the model learns to ignore the influence of lighting and only retain the inherent color of the skin, and restricting the UVd prediction values under different lighting conditions to be close through a consistency loss to improve the generalization ability.

[0035] Furthermore, the loss function for training the DSN-Net algorithm in S2 is where i ∈ {d, s, n} and r i are each reflectance domain and domain indicator, K is the exponent of the color image at different brightness levels, where ε t i is Gaussian noise independently sampled from a multi-scale noise set.

[0036] Furthermore, the three UV texture maps generated by the DSN-Net algorithm in S2 include a diffuse texture map, a specular texture map, and a normal texture map. The diffuse texture map is used to represent the base color of the skin and remove the influence of lighting. The specular texture map is used to represent the gloss information of the skin, such as oil reflection. The normal texture map is used to represent the surface structure of the skin, such as wrinkles and pores.

[0037] Furthermore, the loss function for training the HM-Net algorithm in S3 is where j ∈ {h, m} and r dare the domain indicators of melanin and hemoglobin respectively, where ε t j is Gaussian noise independently sampled from a multi-scale noise set.

[0038] Furthermore, in step S3, the HM-Net algorithm generates two UV maps, namely the hemoglobin map and the melanin map.

[0039] Specific implementation case 1

[0040] Input a normal face image in the wild environment. The running hardware device is an NVIDIA A6000 GPU, and the running programming framework is PyTorch. The inference time is about 25 seconds for each face. The input data is a face image in RGB color in the wild environment, with a size of 768×768. It can be multi-angle data, including different head postures; it can be under uneven lighting conditions, such as outdoor sunlight, indoor lighting, and multi-light source environments. In the first stage, the DSN-Net algorithm is used to predict diffuse reflection, specular reflection, and normal. In the second stage, the HM-Net algorithm is used to predict hemoglobin and melanin. A total of 5 UV maps are output, and 3D skin lighting is rendered to display the final reflectance estimation result.

[0041] Specific implementation case 2

[0042] The process is to input a patient's face photo. The running hardware device is an NVIDIA A6000 GPU, and the running programming framework is PyTorch. The inference time is about 25 seconds for each face. After the DSN-Net algorithm in the first stage predicts diffuse reflection, specular reflection, and normal, then the HM-Net algorithm in the second stage predicts hemoglobin and melanin. Finally, hemoglobin and melanin maps are output. In the attached instructions Figure 4 the first row is the real data collected by the VISIA skin analyzer, the second row is the predicted hemoglobin and melanin maps, which are highly consistent with the real situation, and the third and fourth rows are the inference results on the wild images.

[0043] It can be seen from specific implementation cases 1-2 that the present invention can be applied to face photos in the wild, maintain consistent skin lighting simulation ability under different lighting conditions, improve the authenticity of 3D characters in complex scenes, and can also be used for skin disease diagnosis and skin care assessment, assisting doctors in formulating treatment plans and improving the convenience of large-scale skin disease screening. The present invention adopts the Stable Diffusion prior and two-stage training strategy. Compared with existing methods, it has obvious improvements in terms of SSS compatibility, generalization ability, computational efficiency, detail retention, and lighting adaptability, and can be applied to multiple fields such as games, movies, VR / AR, and medicine, providing a more efficient, more realistic, and more stable facial reflectance estimation solution.

[0044] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for estimating SSS-compatible facial reflectance based on a diffusion model, characterized in that, It includes the following steps: S1. Using a face image as input, after standardization to meet the model input requirements, a frozen encoder is adopted to extract the latent variables of the input image, converting high-dimensional image data into low-dimensional features to reduce computational complexity and improve training efficiency. The single-image - multiple reflection mode is adopted, and under the same input image, diffuse reflection, specular reflection, and normal are predicted simultaneously to ensure consistent illumination calculation of different reflectivity information on the same face. The multi-image - single diffuse reflection mode is adopted to ensure that the diffuse reflection map remains consistent under different lighting conditions and avoid the influence of lighting on skin color; S2. The joint reflection attention mechanism is used to optimize the map, making the illumination calculation of different maps consistent at the same part. The DSN-Net algorithm is used for denoising optimization in the latent space and trained using a loss function. After training, the DSN-Net algorithm generates three UV maps; S3. Using the diffuse reflection map generated by the DSN-Net as input to remove the influence of lighting, making the subsequent prediction of hemoglobin and melanin more stable. The HM-Net algorithm is denoised and optimized through a diffusion model, gradually generating a hemoglobin map and a melanin map. The HM-Net algorithm is denoised and optimized in the latent space, gradually restoring the true distribution of hemoglobin and melanin to conform to the physical characteristics of subsurface scattering. A loss function is used to train the HM-Net algorithm to ensure that the generated UVh and UVm maps conform to the biological characteristics of real skin. After training, the HM-Net algorithm generates two UV maps.

2. The SSS-compatible facial reflectance estimation method based on a diffusion model according to claim 1, wherein: The face image data input in S1 includes: an RGB color image with a size of 768×768, illumination change image data of the same face under different illuminations, and a high-resolution UV map for training the generation of diffuse reflection, specular reflection, and normal maps.

3. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The specific method of adopting the single-image - multiple reflection mode in S1 is to use a switch to control the currently predicted reflectivity type (diffuse reflection, specular reflection, normal), and optimize it through a multi-task loss function to ensure geometric consistency of the three maps during training.

4. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The specific method of adopting the multi-image - single diffuse reflection mode in S1 is to use images of the same face under different illuminations as input, ensuring that the model learns to ignore the influence of lighting, only retaining the inherent skin color, and restricting the predicted values of UVd under different lighting conditions to be close through a consistency loss to improve the generalization ability.

5. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The loss function for training the DSN-Net algorithm in S2 is where i ∈ {d, s, n} and r i are the reflectivity domain and domain indicator respectively, K is the exponent of the color image under different luminances, where is Gaussian noise independently sampled from a multi-scale noise set.

6. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The three UV maps generated by the DSN-Net algorithm in S2 include a diffuse reflection map, a specular reflection map, and a normal map. The diffuse reflection map is used to represent the base color of the skin and remove the influence of lighting. The specular reflection map is used to represent the gloss information of the skin, such as grease reflection. The normal map is used to represent the surface structure of the skin, such as wrinkles and pores.

7. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The loss function for training the HM-Net algorithm in S3 is where j ∈ {h, m} and r d are the domain indicators of melanin and hemoglobin respectively, where is Gaussian noise independently sampled from a multi-scale noise set.

8. A method for estimating SSS-compatible facial reflectance based on a diffusion model according to claim 1, characterized in that: The two UV maps generated by the HM-Net algorithm in S3 are a hemoglobin map and a melanin map.