360-degree drivable three-dimensional head portrait method supporting multi-mode generation

Through a multi-stage generation framework combining three-dimensional generative adversarial network, FLAME model and Gaussian sputtering rendering algorithm, the problem of difficult to generate high-quality, full-view and driveable three-dimensional head models in the existing technology is solved, and a 360-degree three-dimensional head model generation with high fidelity and multi-view consistency is achieved.

CN120147487APending Publication Date: 2025-06-13NANJING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510300655.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently generate high-quality, full-view and driveable three-dimensional human head models, especially in terms of diversification and efficient generation.

Method used

A multi-stage generation framework is adopted, combining three-dimensional generative adversarial networks, FLAME models and Gaussian sputtering rendering algorithms, and a 360-degree driveable three-dimensional head model is generated through multimodal inputs (random noise, text, and pictures). The framework includes optimizations for generating head images, optimizing FLAME model parameters, texture maps, and 3D Gaussian optimizations to improve rendering quality and drive robustness.

Benefits of technology

The 360-degree three-dimensional human head model with high fidelity, driveability and multi-view consistency is achieved, reducing the dependence on high-precision devices and is suitable for virtual reality, meta-universe and digital human fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147487A_ABST
    Figure CN120147487A_ABST
Patent Text Reader

Abstract

The invention discloses a 360-degree drivable three-dimensional head portrait method supporting multi-mode generation. The method comprises the following steps: generating a 360-degree head image in combination with a three-dimensional generative adversarial network; three-dimensional Gaussian flattening is attached to a surface patch of the initialized FLAME model, and parameters of the FLAME model are optimized; the vertex displacement of the FLAME model is further optimized, and a refined grid model structure is obtained; using a differentiatable renderer to optimize the texture map; and the shape and size of the three-dimensional Gaussian are freely optimized, so that the three-dimensional Gaussian captures a high-frequency detail region, and finally, a three-dimensional head image which has high rendering quality and driving performance and supports various generation input forms is generated. By combining the three-dimensional generative adversarial network, the FLAME model and the differentiable renderer, a brand new generation framework is provided, and a high-quality and full-view renderable three-dimensional head generation function is realized from multiple input forms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for a 360-degree drivable three-dimensional avatar supporting multi-modal generation. Background Art

[0002] With the rapid development of virtual reality, the metaverse, and game technologies, the demand for high-quality, full-view three-dimensional digital human heads is increasing day by day. Existing methods mostly rely on high-cost large-scale three-dimensional scanning devices or manual modeling, and it is difficult to meet the diverse and efficient generation requirements.

[0003] In recent years, digital face synthesis models have attracted wide attention. Existing work can be divided into two categories: face modeling and face generation. Techniques based on generative adversarial networks (GANs) have made remarkable progress in the field of face image generation and editing. StyleGAN proposed by Karras et al. (Tero Karras, Samuli Laine, and Timo Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2019, pp. 4401-4410.) has greatly improved the stability and quality of 2D face image generation, but it cannot maintain multi-view consistency. The model for generating two-dimensional images from text proposed by Xia et al. (Weihao Xia, Yujiu Yang, Jing-Hao Xue, Baoyuan Wu. TediGAN: Text-Guided Diverse Face Image Generation and Manipulation. In CVPR, pages 2256-2265, 2021.2) consists of a text encoder and a generative neural network. The text-to-face model has made progress in image quality, but there are still challenges in accurately mapping text to image content.

[0004] Chan et al. introduced a tri-plane representation in EG3D (Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al., “Efficient geometry-aware 3d generative adversarial networks,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2022, pp. 16123-16133.) to achieve preliminary 3D face generation, but it is still difficult to generate dynamic and drivable faces. An et al. introduced a tri-mesh structure on the basis of EG3D in PanoHead (Sizhe An, Hongyi Xu, Yichun Shi, Guoxian Song, Umit Y Ogras, and Linjie Luo, “Panohead: Geometry-aware 3d full-head synthesis in 360 deg,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 20950–20959.) to achieve 360-degree head generation and support the reconstruction of input face images, but the generated heads still cannot be driven.

[0005] For the goal of face modeling, Bernhard Kerbl, Georgios Kopanas, Thomas and George Drettakis, "3D Gaussian Splatting for Real-Time Radiance Field Rendering," ACM Transactions on Graphics, vol. 42, no. 4, July 2023.) Gaussian splatting exhibits excellent performance in novel view synthesis. Shao et al. (Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang, "SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting," in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.) combined Gaussian splatting with parametric models such as SMPL-X and FLAME to learn digital avatar representations. However, the generated digital humans require inputting a video and performing operations such as camera calibration and data preprocessing on the input video, which raises the application threshold. Summary of the Invention

[0006] To generate a drivable 360-degree renderable human head, the present invention provides a method for a 360-degree drivable three-dimensional avatar that supports multi-modal generation.

[0007] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0008] A method for a 360-degree drivable three-dimensional avatar that supports multi-modal generation, comprising the following steps:

[0009] S1, generating a 360-degree human head image by combining a three-dimensional generative adversarial network; including generating multi-view human head images from random noise, text descriptions, or single images;

[0010] S2, fitting three-dimensional Gaussian flattening onto the patches of an initialized FLAME model, calculating the loss between the image generated by Gaussian rendering and the human head image generated in step S1, and optimizing the FLAME model parameters;

[0011] S3, fixing the parameters of the FLAME model and further optimizing the vertex displacement of the FLAME model to obtain a refined mesh model structure;

[0012] S4. Initialize the texture map. Render on the initialized texture map using the UV mapping of the FLAME model on the refined mesh model structure obtained in step S3 through the differentiable renderer, and calculate the loss with the human head image generated in step S1 to optimize the texture map and obtain the corresponding texture map.

[0013] S5. Freely optimize the scale and rotation attributes of the three-dimensional Gaussian on the FLAME model patches to capture more high-frequency detail areas, so as to generate a three-dimensional Gaussian digital human expression with higher rendering quality.

[0014] The present invention aims to provide a full-view renderable three-dimensional human head generation method supporting multi-modal input. By combining an advanced human head generation model, a parametric face model, and a Gaussian sputtering rendering algorithm, a multi-stage generation framework is proposed to learn the distributions of controllable face pictures, face meshes, textures, and three-dimensional Gaussians, thereby generating a high-fidelity and drivable three-dimensional human head model, which solves the problems of strong data dependence, low rendering quality, and poor drivability existing in the prior art. The present invention realizes the generation of a drivable three-dimensional human head model for the first time, fills the research gap in this aspect, and can directly generate and input text or pictures for generation, that is, supports multi-modal input. The proposed method can be widely applied to fields such as digital humans, game creation, and movie special effects, and has high practical value and development prospects. Brief Description of the Drawings

[0015] Figure 1 It is a flowchart of the method of the present invention.

[0016] Figure 2 It is a flowchart of the running stage in the embodiment of the present invention.

[0017] Figure 3 It is a 360-degree rendering diagram in the embodiment of the present invention. Among them, the first row is a three-dimensional human head image generated from random noise, the second row is a three-dimensional human head image generated from text, and the third row is a three-dimensional human head image generated from a picture.

[0018] Figure 4 It is a driving rendering diagram in the embodiment of the present invention. Among them, the first column is an example diagram of the driving face actions and expressions. The first row is the rendering diagram of the three-dimensional human head generated from random noise driven by the example actions and expressions, the second row is the rendering diagram of the three-dimensional human head generated from text driven by the example actions and expressions, and the third row is the rendering diagram of the three-dimensional human head generated from a picture driven by the example actions and expressions. Detailed Embodiment

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] As Figure 1 shown, a method for a 360-degree drivable three-dimensional avatar that supports multi-modal generation in this embodiment is as follows:

[0021] S1. Combine a three-dimensional generative adversarial network to generate a 360-degree head image, including generating multi-view head images from random noise, text descriptions, or single images, and supporting multi-modal input. The three-dimensional generative adversarial network in this embodiment adopts the PanoHead model, which is a model capable of generating 360-degree static head pictures.

[0022] For generating a corresponding head image from a text description, fix the camera parameters of the PanoHead model, generate a large number of head images, and save the latent vector w of the PanoHead model. The generated pictures and the feature vectors Vi of the image encoder of the CLIP model form a paired dataset. Train a diffusion model to map Vi to w. Specifically, the diffusion model disrupts the data distribution by gradually adding noise, and then learns how to reverse this process to recover the original data from the noise. In the prior art, diffusion models are used to map text feature vectors to the image feature space. In the present invention, a diffusion model is used to map the feature vector Vi generated by the image encoder of the CLIP model to the latent vector w of the PanoHead generator. During the inference process, the text encoder of the CLIP model converts the text into a feature vector Vt, and the trained diffusion model maps Vt to w, enabling the PanoHead model to generate a head image corresponding to the input text.

[0023] For generating a corresponding 360-degree head image from a single picture, adopt the GAN inversion technology PTI, that is, first fix the parameters of the PanoHead model, calculate the loss through the input picture, random sampling noise, and the picture generated by the PanoHead model, optimize the latent vector w of the PanoHead model, and then fix the vector w to optimize the PanoHead model. Applying PTI to the PanoHead model can complete the generation of a multi-view head image corresponding to the input head picture.

[0024] Generating a 360-degree head image from random noise involves randomly sampling a noise from a normal distribution and generating a 360-degree head image through the PanoHead model. The generated image is uniformly sampled at 240 viewpoints across the full 360-degree view, and the head image generated through the PanoHead model is used as the training data for the subsequent steps.

[0025] S2. Flatten and fit the 3D Gaussian onto the patches of the initialized FLAME model. Determine the position, rotation, and size attributes of the 3D Gaussian based on the vertex positions of the patches where the 3D Gaussian is located:

[0026] p = α 1 v 1 +α 2 v 2 +α 3 v 3

[0027] S = diag(s 1 ,s 2 ,s 3 )

[0028] where p is the position of the 3D Gaussian, α 1 +α 2 +α 3 = 1 and is a trainable parameter, v 1 ,v 2 ,v 3 are the coordinates of the three vertices of the patch where the 3D Gaussian is located; the rotation matrix R of the 3D Gaussian is [r 1 ,r 2 ,r 3 , r 1 is the normal vector of the patch it is on, r 2 is the unit vector from the first vertex of the patch to the centroid of the patch, r 3 is the unit vector orthogonal to r 1 and r 2 ; S is the size matrix of the 3D Gaussian, s 1 = ε, ε = 1e-5, s 2 = ||mean(v 1 ,v 2 ,v 3 ) - v 1 ||, s 3 = v 2 ·r 3On each patch of the FLAME model, there are 20 three-dimensional Gaussians. The loss is calculated between the image generated by rendering the three-dimensional Gaussians bound to FLAME and the generated face image to optimize the FLAME model parameters. In the stage of optimizing the FLAME model parameters, this embodiment mainly adjusts parameters such as scale, translation, shape, pose, expression, and neck pose, and iteratively optimizes them 5000 times. In this stage, the loss function adopts a combination of L1 loss and SSIM loss:

[0029] L=(1 - λ)L 1 +λL SSIM

[0030] where λ = 0.2, representing the weight of the SSIM loss in the loss function.

[0031] S3. After obtaining the optimized FLAME model parameters, set the learning rate of the FLAME model parameters to zero to freeze the FLAME parameters, and continue to optimize the displacement of each vertex of FLAME along the vertex normal vector. After fixing the displacement direction of the vertex, make the displacement of the vertex more controllable, and obtain a refined mesh model structure after optimization, improving the rendering quality of the human head. In addition, this embodiment adds Laplacian loss and normal consistency loss to the training loss:

[0032] L=(1 - λ)L 1 +λL SSIM +λ 1 L Laplasian +λ 2 L normal

[0033] where λ 1 =λ 2 =0.05 - 0.2, representing the weight of the corresponding loss function. This is different from the loss function of the existing Gaussian rendering. The added Laplacian loss and normal consistency loss enable the FLAME model with displaced vertices to still maintain a certain surface smoothness and overall high quality.

[0034] S4. Initialize the texture map, and use the UV mapping of the FLAME model to render on the initialized texture map through a differentiable renderer on the refined mesh model structure obtained in step S3, and calculate the loss with the generated 360-degree human head image to optimize the texture map and obtain the corresponding texture map. In this embodiment, the initialized texture map is an image with a resolution of 512*512, and the differentiable renderer is Nvdiffrast.

[0035] S5. Freely optimize the scale and rotation attributes of the three-dimensional Gaussians on the FLAME model patches to capture more high-frequency detail areas (such as hair) to generate a three-dimensional Gaussian digital human expression with higher rendering quality.

[0036] Before starting the optimization, the training parameters must be re-initialized to avoid getting stuck in local minima. At this stage, the grid-related regularization loss is not required in the loss function, and the opacity loss L opacity is added to prevent the Gaussian from becoming transparent:

[0037] L = (1 - λ)L 1 + λL SSIM + λ 3 L opacity

[0038] where the weight λ 3 = 0.05 - 0.2.

[0039] Through the above steps, the present invention can generate a drivable 360-degree renderable three-dimensional human head from random noise, text, or pictures, and the generation results are as shown in Figure 3 , 4 . It can be seen from Figure 3 that the generated human head can be rendered 360 degrees and has multi-view consistency. Figure 4 It can be seen that the generated human head can be driven to make different actions.

[0040] By combining Gaussian sputtering, the FLAME model, and the human head generation model, the present invention proposes a multi-stage generation framework to learn the expression of digital human heads, thereby generating a three-dimensional human head model with driving robustness and high rendering quality. The present invention proposes a method for generating a full-view drivable three-dimensional human head, which can generate a high-quality, full-view three-dimensional human head through multi-modal input. This method reduces the dependence on high-precision devices and has broad application prospects, and is applicable to the fields of virtual reality, the metaverse, and digital humans.

Claims

1. A method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar, characterized in that: The steps include: S1, generates 360-degree head images by combining a 3D generative adversarial network; including generating multi-view head images from random noise, text description or a single image; S2, flattening the three-dimensional Gaussian onto the surface of the initialized FLAME model, performing loss calculation on the image generated by Gaussian rendering and the head image generated in step S1, and optimizing the FLAME model parameters; S3, fix the parameters of the FLAME model, further optimize the vertex displacement of the FLAME model, and obtain a refined mesh model structure; S4, initializing the texture map, rendering on the initialized texture map using the UV mapping of the FLAME model on the refined mesh model structure obtained in step S3 through a differentiable renderer, performing loss calculation on the head image generated in step S1, optimizing the texture map, and obtaining a corresponding texture map; S5, freely optimizes the scale and rotation properties of the 3D Gaussian on the FLAME model patch to capture more high-frequency detail areas to generate a 3D Gaussian digital human expression with higher rendering quality.

2. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In the step S1, when a multi-view head image is generated from a text description, the latent vector w of the three-dimensional generative adversarial network and the feature vector Vi of the image encoder of the CLIP model are used to form a paired data set; then the diffusion model is used to map the feature vector Vi to the latent vector w; during the inference process, the text encoder of the CLIP model converts the text into a feature vector Vt, and the diffusion model is used to map the feature vector Vt to the latent vector w.

3. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In the step S1, when generating a multi-view head image from a single image, the parameters of the three-dimensional generative adversarial network are fixed, the loss of the input single image and the generated image is calculated, the latent vector w of the three-dimensional generative adversarial network is optimized, the optimized latent vector is fixed, the three-dimensional generative adversarial network is optimized, and the optimized three-dimensional generative adversarial network is inverted using the generative adversarial network inversion, thereby generating a 360-degree head image consistent with the input image.

4. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In the step S2, flattening the three-dimensional Gaussian specifically includes: determining the position, rotation and size attributes of the three-dimensional Gaussian according to the vertex positions of the facet where the three-dimensional Gaussian is located: p=α1v1+α2v2+α3v3 S=diag(s1,s2,s3) Where p is the position of the three-dimensional Gaussian, α1, α2, α3 are trainable parameters and α1+α2+α3=1 and are trainable parameters, v1, v2, v3 are the coordinates of the three vertices of the patch where the three-dimensional Gaussian is located; the rotation matrix of the three-dimensional Gaussian R=[r1, r2, r3], r1 is the normal vector of the patch, r2 is the unit vector from the first vertex of the patch to the center of gravity of the patch, and r3 is the unit vector orthogonal to r1 and r2; S is the size matrix of the three-dimensional Gaussian, s1=ε, ε=1e-5, s2=||mean(v1,v2,v3)-v1||, s3=v2·r3.

5. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In step S2, the loss function L includes the L1 loss function and the SSIM loss function L SSIM : L=(1-λ)L1+λL SSIM Where λ=0.

2.

6. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In step S3, the vertex displacement of the optimized FLAME model is the displacement along the normal vector of the corresponding vertex.

7. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In step S3, the training loss function L includes the L1 loss function, the SSIM loss function L SSIM , Laplace loss function L Laplasian And the normal consistency loss function L normal : L=(1-λ)L1+λL SSIM +λ1L Laplasian +λ2L normal Among them, λ=0.2, λ1=λ2=0.05~0.

2.

8. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In step S4, the initialized texture map is a picture with a resolution of 512*512, and the differentiable renderer is Nvdiffrast.

9. The method for supporting multi-modal generation of a 360-degree drivable three-dimensional avatar according to claim 1, characterized in that: In step S5, the image generated by Gaussian rendering is subjected to loss calculation with the head image generated in step S1 to optimize the rotation and size parameters of the three-dimensional Gaussian. The loss function L includes the L1 loss function, the SSIM loss function L SSIM And the opacity loss function L opacity to prevent the Gaussian from becoming transparent: L=(1-λ)L1+λL SSIM +λ3L opacity Among them, λ=0.2 and λ3=0.05~0.2.

Citation Information

Cited By

  • Object-level data acquisition and reconstruction method based on three-dimensional Gaussian splashing

    CN122176169A