Method and apparatus for reconstructing virtual avatar of human body, terminal and storage medium

By combining 3D Gaussian sputtering deep learning with neural networks and pre-trained generative models, the problems of high data acquisition cost, low rendering quality, and insufficient computational efficiency in human virtual avatar reconstruction are solved, enabling the rapid generation of realistic virtual human models suitable for applications such as virtual reality and augmented reality.

CN122115697APending Publication Date: 2026-05-29BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies for human virtual avatar reconstruction suffer from high data acquisition costs, long training times, low rendering quality, and insufficient computational efficiency, making them particularly difficult to meet the needs of applications requiring rapid feedback.

Method used

We employ a 3D Gaussian sputtering deep learning method based on a small number of images, combining neural networks and pre-trained generative models. By acquiring multiple image, depth map, and virtual human datasets, we utilize a trained neural network model for 3D reconstruction and rendering. Gaussian mixture models and various loss functions are introduced to optimize model performance.

Benefits of technology

It achieves a significant improvement in computational efficiency while ensuring rendering quality, and can quickly generate realistic virtual human models, suitable for application scenarios that require rapid feedback, such as virtual reality, augmented reality, and games.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115697A_ABST
    Figure CN122115697A_ABST
Patent Text Reader

Abstract

The present disclosure provides a reconstruction method and device of a human virtual avatar, a terminal and a storage medium. The reconstruction method of the human virtual avatar comprises: acquiring a plurality of images including a human body and a virtual human data set; acquiring a depth map from the plurality of images; and performing three-dimensional reconstruction and rendering based on the plurality of images, the depth map and the virtual human data set by using a trained neural network model to obtain a human three-dimensional model and a corresponding rendered image. The present disclosure can reconstruct the human virtual avatar by using a small number of images, thereby simplifying the data acquisition process. In addition, the reconstruction method of the present disclosure can ensure high calculation efficiency while ensuring rendering quality, and can be used for real-time interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of information technology, and in particular to methods, apparatus, terminals and storage media for reconstructing virtual human avatars. Background Technology

[0002] Research on human virtual avatar technology has significant academic and applied implications. With the development of technologies such as augmented reality (AR), virtual reality (VR), and the metaverse, creating high-quality digital human avatars has become a key factor in enhancing user experience and interactive effects. This not only provides users with a more immersive experience but also plays an important role in fields such as fashion design and virtual social interaction. However, one of the main challenges facing this field is data acquisition. Traditional reconstruction methods typically rely on expensive multi-view data, which is not only costly to acquire but also difficult to scale up in practical applications. Recent research has begun to favor more readily available data sources, such as a small number of RGB images. While these methods reduce costs to some extent, they still suffer from long training times. Furthermore, a small number of images often fails to provide sufficient information. Additionally, in the 3D reconstruction of human virtual avatars, ensuring both rendering quality and computational efficiency is a pressing issue that needs to be addressed. Summary of the Invention

[0003] To address the existing problems, this disclosure provides a method, apparatus, terminal, and storage medium for reconstructing a human virtual avatar.

[0004] The following technical solution is adopted in this disclosure.

[0005] The embodiments of this disclosure provide a method for reconstructing a human virtual avatar, the method comprising: acquiring multiple images including a human body and a virtual human body dataset; acquiring depth maps from the multiple images; and performing three-dimensional reconstruction and rendering using a trained neural network model based on the multiple images, the depth maps and the virtual human body dataset, to obtain a three-dimensional human body model and a corresponding rendered image.

[0006] Another embodiment of this disclosure provides a reconstruction apparatus for a human virtual avatar. The processing apparatus includes: an image and dataset acquisition module configured to acquire multiple images including a human body and a virtual human body dataset; a depth map acquisition module configured to acquire depth maps from the multiple images; and a model building and rendering module configured to perform three-dimensional reconstruction and rendering based on the multiple images, the depth maps, and the virtual human body dataset using a trained neural network model to obtain a three-dimensional human body model and a corresponding rendered image.

[0007] In some embodiments, this disclosure provides a terminal, including: at least one memory and at least one processor; wherein the memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above-described method for reconstructing a human virtual avatar.

[0008] In some embodiments, this disclosure provides a storage medium for storing program code for executing the above-described method for reconstructing a human virtual avatar.

[0009] This disclosure utilizes a trained neural network model to perform 3D reconstruction and rendering based on multiple images, depth maps, and virtual human datasets, resulting in a 3D human body model and corresponding rendered images. The reconstruction of a virtual human avatar can be performed using only a few images, simplifying the data acquisition process. Furthermore, the reconstruction method of this disclosure ensures high computational efficiency while maintaining rendering quality, thus enabling real-time interaction. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a flowchart of a method for reconstructing a human virtual avatar according to an embodiment of the present disclosure.

[0012] Figure 2 A structural diagram of a neural network model according to some embodiments is shown.

[0013] Figure 3 This is a partial module of a human virtual avatar reconstruction device according to another embodiment of this disclosure.

[0014] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be understood that the various steps described in the method embodiments of this disclosure can be performed in sequence and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0017] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the use of the word "a" in this disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as "one or more" unless otherwise expressly indicated in the context.

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] Parametric modeling methods based on a limited number of images (such as the Standard Human Model for Perception and Learning (SMPL) and (SMPL-X)) are among the most widely used human reconstruction techniques. They can generate accurate 3D models by parametrically configuring the shape and pose of the human body, and quickly fit the human body in an image by optimizing these parameters. Their advantages include speed and compatibility with traditional rendering pipelines, making them suitable for real-time applications. However, the rendering of these models is primarily mesh-based, resulting in limited expressive power and lower rendering quality. To overcome these limitations, combining them with neural rendering or voxel representation can be considered to improve rendering quality and expressiveness.

[0022] Implicit methods based on a small number of images (such as Depth Tetrahedral Moving Mesh (DMTet) and Neural Radiation Field (NeRF)) have made significant progress in the field of 3D reconstruction in recent years. NeRF represents the color and density of a scene through implicit functions, enabling the generation of highly realistic 3D renderings. It models the lighting in a scene by training a network, allowing for the rendering of high-quality images from any viewpoint. Furthermore, NeRF effectively captures details and lighting variations in complex scenes, resulting in more realistic and detailed 3D scenes. This method is applicable not only to static scenes but also to dynamic effects, providing strong support for applications such as virtual reality and augmented reality. DMTet is an advanced algorithm specifically designed to generate high-quality 3D meshes. It represents 3D shapes as tetrahedral meshes and utilizes dense matching techniques to optimize the quality and geometric details of the mesh. DMTet's core advantage lies in its ability to generate meshes with high fidelity and low geometric errors, making it outstanding in fields such as 3D reconstruction, computer graphics, and computational geometry.

[0023] However, while these two implicit methods can produce high-quality rendering results, they still face significant challenges in terms of time cost. Generating a single instance typically takes about 4.5 hours, which is clearly a major bottleneck for applications requiring rapid feedback or frequent updates. This high time consumption not only limits their widespread use in highly interactive virtual reality and augmented reality applications but also places higher demands on the efficient allocation of resources. Therefore, how to reduce computation time while maintaining high-quality rendering effects will become an important direction for future research.

[0024] Pre-trained generative large models are a class of deep learning models that are pre-trained on large-scale datasets and have the ability to generate new data. They are widely used in various generative tasks such as text and images. These models typically employ self-supervised or unsupervised learning methods to learn features and patterns from data without explicit labels, and then achieve generative tasks through a small amount of task-specific fine-tuning.

[0025] DINAR aims to reconstruct virtual human avatars from input images, primarily consisting of two components: a generative model and a refinement model. The generative model reconstructs the neural texture of the input image and synthesizes a rendered image through neural rendering. This model takes an RGB image and a parameterized SMPL-X human model as input, combining the SMPLifyX method and segmentation loss to optimize human contour matching. The generated neural texture includes textures generated by StyleGAN2 and textures sampled from the input image, resulting in a final texture dimension of 256×256×21. Furthermore, the refinement model is trained based on a denoising diffusion probability model (DDPM) to recover missing human regions in the input image. During training, the model is optimized by minimizing difference loss, perceptual loss, and adversarial loss. However, this method suffers from lower rendering quality and a bias towards generating tight-fitting human figures, which may lead to poor performance when handling complex poses or human figures of different body types, thus limiting its application scope.

[0026] HAVEFUN proposes an implicit method for reconstructing human virtual avatars, aiming to support free-view rendering and free-pose animation. Given a small number of RGB images, the method generates a 3D representation that includes a triangular mesh, texture field, skinning weights, and a hybrid shape for facial expressions. The core idea is to utilize DMTet, combined with references and guidance from a small number of images for training. The method involves initialization through a hybrid representation, allowing arbitrary human figures to be represented by defined tetrahedral meshes and learnable vertex displacements. Then, visual features and camera viewpoints are used as conditions, and Zero123 is used to supervise the generation of images from different viewpoints. Finally, dynamic poses are handled through parametric meshing and skinning mechanisms. While this method performs well in generating high-quality images and maintaining consistency across multiple viewpoints, it suffers from time efficiency limitations. The optimization process requires significant computation time, which restricts its application in scenarios requiring real-time feedback.

[0027] ihuman proposes a simple and efficient method for creating driveable 3D human virtual avatars from monocular video. ihuman leverages the efficiency of Gaussian sputtering to model the dynamic 3D geometry and appearance of the human body. This work binds 3D Gaussian sputtering to corresponding triangular patches of a body template, achieving an accurate and efficient method for human body modeling. However, the input to this work is monocular video, which is not as readily available to users as a small number of images.

[0028] HumanSplat proposes a Gaussian sputtering method for predicting static 3D human bodies from a limited number of images. Specifically, HumanSplat includes a 2D multi-view diffusion model and a transformation model with prior knowledge of human structure. It integrates geometric priors and semantic features, enabling high-fidelity texture modeling. However, HumanSplat can only generate static 3D human body models and cannot be directly animated, thus failing to meet the requirements of applications requiring dynamic interaction, such as virtual avatars and virtual social interaction.

[0029] Therefore, although DINAR achieves the reconstruction of human virtual avatars, its rendering quality is low, especially when generating complex poses, which may lead to distortion or unnatural effects. This low-quality rendering affects the user experience, making the generated animations appear unrealistic in practical applications. Furthermore, this method tends to generate tight-fitting human structures, lacking adaptability to different body shapes and poses. This results in poor performance in application scenarios that require handling various body shapes or dynamic poses, limiting its application scope.

[0030] While HAVEFUN excels at generating high-quality images, its optimization process requires significant computation time. This is especially true when dealing with complex reference images, where computational demands increase dramatically, slowing down model training and rendering. This computational burden makes implicit methods unsuitable for scenarios requiring rapid feedback, such as fast animation generation and games. The inability to meet user expectations for quick responses limits the practical application of this technology in dynamic interactions.

[0031] Existing explicit methods for generating human virtual avatars suffer from relatively low rendering quality and are limited to tight-fitting human structures. This disclosure enhances rendering quality by introducing a more refined optimization mechanism, ensuring the robustness of the generated virtual character across different dynamic poses to meet the needs of various application scenarios. Implicit methods suffer from insufficient time efficiency during the optimization process, limiting their use in some applications. This disclosure employs more efficient algorithms and model architectures to reduce the computation time required for training and rendering, ensuring efficient and rapid feedback capabilities even when handling complex scenes. This allows the technology to be better applied to fields requiring fast responses, such as virtual reality and games. Therefore, this disclosure pursues both high rendering quality and computational efficiency.

[0032] Figure 1 A flowchart of a method for reconstructing a human virtual avatar according to embodiments of the present disclosure is provided. Figure 2A structural diagram of a neural network model according to some embodiments is shown. The method for reconstructing a human virtual avatar according to this disclosure may include step S101, acquiring multiple images including a human body and a virtual human body dataset. In some embodiments, the multiple images may be images captured by a typical RGB camera, and the number of multiple images is typically 3 to 10, but this disclosure is not limited thereto. In some embodiments, the virtual human body dataset is existing, collected and organized human-related data, such as human-related shape parameters and pose parameters, which can be used to train machine learning models, perform computer vision tasks, conduct biomechanical research, etc. In some embodiments, the virtual human body dataset can be obtained from this existing data through random sampling. Thus, virtual human body datasets representing different shapes and poses can be used to supervise the training of models, enabling them to better generalize to real-world scenes, thereby enriching the poses of the reconstructed 3D human body model and improving generalization.

[0033] In some embodiments, the method of this disclosure may further include step S102, acquiring depth maps from multiple images. In some embodiments, depth maps can be estimated from the multiple images using various existing pre-trained large models. In some embodiments, the input multiple images, depth maps, virtual human datasets, etc., are normalized to unify the data format and adapt it to the input requirements of the model, including adjusting the resolution, normalizing pixel values, and depth values, etc.

[0034] In some embodiments, the method of this disclosure may further include step S103, which involves performing 3D reconstruction and rendering using a trained neural network model based on multiple images, depth maps, and a virtual human body dataset to obtain a 3D human body model and a corresponding rendered image. In some embodiments, this disclosure obtains a 3D human body model and a corresponding rendered image by performing 3D reconstruction and rendering using a trained neural network model based on multiple input images, depth maps obtained from the multiple images, and a virtual human body dataset.

[0035] This disclosure simplifies the data acquisition process by using a few images instead of providing videos, and by using a virtual human dataset, it can provide images from different perspectives, improving rendering quality and generalization.

[0036] In some embodiments, reference Figure 2Based on multiple images, depth maps, and virtual human datasets, a trained neural network model is used for 3D reconstruction and rendering to obtain a 3D human body model and corresponding rendered images. The process includes: acquiring shape parameters and first pose parameters of the human body from multiple images; generating an initial 3D human body template mesh based on the shape parameters and first pose parameters; generating initial Gaussian mixture model parameters based on the initial 3D human body template mesh; generating the density field and color field of the 3D scene based on the initial Gaussian mixture model parameters; generating deformed Gaussian mixture model parameters based on the initial Gaussian mixture model parameters, the first pose parameters, and second pose parameters from the virtual human body dataset; and obtaining the 3D human body model and corresponding rendered images based on the deformed Gaussian mixture model parameters, the external and internal parameters of the camera used to acquire the multiple images.

[0037] In some embodiments, existing SMPL models can be used for parameter estimation to estimate the human body's pose parameters θ and shape parameters β from multiple input images. These parameters can be used to construct a three-dimensional model of the human body and generate an initial three-dimensional human body template mesh.

[0038] In some embodiments, generating initial Gaussian mixture model parameters based on an initial 3D human body template mesh includes: using the vertices of the initial 3D human body template mesh as the point cloud positions of the Gaussian mixture model, and using the normal directions of the mesh's faces as the depth directions to initialize the parameters of the Gaussian mixture model, thereby generating the initial Gaussian mixture model parameters. In some embodiments, as a Gaussian initialization module in a driveable Gaussian sputtering module, the input is the initial 3D human body template mesh and the skeleton structure B estimated by SMPL and its corresponding weights W, and the output is the 3D human body represented by the initial Gaussian mixture model parameters G{μ, R, σ, c}. In some embodiments, the purpose of the Gaussian initialization module is to convert the input 3D template into a Gaussian representation suitable for subsequent deformation and rendering. This process provides basic shape and appearance information for subsequent modules. Specifically, a predefined template mesh M is used to initialize the Gaussian mixture model parameters, that is, the predefined template mesh vertex Vc is used as the position μ of the Gaussian mixture model, the normal directions of the mesh's faces are used as the depth directions R, and the density σ and color c are randomly initialized; in addition, the joint information and weights W of the predefined skeleton B are used to determine the skinning weights of each Gaussian component.

[0039] In some embodiments, generating the density field and color field of a 3D scene based on initial Gaussian mixture model parameters includes: determining the Gaussian density for each Gaussian component; determining the color of the corresponding Gaussian component based on the Gaussian density; and determining the density field and color field of the 3D scene based on the Gaussian density and color of each Gaussian component. This is achieved through... Figure 2This is implemented using the shape and appearance representation module, which is responsible for explicitly representing the shape and appearance of the human body using a Gaussian model, calculating the density and color information at a given location, thereby providing the necessary visual features for subsequent rendering. Specifically, for each Gaussian component i, its density contribution is calculated; the color of each Gaussian component is calculated based on the Gaussian density; and then the density and color of the entire scene are obtained by summing the contributions of all Gaussian components.

[0040] In some embodiments, generating deformed Gaussian mixture model parameters based on initial Gaussian mixture model parameters, first pose parameters, and second pose parameters from a virtual human dataset includes: utilizing learnable skin weights such that each Gaussian component deforms using the first and second pose parameters; determining the transformation of the Gaussian component and determining the deformed mean and rotation to determine the deformed Gaussian mixture model parameters. This is achieved through... Figure 2 This is achieved using the Deformable Gaussian module. The Deformable Gaussian module achieves the deformation of the Gaussian model through forward skin deformation and learnable bone adjustment, capturing the dynamic changes of the character and clothing. The transformation of each Gaussian component includes position μ(i) and rotation R(i). Specifically, learnable skin weights W are used to enable each Gaussian component to deform using a human parametric template; the transformation of the Gaussian component is calculated, along with the mean and rotation after deformation; the deformation capability is enhanced by introducing a second pose parameter or latent bone.

[0041] In some embodiments, obtaining a 3D human body model and corresponding rendered image based on deformed Gaussian mixture model parameters, extrinsic and intrinsic parameters of the camera acquiring multiple images includes: converting the mean, covariance, color, and density of the 3D Gaussian mixture model into a 2D representation. This is achieved through... Figure 2 This is achieved through a Gaussian rendering module. In some embodiments, the Gaussian rendering module is responsible for rendering a 3D Gaussian mixture model into a 2D image, realizing the conversion from a 3D scene to a 2D image through differentiable rendering technology. Specifically, it converts the mean, covariance, color, and density of the 3D Gaussian mixture into a 2D representation. Then, it can be rendered using a commonly used sorted color accumulation method. In this way, the rendered output, i.e., the rendered 2D image, can be obtained.

[0042] This disclosure employs 3D Gaussian sputtering, an emerging method for rendering 3D scenes using Gaussian sputtering technology. It combines the advantages of Gaussian distributions and point cloud data, achieving efficient, realistic, and continuous 3D object and scene representations by rendering point clouds as Gaussian specks in 3D space. In 3D rendering, traditional mesh models require triangles to represent 3D surfaces. While accurate, this method is inefficient when handling very complex or sparse 3D data (such as point clouds). 3D Gaussian sputtering provides an efficient solution to this problem by expanding each point in the point cloud into a Gaussian distribution, generating continuous surfaces and textures through the superposition of these distributions. The core idea of ​​3D Gaussian sputtering is to transform the point cloud into Gaussian specks, where each point is not just a spatial coordinate but is expanded into a 3D "spot" in space through a Gaussian function; rendering is achieved through the superposition of these specks: multiple Gaussian specks superimposed on each other can generate continuous surfaces and capture the details and lighting effects of object surfaces.

[0043] Compared to traditional triangular meshing methods, 3D Gaussian sputtering directly processes point clouds, avoiding complex mesh generation and ray tracing processes, enabling high-quality rendering with low computational resources. 3D Gaussian sputtering offers significant advantages in point cloud rendering, particularly in efficiency, quality, and applicability. It can render directly from point cloud data without generating complex triangular meshes. This significantly improves rendering speed when processing large-scale point cloud data, making it particularly suitable for real-time rendering applications such as virtual reality (VR), augmented reality (AR), and game engines. Furthermore, 3D Gaussian sputtering can handle both sparse and dense point cloud data, generating high-quality rendering results without excessive reliance on high-precision point cloud density. Additionally, 3D Gaussian sputtering handles objects with complex shapes and textures well. When processing objects with complex geometry or non-rigid objects (such as clothing and the human body), it can generate realistic dynamic effects without the limitations of traditional polygonal meshing methods. This disclosure significantly improves computational and rendering efficiency by employing a Gaussian model.

[0044] In some embodiments, during the training of the neural network model, reconstruction loss function, depth consistency loss function, fractional distillation sampling generation loss function, and color gamut consistency loss function are used for model supervision. In some embodiments, the model supervision module optimizes model performance by calculating multiple losses to ensure that the generated image matches the input image in multiple dimensions, thereby improving rendering quality and generalization under different poses. In some embodiments, reconstruction loss measures the difference between the Gaussian-rendered original viewpoint image Irender and the original image I_real (i.e., multiple input images) by using L1 loss and structural similarity index (SSIM) loss. L1 loss focuses on absolute differences at the pixel level, while SSIM considers the structural information of the image; the combination of the two helps to improve image quality. In some embodiments, depth consistency loss evaluates the model's accuracy in depth information by comparing the Gaussian-rendered depth map D_render with the original depth map D_real. Depth consistency loss ensures that the model performs well not only in two-dimensional space but also captures the correct structure in three-dimensional space. In some embodiments, the fractional distillation sampling (SDS) generation loss aims to leverage the difference between the input image I_real and the Gaussian-rendered images I_diff from different viewpoints to ensure that the model maintains a consistent spatial structure during generation, particularly producing natural and coherent results across different poses. In some embodiments, the color gamut consistency loss calculates the consistency in color space between the input image I_real and the Gaussian-rendered images I_diff from different viewpoints. This loss helps ensure that the generated image matches the real image in color representation, enhancing not only the visual appeal of the image but also the realism of the model.

[0045] By comprehensively calculating these losses, the model supervision module effectively guides model optimization, ensuring that the final generated image maintains consistency with the original image in terms of structure, depth, and color, thereby improving the quality of visual effects and generalization.

[0046] In some embodiments, SDS generation loss is a loss function used to train generative models and is widely applied in the field of 3D generation, such as tasks that generate 3D models from natural language descriptions or images. SDS loss guides the consistency between generated 3D data and target text or image descriptions by incorporating a pre-trained diffusion model. This method was initially designed to optimize latent representations in 3D shape generation but can be extended to other generative tasks. The core idea of ​​SDS loss is to distill the knowledge of a pre-trained diffusion model into the generative model through fractional distillation. Specifically, pre-trained diffusion models (such as stable diffusion) typically perform well on 2D image generation tasks; therefore, in 3D generation tasks, SDS loss utilizes the gradient information of the diffusion model to ensure that the generated 3D objects are consistent with the target description (e.g., natural language or images). In some embodiments, the SDS generation loss process can be summarized in the following steps: Use of a pre-trained diffusion model: First, a diffusion model already trained on a 2D task (e.g., a text-image diffusion model) is used as a teacher model. This teacher model provides a score in a high-dimensional space, representing the degree of matching between the target description and the generated result; Generation model and optimization: In a 3D generation task, there is a generation model that generates 3D objects based on certain latent parameters (such as the shape and texture of the 3D model). Through SDS loss, the generation model learns how to adjust its parameters to maximize consistency with the target description; Fraction distillation: The diffusion model generates a fractional gradient during its sampling process, representing the degree of deviation between the generated 2D representation and the target text description or image. SDS loss guides the 3D generation model to adjust along this gradient, making the generated 3D shape more consistent with the target description in the 2D projection of the diffusion model; Loss calculation: SDS loss calculates the loss value by comparing the similarity between the 2D projection of the generated 3D object and the target description from different viewpoints. Through backpropagation, the parameters of the generation model are updated, making the generated result better match the target description. SDS loss can be efficiently combined with 3D generation tasks, enabling generative models to generate realistic 3D models based on text. Furthermore, SDS fully utilizes the knowledge already learned in pre-trained diffusion models, avoiding the need to retrain numerous complex models and improving training efficiency. In addition, SDS loss compares the generated 3D model with the target description from multiple perspectives, ensuring that the generated 3D model meets requirements from different viewpoints, thus generating higher-quality 3D shapes.

[0047] In some embodiments, a deep learning framework (PyTorch) is used to construct and train the neural network model disclosed herein. The optimizer employs a common optimization algorithm (Adam) to adjust the model's weights, gradually reducing errors. Learning rate settings and weight decay strategies are adjusted to ensure the model does not overfit during convergence. In some embodiments, the neural network model of this disclosure takes as input multiple preprocessed RGB images, depth maps, and corresponding virtual human datasets (pose and shape), and outputs a reconstructed 3D human model and a corresponding rendered image generated by the model. The goal of the neural network model of this disclosure is to generate rendered images that are consistent with the input data and can generalize across different poses. The neural network model of this disclosure gradually improves the accuracy of its output by optimizing the errors between modules (such as depth consistency errors), and optimizes model parameters through backpropagation during training, making the generated results gradually approximate real-world scenes.

[0048] In some embodiments, the neural network model of this disclosure can be subjected to algorithm testing, including model validation, performance evaluation, result display, and iterative optimization. In model validation, the model's generalization ability is verified by inputting unseen test data, ensuring good generalization even outside the training set. During validation, the model's ability to handle changes in human pose, shape, and viewpoint is examined, and its performance in different scenarios is evaluated. In performance evaluation, various quantitative metrics (such as depth) are used to evaluate the model's performance. Figure 1 The model's performance is evaluated based on factors such as consistency and reconstruction accuracy. Furthermore, qualitative evaluation is combined with comparisons of the input and generated 3D models and rendering results to ensure visual quality. The results presentation showcases the model's generated 3D human reconstruction and rendering effects, comparing them with the input image and real-world scenes to intuitively demonstrate the algorithm's effectiveness and performance. Multi-view and multi-pose result presentations help comprehensively evaluate the model's applicability in practical applications. In iterative optimization, the model architecture or training strategy is further adjusted based on test results and performance evaluation feedback. Through iterative optimization, the model's performance is continuously improved, ensuring its stability and reliability in practical applications.

[0049] The reconstruction method disclosed herein is highly practical, with one of its biggest highlights being its ability to be directly applied to real-world scenarios based on a small number of images without any post-processing steps. This means users can directly obtain results, improving overall efficiency. Especially in applications requiring rapid feedback, such as virtual avatars and virtual social interactions, it can quickly generate realistic and usable virtual characters. In terms of algorithm design, this disclosure introduces 3D Gaussian sputtering and SMPL-X (Scalable Statistical Human Model) techniques. It can accurately capture and realistically reconstruct 3D human bodies in various everyday scenarios. This accuracy and robustness enable the algorithm to work stably and is highly adaptable, particularly suitable for environments requiring rapid interaction, such as virtual reality and augmented reality. Furthermore, this disclosure has deeply optimized the design of the loss function, improving the realism of the reconstruction results from multiple dimensions. Specifically, the loss functions include the following: SDS generation loss: ensuring the model maintains a consistent spatial structure during generation, especially producing natural and coherent results in different pose scenarios; Depth consistency loss: by introducing consistency constraints on depth information, ensuring the generated 3D human body has the same depth structure as the real image, improving the accuracy of 3D reconstruction; Color gamut consistency loss: further optimizing the color distribution, making the generated image more natural in color representation and avoiding color cast; Reconstruction loss: ensuring the model is highly consistent with the original data in pose and shape reconstruction, avoiding damage to the overall reconstruction effect due to over-optimization of a certain indicator. Through the comprehensive optimization of these loss functions, the model is improved at all levels, ensuring high realism of the generated images.

[0050] The optimized algorithm disclosed herein can generate highly realistic 3D images and poses in various environments. This high realism allows the algorithm to be widely applied in fields requiring high visual fidelity, such as virtual avatars, virtual social interaction, and digital entertainment. Furthermore, the algorithm exhibits excellent robustness when handling various complex pose changes; even with extreme movements, the model can stably output high-precision reconstruction results. This robustness ensures the reliability of the algorithm in real-world applications, particularly in scenarios requiring dynamic feedback, such as motion capture, virtual social interaction, and motion analysis, providing an excellent user experience. Through these improvements, the reconstruction method disclosed herein shows significant enhancements in practicality, algorithmic detail, and generation quality.

[0051] This disclosure proposes a novel deep learning framework for 3D-driven Gaussian sputtering of human virtual avatars based on a limited number of images. This framework combines the advantages of implicit and explicit methods, significantly improving computational efficiency while maintaining rendering quality. 3D Gaussian sputtering transforms the 3D human template and pose information into a Gaussian mixture representation suitable for deformation and rendering, enabling the framework to quickly adapt to different pose variations while maintaining high rendering speed. Furthermore, by introducing a pre-trained generative large model prior and color consistency loss, this disclosure addresses the sparsity problem of appearance information in a limited number of images. By combining the advantages of generative models, the ability to capture image details is enhanced during appearance reconstruction, while the color consistency loss ensures the color uniformity of the reconstructed human body at different angles, thus greatly improving the quality of appearance reconstruction. In addition, existing reconstruction methods have poor generalization ability, especially when facing different poses, where the reconstruction effect is not stable enough and cannot guarantee consistency. This disclosure significantly improves the generalization ability of the algorithm by constructing a virtual human dataset and introducing a depth consistency loss for different poses. The diversity of predefined virtual human datasets provides abundant training samples, while the depth consistency loss ensures the stability and consistency of the reconstructed human body across different poses through cross-pose depth constraints. This enables the reconstruction method to better adapt to various complex pose scenarios.

[0052] Therefore, the reconstruction method disclosed herein does not require any post-processing and can be directly used in real-world scenarios; by combining 3D Gaussian sputtering with deep learning, the training time of the model is greatly reduced; by combining pre-trained generative large model priors with color consistency loss, the reconstruction quality of the appearance is greatly improved; and by combining predefined virtual human datasets with depth consistency loss, the generalization ability for different poses is greatly improved.

[0053] Embodiments of this disclosure also provide a human virtual avatar reconstruction device 400. Figure 3 A human virtual avatar reconstruction apparatus 400 according to some embodiments is shown. The human virtual avatar reconstruction apparatus 400 includes an image and dataset acquisition module 401, a depth map acquisition module 402, and a model building and rendering module 403. In some embodiments, the image and dataset acquisition module 401 is configured to acquire multiple images including a human body and a virtual human body dataset. In some embodiments, the depth map acquisition module 402 is configured to acquire depth maps from multiple images. In some embodiments, the model building and rendering module 403 is configured to perform 3D reconstruction and rendering based on the multiple images, depth maps, and the virtual human body dataset using a trained neural network model to obtain a 3D human body model and a corresponding rendered image.

[0054] It should be understood that the description of the method for reconstructing a human virtual avatar also applies to the reconstruction device 400 for human virtual avatars described herein, but will not be described in detail here for simplicity.

[0055] In some embodiments, based on multiple images, depth maps, and a virtual human dataset, 3D reconstruction and rendering are performed using a trained neural network model to obtain a 3D human body model and a corresponding rendered image. This includes: acquiring shape parameters and first pose parameters of the human body from multiple images; generating an initial 3D human body template mesh based on the shape parameters and first pose parameters; generating initial Gaussian mixture model parameters based on the initial 3D human body template mesh; generating a density field and color field of a 3D scene based on the initial Gaussian mixture model parameters; generating deformed Gaussian mixture model parameters based on the initial Gaussian mixture model parameters, the first pose parameters, and a second pose parameter from the virtual human dataset; and obtaining the 3D human body model and the corresponding rendered image based on the deformed Gaussian mixture model parameters, the external parameters and internal parameters of the camera that acquired the multiple images. In some embodiments, generating initial Gaussian mixture model parameters based on the initial 3D human body template mesh includes: using the vertices of the initial 3D human body template mesh as the point cloud positions of the Gaussian mixture model, and using the normal direction of the mesh's facets as the depth direction to initialize the parameters of the Gaussian mixture model, thereby generating the initial Gaussian mixture model parameters. In some embodiments, generating the density field and color field of a 3D scene based on initial Gaussian mixture model parameters includes: determining the Gaussian density for each Gaussian component; determining the color of the corresponding Gaussian component based on the Gaussian density; and determining the density field and color field of the 3D scene based on the Gaussian density and color of each Gaussian component. In some embodiments, generating deformed Gaussian mixture model parameters based on initial Gaussian mixture model parameters, first pose parameters, and second pose parameters from a virtual human dataset includes: utilizing learnable skinning weights to deform each Gaussian component using the first and second pose parameters; determining the transformation of the Gaussian component and determining the deformed mean and rotation to determine the deformed Gaussian mixture model parameters. In some embodiments, obtaining a 3D human body model and corresponding rendered images based on the deformed Gaussian mixture model parameters, extrinsic and intrinsic parameters of a camera acquiring multiple images includes: converting the mean, covariance, color, and density of the 3D Gaussian model into a 2D representation. In some embodiments, during the training of the neural network model, a reconstruction loss function, a depth consistency loss function, a fractional distillation sampling generation loss function, and a color gamut consistency loss function are used for model supervision.

[0056] Furthermore, this disclosure also provides a terminal, comprising: at least one memory and at least one processor; wherein the memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above-described method for reconstructing a human virtual avatar.

[0057] In addition, this disclosure also provides a computer storage medium storing program code for executing the above-described method for reconstructing a human virtual avatar.

[0058] The above describes the method and apparatus for reconstructing a virtual human avatar according to the present disclosure, based on embodiments and application examples. Furthermore, the present disclosure also provides a terminal and a storage medium, which are described below.

[0059] The following is for reference. Figure 4 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 500 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0060] like Figure 4 As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0061] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0062] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0063] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0064] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0065] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0066] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure.

[0067] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0069] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0070] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0071] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0072] According to one or more embodiments of this disclosure, a method for reconstructing a human virtual avatar is provided. The method includes: acquiring multiple images including a human body and a virtual human body dataset; acquiring depth maps from the multiple images; and performing three-dimensional reconstruction and rendering using a trained neural network model based on the multiple images, the depth maps, and the virtual human body dataset to obtain a three-dimensional human body model and a corresponding rendered image.

[0073] According to one or more embodiments of this disclosure, based on the plurality of images, the depth map, and the virtual human dataset, 3D reconstruction and rendering are performed using a trained neural network model to obtain a 3D human body model and a corresponding rendered image, including: obtaining shape parameters and a first pose parameter of the human body from the plurality of images; generating an initial 3D human body template mesh based on the shape parameters and the first pose parameter; generating initial Gaussian mixture model parameters based on the initial 3D human body template mesh; generating a density field and a color field of the 3D scene based on the initial Gaussian mixture model parameters; generating deformed Gaussian mixture model parameters based on the initial Gaussian mixture model parameters, the first pose parameter, and a second pose parameter from the virtual human dataset; and obtaining a 3D human body model and a corresponding rendered image based on the deformed Gaussian mixture model parameters, the external parameters and internal parameters of the camera used to obtain the plurality of images.

[0074] According to one or more embodiments of this disclosure, generating initial Gaussian mixture model parameters based on the initial three-dimensional human body template mesh includes: using the vertices of the initial three-dimensional human body template mesh as the point cloud positions of the Gaussian mixture model, using the normal direction of the mesh's facets as the depth direction to initialize the parameters of the Gaussian mixture model, and generating initial Gaussian mixture model parameters.

[0075] According to one or more embodiments of this disclosure, generating the density field and color field of a three-dimensional scene based on the initial Gaussian mixture model parameters includes: determining the Gaussian density for each Gaussian component; determining the color of the corresponding Gaussian component based on the Gaussian density; and determining the density field and color field of the three-dimensional scene based on the Gaussian density and color of each Gaussian component.

[0076] According to one or more embodiments of this disclosure, generating deformed Gaussian mixture model parameters based on the initial Gaussian mixture model parameters, the first pose parameters, and the second pose parameters from the virtual human dataset includes: using learnable skin weights to deform each Gaussian component using the first pose parameters and the second pose parameters; determining the transformation of the Gaussian component and determining the deformed mean and rotation to determine the deformed Gaussian mixture model parameters.

[0077] According to one or more embodiments of this disclosure, obtaining a three-dimensional human body model and a corresponding rendered image based on the deformed Gaussian mixture model parameters, the external parameters and internal parameters of the camera that acquires the plurality of images includes: converting the mean, covariance, color and density of the three-dimensional Gaussian mixture model into a two-dimensional representation.

[0078] According to one or more embodiments of this disclosure, during the training of a neural network model, a reconstruction loss function, a depth consistency loss function, a fractional distillation sampling generation loss function, and a color gamut consistency loss function are used for model supervision.

[0079] According to one or more embodiments of this disclosure, a human virtual avatar reconstruction apparatus is provided, characterized in that the human virtual avatar reconstruction apparatus includes: an image and dataset acquisition module configured to acquire multiple images including a human body and a virtual human body dataset; a depth map acquisition module configured to acquire depth maps from the multiple images; and a model building and rendering module configured to perform three-dimensional reconstruction and rendering based on the multiple images, the depth maps, and the virtual human body dataset using a trained neural network model to obtain a three-dimensional human body model and a corresponding rendered image.

[0080] According to one or more embodiments of the present disclosure, a terminal is provided, comprising: at least one memory and at least one processor; wherein the at least one memory is used to store program code, and the at least one processor is used to invoke the program code stored in the at least one memory to execute the method described in any one of the above descriptions.

[0081] According to one or more embodiments of the present disclosure, a storage medium is provided for storing program code for performing the methods described above.

[0082] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0083] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0084] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for reconstructing a virtual human avatar, characterized in that, The method for reconstructing the human virtual avatar includes: Acquire multiple images and virtual human body datasets, including those of the human body; Obtain depth maps from the plurality of images; Based on the multiple images, the depth map, and the virtual human dataset, a trained neural network model is used for 3D reconstruction and rendering to obtain a 3D human model and corresponding rendered images.

2. The method for reconstructing a human virtual avatar according to claim 1, characterized in that, Based on the multiple images, the depth map, and the virtual human dataset, a trained neural network model is used for 3D reconstruction and rendering to obtain a 3D human model and corresponding rendered images, including: The shape parameters and first pose parameters of the human body are obtained from the multiple images; Based on the shape parameters and the first pose parameters, an initial three-dimensional human body template mesh is generated; Based on the initial 3D human body template mesh, generate initial Gaussian mixture model parameters; Based on the initial Gaussian mixture model parameters, the density field and color field of the three-dimensional scene are generated; Based on the initial Gaussian mixture model parameters, the first pose parameters, and the second pose parameters from the virtual human dataset, deformed Gaussian mixture model parameters are generated. Based on the deformed Gaussian mixture model parameters, and the external and internal parameters of the camera used to obtain the multiple images, a 3D human body model and corresponding rendered images are obtained.

3. The method for reconstructing a human virtual avatar according to claim 2, characterized in that, Based on the initial 3D human body template mesh, the parameters for generating the initial Gaussian mixture model include: The vertices of the initial 3D human body template mesh are used as the point cloud positions of the Gaussian mixture model, and the normal direction of the mesh facets is used as the depth direction to initialize the parameters of the Gaussian mixture model, thus generating the initial Gaussian mixture model parameters.

4. The method for reconstructing a human virtual avatar according to claim 2, characterized in that, Based on the initial Gaussian mixture model parameters, the density field and color field of the 3D scene are generated as follows: For each Gaussian component, determine the Gaussian density; The color of the corresponding Gaussian component is determined based on the Gaussian density; The density field and color field of the 3D scene are determined based on the Gaussian density and color of each Gaussian component.

5. The method for reconstructing a virtual human avatar according to claim 2, characterized in that, Based on the initial Gaussian mixture model parameters, the first pose parameters, and the second pose parameters from the virtual human dataset, the deformed Gaussian mixture model parameters are generated as follows: Learnable skin weights are used to deform each Gaussian component using the first pose parameter and the second pose parameter; The transformation of the Gaussian component is determined, and the deformed mean and rotation are determined to determine the parameters of the deformed Gaussian mixture model.

6. The method for reconstructing a human virtual avatar according to claim 2, characterized in that, Based on the deformed Gaussian mixture model parameters, and the external and internal parameters of the camera used to obtain the multiple images, the resulting 3D human body model and corresponding rendered images are obtained, including: Convert the mean, covariance, color, and density of a 3D Gaussian into a 2D representation.

7. The method for reconstructing a virtual human avatar according to claim 1, characterized in that, During the training of the neural network model, reconstruction loss function, depth consistency loss function, fractional distillation sampling generation loss function, and color gamut consistency loss function are used for model supervision.

8. A device for reconstructing a virtual human avatar, characterized in that, The device for reconstructing the human virtual avatar includes: The image and dataset acquisition module is configured to acquire multiple images of the human body and a virtual human body dataset. A depth map acquisition module is configured to acquire depth maps from the plurality of images; The model building and rendering module is configured to perform 3D reconstruction and rendering using a trained neural network model based on the multiple images, the depth map, and the virtual human dataset, to obtain a 3D human model and the corresponding rendered image.

9. A terminal, comprising: At least one memory and at least one processor; Wherein, the at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the method for reconstructing a human virtual avatar as described in any one of claims 1 to 7.

10. A storage medium for storing program code for executing the method for reconstructing a human virtual avatar according to any one of claims 1 to 7.