Three-dimensional reconstruction method, device, terminal and storage medium

By combining explicit geometric meshes with implicit representations, the problems of insufficient accuracy in clothing geometric reconstruction and insufficient flexibility in texture coupling are solved, achieving high-precision clothing reconstruction and flexible clothing editing, which is suitable for virtual reality and fashion design.

CN122115696APending Publication Date: 2026-05-29BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies have limitations in decoupling the human body from clothing and reconstructing high-quality clothing geometric details. Implicit representation leads to insufficient accuracy in clothing geometry reconstruction, and the coupling of clothing geometry with texture limits the user's flexible processing.

Method used

A method combining explicit geometric meshes and implicit representations is introduced to improve the accuracy of clothing geometry reconstruction through explicit geometric constraints, and to decouple the geometric and texture information of clothing, allowing for independent editing and adjustment.

Benefits of technology

It improves the representation of geometric details in clothing, ensures accurate capture of geometric features under complex shapes, enhances the flexibility of virtual try-on and style editing, and meets the needs of virtual reality and fashion design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115696A_ABST
    Figure CN122115696A_ABST
Patent Text Reader

Abstract

The disclosure provides a three-dimensional reconstruction method and device, a terminal and a storage medium. The three-dimensional reconstruction method comprises: acquiring a video including a human body and clothes worn by the human body; processing the video by using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing the corresponding clothes, wherein processing the video by using the trained neural network model comprises: performing human three-dimensional reconstruction based on the video; performing clothes capturing and rendering based on the video; wherein performing clothes capturing and rendering based on the video comprises: obtaining a three-dimensional position of a clothes standard space point in a video frame of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric grid information from the density field; and generating a clothes model in the three-dimensional reconstruction model based on the density field, the color field and the explicit geometric grid information. The disclosure can effectively improve the detail performance of clothes and ensure that the geometric characteristics of clothes under complex shapes can be accurately captured through explicit geometric constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of information technology, and in particular to three-dimensional reconstruction methods, apparatus, terminals and storage media. Background Technology

[0002] In fields such as virtual reality (VR), augmented reality (AR), fashion design, virtual try-on, and virtual social interaction, automatically reconstructing virtual avatars wearing clothing has become an important research direction. Reconstructing virtual avatars can not only provide users with a more immersive experience but also significantly improve the efficiency of personalized clothing design and virtual try-on. However, despite extensive research in this area, existing solutions still have many limitations, particularly in decoupling the human body from clothing and reconstructing high-quality garment geometric details. Therefore, further improvements in these areas are expected. Summary of the Invention

[0003] To address the existing problems, this disclosure provides a three-dimensional reconstruction method, apparatus, terminal, and storage medium.

[0004] The following technical solution is adopted in this disclosure.

[0005] This disclosure provides a three-dimensional reconstruction method, comprising: acquiring a video including a human body and its clothing; processing the video using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing corresponding clothing, wherein processing the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and capturing and rendering clothing based on the video; wherein capturing and rendering clothing based on the video includes: acquiring the three-dimensional position of standard spatial points of clothing in video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

[0006] Another embodiment of this disclosure provides a three-dimensional reconstruction apparatus. The processing apparatus includes: a video acquisition module configured to acquire a video including a human body and its clothing; and a processing module configured to process the video using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing corresponding clothing. The processing of the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and capturing and rendering clothing based on the video. The capturing and rendering of clothing based on the video includes: acquiring the three-dimensional positions of standard spatial points of clothing in video frames of the video; obtaining a density field and a color field based on the three-dimensional positions; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

[0007] In some embodiments, this disclosure provides a terminal, including: at least one memory and at least one processor; wherein the memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above-described three-dimensional reconstruction method.

[0008] In some embodiments, this disclosure provides a storage medium for storing program code for executing the above-described three-dimensional reconstruction method.

[0009] This disclosure addresses the issue of insufficient accuracy in clothing geometry reconstruction caused by implicit representation by introducing explicit geometric mesh information and combining explicit mesh with implicit representation. Through explicit geometric constraints, it effectively improves the detail representation of clothing, ensuring that the geometric features of clothing in complex shapes can be accurately captured. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a flowchart of a three-dimensional reconstruction method according to an embodiment of the present disclosure.

[0012] Figure 2 A structural diagram of a neural network model according to some embodiments is shown.

[0013] Figure 3 This is a partial module of a three-dimensional reconstruction apparatus according to another embodiment of this disclosure.

[0014] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should be understood that the various steps described in the method embodiments of this disclosure can be performed in sequence and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0017] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0019] It should be noted that the use of the word "a" in this disclosure is illustrative rather than restrictive, and those skilled in the art should understand that it should be understood as "one or more" unless otherwise expressly indicated in the context.

[0020] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0021] Currently, 3D virtual character reconstruction technology is mainly divided into two categories: reconstruction methods based on multi-view cameras and reconstruction methods based on monocular video. Multi-view camera methods rely on multiple cameras simultaneously capturing images of a person from different angles to generate a high-precision 3D model. This type of method can reconstruct very detailed geometry, especially excelling in the reconstruction of complex clothing. However, this method requires a large amount of equipment, is costly, and complex to operate, making it difficult to scale up for large-scale applications.

[0022] In contrast, monocular video-based reconstruction methods have become a research hotspot due to their lower equipment requirements, lower cost, and simpler operation. By recording video of a person using a single camera and combining it with advanced computer vision and machine learning techniques, monocular video reconstruction methods can generate 3D models. In recent years, many algorithms have made significant progress in improving the quality of monocular video reconstruction, especially those relying on human parametric models and implicit representations.

[0023] Parametric modeling methods, such as SMPL and SMPL-X, are among the most widely used 3D human reconstruction techniques. They generate accurate 3D human models by parametrically configuring the shape and pose of the human body, and then fit these parameters to the human pose in a video. However, the topology of these models is fixed, which significantly limits their ability to model complex clothing, making it difficult to capture the shape and details of intricate garments, such as folds or loose sections.

[0024] Implicit representation techniques, such as Neural Radiation Fields (NeRF), have made significant progress in the field of 3D reconstruction in recent years. NeRF uses implicit functions to represent color and density in a scene, enabling the generation of highly realistic 3D renderings, particularly adept at capturing complex shapes and details. Therefore, it has been increasingly applied to the reconstruction of figures and clothing. However, due to the lack of geometric constraints, NeRF often lacks precision in generating geometric structures, especially in areas where the shape and edge details of clothing may be blurry and inaccurate, failing to reconstruct the geometric details of clothing effectively.

[0025] In addition, there are some two-layer reconstruction methods, such as SCARF. However, because the NeRF of this method lacks explicit geometric constraints, the generated clothing geometry details are often not accurate enough, especially in some clothing details, resulting in poor reconstruction quality. Furthermore, this method couples the geometry and texture of clothing together for modeling, which limits the user's independent modification of the appearance of clothing and makes it impossible to flexibly adjust the texture, color, and other features of clothing, resulting in insufficient operational flexibility.

[0026] Implicit representation techniques (such as Neural Radiation Field, NeRF) are flexible in capturing complex shapes, but they often perform poorly in the geometric reconstruction of clothing. Because implicit representations lack explicit geometric constraints, the generated clothing models often suffer from blurry details and edge processing, resulting in insufficient geometric accuracy. This is especially true when dealing with complex clothing, such as skirts or coats with folds; implicit representations may fail to effectively capture the key geometric features of these garments, affecting the final reconstruction results. The geometry of clothing is often influenced by implicit representations, causing the generated models to lack the necessary clarity in edges and details. For example, geometric features in areas such as folds are often ignored or improperly processed, resulting in unsatisfactory clothing reconstruction quality. Therefore, addressing the shortcomings of implicit representations in clothing geometric reconstruction and improving the geometric accuracy of clothing has become an urgent technical problem to be solved.

[0027] Furthermore, existing reconstruction methods typically couple the geometric and textural information of clothing together for modeling, which limits the user's flexibility in manipulating clothing details. In many applications, such as virtual try-on, clothing style editing, and texture conversion, users want to be able to freely adjust other features of clothing while maintaining its geometric shape. However, because existing methods tightly couple geometry and texture, users often need to reconsider the entire clothing geometry when modifying textures, thus reducing the flexibility and efficiency of the operation. Therefore, decoupling clothing geometry and texture information to allow users to edit and adjust independently has become an important issue for improving user experience.

[0028] This disclosure addresses the issue of insufficient accuracy in clothing geometry reconstruction caused by implicit representation by introducing a technique that combines explicit meshes with implicit representation. Through explicit geometric constraints, the detail representation of clothing can be effectively improved, ensuring that the geometric features of clothing in complex shapes can be accurately captured. Furthermore, this disclosure aims to address the lack of flexibility caused by the coupling of geometry and texture in existing technologies by decoupling the geometric and texture information of clothing. By allowing users to independently process the geometric and texture features of clothing, the flexibility of applications such as virtual try-on and style editing can be improved, enabling users to freely adjust the appearance of clothing without affecting its geometric structure. By combining the advantages of explicit meshes and implicit representation, this disclosure improves the representation of clothing geometric details while supporting more flexible texture processing, thereby overcoming the shortcomings of existing technologies and meeting the needs of users in virtual reality, fashion design, and personalized try-on fields.

[0029] Figure 1 A flowchart of a three-dimensional reconstruction method according to embodiments of the present disclosure is provided. Figure 2A structural diagram of a neural network model according to some embodiments is shown. The 3D reconstruction method of this disclosure may include step S101, acquiring a video including a human body and its clothing. In some embodiments, the video is a monocular video. In some embodiments, for example, the video includes content of a human body wearing clothing rotating, thus facilitating the extraction of 3D information of the human body and clothing from the video, thereby facilitating 3D model reconstruction.

[0030] In some embodiments, the method of this disclosure may further include step S102, processing the video using a trained neural network model to obtain a three-dimensional reconstruction model of a human body wearing corresponding clothing. Therefore, the trained neural network model of this disclosure is capable of three-dimensional reconstruction of the human body and clothing based on monocular video.

[0031] In some embodiments, processing the video using a trained neural network model includes: performing 3D reconstruction of the human body based on the video; and capturing and rendering clothing based on the video. In some embodiments, capturing and rendering clothing based on the video includes: acquiring the 3D position of standard spatial points of clothing in video frames; obtaining a density field and a color field based on the 3D position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the 3D reconstruction model based on the density field, color field, and explicit geometric mesh information.

[0032] This disclosure introduces a reconstruction method that combines explicit meshes with implicit representations. The explicit mesh provides clear geometric constraints, while the implicit representation effectively improves the accuracy of clothing geometric reconstruction, capturing the geometric features of complex clothing shapes and edges. In particular, the explicit mesh provides prior constraints on clothing shapes, significantly improving reconstruction results and ensuring clearer and more accurate geometric details of the clothing.

[0033] In some embodiments, human body 3D reconstruction based on video includes: extracting human body shape parameters, posture parameters, and expression parameters from the video; generating a basic human body mesh based on the shape parameters, posture parameters, and expression parameters; determining vertex offsets based on the vertex positions of the basic human body mesh; adding the vertex offsets to the basic human body mesh to generate a human body mesh; and determining vertex color information based on the vertex positions of the basic human body mesh. In some embodiments, commonly used existing human pose estimation algorithms can be used to extract human body shape parameters, posture parameters, and expression parameters from the video. Then, the human body shape parameters, posture parameters, and expression parameters are input into a model such as SMPL-X to generate a basic human body mesh. SMPL-X is a parameterized human body model that provides efficient modeling of human body shape and posture. The model input includes shape parameters, posture parameters, and expression parameters to generate the corresponding 3D mesh. In some embodiments, SMPL-X integrates linear hybrid skinning to calculate the influence of joints on mesh vertices during model generation, and calculates the new position of the vertex through the weight matrix W and joint rotation θ. In some embodiments, the dimension of the shape parameter β is typically (10) or (300), used to describe the individual's body characteristics, and the overall shape of the model can be changed by controlling the body proportions; posture parameters Its dimension K is usually 54, which represents the rotation of each joint, and the overall dimension is, for example, (165). By controlling the rotation of each joint, a specific body posture is formed; the dimension of the expression parameter ψ is usually (10) or (20), which represents the changes in facial expression. Its role is to enable the reconstruction model to express different facial expressions.

[0034] In some embodiments, based on the vertex positions of a base human body mesh, a multilayer perceptron (MLP) is used to predict the offset O of each vertex. The MLP receives the base vertex positions of the mesh and outputs the corresponding offset. These vertex offsets are then added to the base human body mesh to generate the human body mesh. In some embodiments, the vertex offsets... Dimension N is typically 10475, representing the offset of each vertex in 3D space. Its function is to adjust the base mesh, enhancing detail and realism. In some embodiments, vertex color... It provides detailed visual information, making the model more realistic. In some embodiments, a multilayer perceptron is used to predict the color value C_T of each vertex, with the spatial position of the vertex as input and RGB color as output.

[0035] In some embodiments, Figure 2 The human body reconstruction module outputs a human body mesh. The dimension represents the position coordinates of each vertex, indicating the entire 3D human model, i.e., generating a 3D human model that can be rendered. The main goal of the human body reconstruction module is to generate a high-quality 3D human model from the input shape, pose, and expression parameters, accurately capturing the model's pose changes and details. Therefore, in the human body reconstruction module, the inputs are shape parameters, pose parameters, and expression parameters; a basic human mesh M(β,θ,ψ) is generated using the SMPL-X model; the offset O of each vertex is predicted using an MLP network, and by adding the offsets to the basic mesh, a detailed mesh M(β,θ,ψ,O) is generated; the RGB color value C_T of each vertex is predicted using an MLP network; the output returns the final 3D human mesh and color information.

[0036] In some embodiments, the three-dimensional position of the standard spatial point of the garment This represents the position of each vertex in 3D space, used for implicit modeling in the Neural Radiation Field (NeRF) model. In some embodiments, pixel information extracted from video frames (including normal maps and depth maps) is used to supervise clothing geometry and texture. In some embodiments, the density field... A scalar value describing the geometric density at a specific location for generating the geometry of the garment. In some embodiments, the color field... These are RGB color values ​​used to represent the texture information of the clothing. In some embodiments, explicit geometric mesh information is also included. Corresponding to the vertex coordinates of the clothing mesh, this provides more accurate geometric information, facilitating high-precision reconstruction. Therefore, Figure 2 The main function of the clothing capture module is to capture the geometric shape and texture information of clothing from monocular video, generate high-quality clothing models, and decouple the texture and geometry of clothing. In some embodiments, Flexicubes is used to extract explicit geometric mesh information F_mesh from the density field σ of NeRF to generate high-quality clothing geometry.

[0037] In some embodiments, obtaining the density field and color field based on the three-dimensional position includes: obtaining the density field based on the three-dimensional position using a first multilayer perceptron; and obtaining the color field based on the three-dimensional position using a second multilayer perceptron. This decouples the geometry and texture of the clothing, supporting applications such as texture transfer.

[0038] In the clothing capture module, the 3D position of the standard spatial points of the clothing is obtained, and the density field and color field are generated using the NeRF model to represent the clothing in an implicit form. At the same time, the color and density are estimated by two MLPs respectively to achieve decoupling, which facilitates subsequent applications such as texture transfer. The implicit density and color information are combined with the explicit geometric mesh to generate a complete clothing model, that is, outputting complete clothing geometric mesh and texture information.

[0039] In some embodiments, clothing loss functions and body loss functions are used for model supervision during the training of the neural network model. In some embodiments, the clothing loss function includes a normal loss function, a mask loss function, and a reconstruction loss function. In some embodiments, the body loss function includes a color loss function, a skin color loss function, an outer mask loss function, an inner mask loss function, a contour loss function, and a regularization loss function. In some embodiments, in Figure 2 In the model supervision module, the inputs are simulated normal maps, reconstructed images, and masks, and the outputs are optimized geometric and texture information to improve the overall performance of the model and ensure the quality of the reconstruction results. In some embodiments, the pseudo-ground normal map represents the normal information in the video frame and is used to supervise the accuracy of geometric reconstruction; the reconstructed image is an image generated by the model, representing the final reconstruction result; the skin color mask in the mask is used to mark skin color areas, the clothing mask is used to mark clothing areas, and the contour map is used to describe the binary image of the human body contour.

[0040] In some embodiments, the clothing loss function is used to generate high-quality clothing geometry and texture, ensuring consistency with the input data. Specifically, the normal loss function is used to constrain the geometry of the implicit radiation field, ensuring that the difference between the generated normal map and the real normal map is minimized; the mask loss function is used to constrain the reconstruction accuracy of the clothing region, ensuring that the generated mask is consistent with the real mask; and the reconstruction loss function is used to minimize the pixel-level distance between the generated clothing image and the real clothing image to supervise the geometry and texture.

[0041] In some embodiments, the body loss function is used to simultaneously reconstruct the human body and clothing, ensuring the accuracy of their separation and reconstruction. Specifically, the color loss function is used to supervise the difference between the generated body color and the real image, using a mask of visible body parts; the skin color loss function is used to assume that the hand color is visible, ensuring that the color of the occluded body area is consistent with the hand color; the outer mask loss function is used to ensure that the rendered body mask is similar to the real mask; the inner mask loss function is used to ensure that the rendered inner body mask is consistent with the clothing mask; the contour loss function is used to ensure the similarity between the body mask and the contour map; and the regularization loss function is used to constrain the smoothness of the reconstructed mesh surface, ensuring that the vertex offset is minimized to maintain the smoothness of the mesh.

[0042] Therefore, in Figure 2In the model supervision module, normal maps, depth maps, reconstructed images, and ground truth images and masks are acquired. Then, the following supervision is performed: Normal supervision: calculates normal loss to supervise the accuracy of geometric reconstruction; Image reconstruction supervision: calculates reconstruction loss to supervise the quality of image generation; Skin color supervision: calculates skin color loss to ensure the accuracy of skin color regions; Clothing supervision: calculates clothing loss to ensure the accuracy of clothing regions; External mask supervision: calculates external mask loss to ensure the appearance of the body is consistent with the ground truth data; Internal mask supervision: calculates internal mask loss to ensure the internal mask matches the clothing mask; Contour supervision: calculates contour loss to ensure the generated contours are consistent with the ground truth contours; Regularization supervision: calculates regularization loss to control the smoothness of the mesh; Total loss calculation: calculates total loss = clothing loss + body loss, used to optimize the performance of the entire model.

[0043] In some embodiments, for each frame of monocular video, common human pose estimation methods are used to estimate pose and shape parameters. These parameters represent the 3D rotation of human joints and the principal component analysis (PCA) coefficients of the T-pose model shape space, respectively. Additionally, weak perspective camera parameters are estimated for each frame to ensure accurate geometric information in the image. Furthermore, a common robust human parsing method is employed to obtain a segmentation mask for the input image. This mask is used to distinguish different body parts and clothing regions, allowing for better processing of each part in subsequent steps. Furthermore, common methods are used to estimate the normals of the input image, and by multiplying them with the clothing mask, a normal map of the clothing is obtained. This step helps capture the geometric details of the clothing, making subsequent modeling more accurate. All extracted parameters and masks are standardized and formatted for input into the model for training and testing; ensuring the consistency and quality of the input data is crucial for achieving high-precision reconstruction.

[0044] The model training process, including the optimization of model parameters, is implemented using the PyTorch framework. Specifically, the Adam optimizer can be used, as it has good convergence characteristics and is suitable for large-scale deep learning tasks. Furthermore, the optimization process can be divided into stages. In the first stage, the standard NeRF model is used to jointly optimize the entire dressed human model, while simultaneously refining the pose of the SMPL-X model. Training is performed, for example, 100k iterations, with a learning rate set to 5×.

[0045] In the second stage, the human body and clothing are decoupled, and an additional 50k iterations are performed. During this stage, the learning rate for NeRF and pose refinement can be, for example, 10^-4, and the learning rate for the human body color model and offsets can be, for example, 10^-5. In this stage, the model parameters are gradually adjusted to ensure the final reconstruction quality. Afterwards, end-to-end training can be performed using all supervision.

[0046] In some embodiments, the trained neural network model outputs the following information: a 3D human body mesh, a clothing mesh, a normal map, mask information, a density field, and a color field. The 3D human body mesh contains complete human geometry information, enabling rendering and display; the clothing mesh describes the generated clothing geometry, precisely conforming to the input data; the normal map describes the normal information of the clothing and human body surfaces, providing a level of detail; the mask information includes clothing masks and human body masks, facilitating subsequent processing and analysis; the density field output by NeRF describes the geometric density at each spatial point, supporting the generation of fine 3D structures; the color field output by NeRF provides color information at each spatial point for generating the final image.

[0047] After training, algorithm testing can be performed to evaluate the model's performance. Specifically, this can include model validation, performance evaluation, results presentation, and iterative optimization. Model validation uses unseen test data to validate the trained model and evaluate its performance in real-world scenarios. Performance evaluation quantifies the model's effectiveness by calculating different metrics (such as reconstruction accuracy and visual quality), and can use a combination of quantitative evaluation (such as IoU and PSNR) and qualitative evaluation (such as visual effects). Results presentation can show the model's reconstruction performance in different test videos and compare it with real-world scenes to verify the model's effectiveness. Iterative optimization can fine-tune the model based on the test results, improving the network structure or loss function to enhance the model's performance.

[0048] The solution disclosed herein requires no post-processing and can be directly applied to real-world scenarios. This feature significantly improves the user experience, especially in applications requiring real-time feedback (such as virtual try-on and augmented reality), enabling rapid generation of usable results and lowering the technical barrier. Furthermore, by combining human pose estimation and the SMPL-X model, this disclosure quickly and accurately estimates pose and shape parameters, ensuring the model's accuracy and robustness in dynamic scenes. This disclosure also features deep optimization in the loss function design, employing a decoupling method for clothing and body losses to ensure the model simultaneously considers the geometric and textural features of both the human body and clothing, thereby enhancing the realism of the reconstruction.

[0049] Furthermore, this disclosure combines explicit geometric mesh information with implicit representation, resulting in higher accuracy and detail in the reconstruction, solving the problem of capturing complex geometric shapes in traditional methods. Through implicit representation, the model can flexibly handle diverse scenes, thereby improving the overall reconstruction quality. This disclosure effectively decouples the geometric structure and texture information of clothing, allowing for independent optimization of both. This strategy not only enhances modeling flexibility but also makes the generated clothing visually more realistic.

[0050] This disclosure provides explicit geometric constraints through an explicit mesh, which, based on implicit representation, effectively improves the accuracy of clothing geometric reconstruction and captures the geometric features of complex clothing shapes and edges. In particular, the explicit mesh for prior constraints on clothing shapes significantly improves reconstruction results, ensuring clearer and more accurate geometric details. By decoupling the geometric and texture information of clothing during modeling, this disclosure solves the problem of insufficient flexibility caused by the coupling of geometry and texture. This decoupling allows users to freely adjust texture, color, and other features while maintaining the clothing's geometric structure, greatly enhancing the flexibility of virtual try-on and personalized design, and enabling texture transfer between different people.

[0051] Embodiments of this disclosure also provide a three-dimensional reconstruction apparatus 400. Figure 3 A three-dimensional reconstruction apparatus 400 according to some embodiments is shown. The three-dimensional reconstruction apparatus 400 includes a video acquisition module 401 and a processing module 402. In some embodiments, the video acquisition module 401 is configured to acquire a video including a human body and its clothing. In some embodiments, the processing module 402 is configured to process the video using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing corresponding clothing, wherein processing the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and capturing and rendering clothing based on the video; wherein capturing and rendering clothing based on the video includes: acquiring the three-dimensional position of standard spatial points of clothing in video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

[0052] It should be understood that the description of the three-dimensional reconstruction method also applies to the three-dimensional reconstruction device 400 described here, but for simplicity, it will not be described in detail here.

[0053] In some embodiments, obtaining the density field and color field based on the three-dimensional position includes: obtaining the density field based on the three-dimensional position using a first multilayer perceptron; and obtaining the color field based on the three-dimensional position using a second multilayer perceptron. In some embodiments, performing three-dimensional human body reconstruction based on video includes: extracting shape parameters, pose parameters, and expression parameters of the human body from the video; generating a basic human body mesh based on the shape parameters, pose parameters, and expression parameters; determining vertex offsets based on the vertex positions of the basic human body mesh; adding the vertex offsets to the basic human body mesh to generate a human body mesh; and determining vertex color information based on the vertex positions of the basic human body mesh. In some embodiments, during the training of the neural network model, a clothing loss function and a body loss function are used for model supervision. In some embodiments, the clothing loss function includes a normal loss function, a mask loss function, and a reconstruction loss function. In some embodiments, the body loss function includes a color loss function, a skin color loss function, an external mask loss function, an internal mask loss function, a contour loss function, and a regularization loss function. In some embodiments, the trained neural network model outputs the following information: a three-dimensional human body mesh, a clothing mesh, a normal map, mask information, a density field, and a color field.

[0054] In addition, this disclosure also provides a terminal, including: at least one memory and at least one processor; wherein, the memory is used to store program code, and the processor is used to call the program code stored in the memory to execute the above-described three-dimensional reconstruction method.

[0055] In addition, this disclosure also provides a computer storage medium storing program code for executing the above-described three-dimensional reconstruction method.

[0056] The above description, based on embodiments and application examples, illustrates the three-dimensional reconstruction method and apparatus of this disclosure. Furthermore, this disclosure also provides a terminal and a storage medium, which are described below.

[0057] The following is for reference. Figure 4 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 500 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0058] like Figure 4As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0059] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0060] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0061] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0062] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0063] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0064] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods of the present disclosure.

[0065] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0066] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0067] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0068] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0069] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0070] According to one or more embodiments of this disclosure, a three-dimensional reconstruction method is provided, the three-dimensional reconstruction method comprising: acquiring a video including a human body and its clothing; processing the video using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing corresponding clothing, wherein processing the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and capturing and rendering clothing based on the video; wherein capturing and rendering clothing based on the video includes: acquiring the three-dimensional position of standard spatial points of clothing in video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

[0071] According to one or more embodiments of this disclosure, obtaining a density field and a color field based on the three-dimensional position includes: obtaining the density field based on the three-dimensional position using a first multilayer perceptron; and obtaining the color field based on the three-dimensional position using a second multilayer perceptron.

[0072] According to one or more embodiments of this disclosure, performing three-dimensional reconstruction of the human body based on the video includes: extracting shape parameters, posture parameters, and expression parameters of the human body from the video; generating a basic human body mesh based on the shape parameters, posture parameters, and expression parameters; determining vertex offsets based on the vertex positions of the basic human body mesh; adding the vertex offsets to the basic human body mesh to generate a human body mesh; and determining vertex color information based on the vertex positions of the basic human body mesh.

[0073] According to one or more embodiments of this disclosure, clothing loss function and body loss function are used for model supervision during the training process of a neural network model.

[0074] According to one or more embodiments of this disclosure, the clothing loss function includes a normal loss function, a mask loss function, and a reconstruction loss function.

[0075] According to one or more embodiments of this disclosure, the body loss function includes a color loss function, a skin color loss function, an outer mask loss function, an inner mask loss function, a contour loss function, and a regularization loss function.

[0076] According to one or more embodiments of this disclosure, the trained neural network model outputs the following information: a three-dimensional human body mesh, a clothing mesh, a normal map, mask information, a density field, and a color field.

[0077] According to one or more embodiments of this disclosure, a three-dimensional reconstruction apparatus is provided, the apparatus comprising: a video acquisition module configured to acquire a video including a human body and its clothing; and a processing module configured to process the video using a trained neural network model to obtain a three-dimensional reconstruction model of the human body wearing corresponding clothing, wherein processing the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and performing clothing capture and rendering based on the video; wherein performing clothing capture and rendering based on the video includes: acquiring the three-dimensional position of standard spatial points of clothing in video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

[0078] According to one or more embodiments of the present disclosure, a terminal is provided, comprising: at least one memory and at least one processor; wherein the at least one memory is used to store program code, and the at least one processor is used to invoke the program code stored in the at least one memory to execute the method described in any one of the above descriptions.

[0079] According to one or more embodiments of the present disclosure, a storage medium is provided for storing program code for performing the methods described above.

[0080] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0081] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0082] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A three-dimensional reconstruction method, characterized in that, The three-dimensional reconstruction method includes: Acquire videos including the human body and its clothing; The video is processed using a trained neural network model to obtain a 3D reconstructed model of a human body wearing corresponding clothing. The processing of the video using the trained neural network model includes: performing 3D human body reconstruction based on the video; and performing clothing capture and rendering based on the video. The process of capturing and rendering clothing based on the video includes: obtaining the three-dimensional position of standard spatial points of clothing in the video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

2. The three-dimensional reconstruction method according to claim 1, characterized in that, The density field and color field obtained based on the three-dimensional position include: The density field is obtained based on the three-dimensional position using a first multilayer perceptron. The color field is obtained based on the three-dimensional position using a second multilayer perceptron.

3. The three-dimensional reconstruction method according to claim 1, characterized in that, Human body 3D reconstruction based on the video includes: Extract human shape parameters, posture parameters, and facial expression parameters from the video; A basic human body mesh is generated based on the shape parameters, the pose parameters, and the expression parameters. Based on the vertex positions of the basic human body mesh, determine the vertex offset; Add the vertex offsets to the base human body mesh to generate a human body mesh; Vertex color information is determined based on the vertex positions of the basic human body mesh.

4. The three-dimensional reconstruction method according to claim 1, characterized in that, During the training of the neural network model, clothing loss function and body loss function are used for model supervision.

5. The three-dimensional reconstruction method according to claim 4, characterized in that, The clothing loss function includes a normal loss function, a mask loss function, and a reconstruction loss function.

6. The three-dimensional reconstruction method according to claim 4, characterized in that, The body loss function includes a color loss function, a skin color loss function, an outer mask loss function, an inner mask loss function, a contour loss function, and a regularization loss function.

7. The three-dimensional reconstruction method according to claim 1, characterized in that, The trained neural network model outputs the following information: 3D human body mesh, clothing mesh, normal map, mask information, density field, and color field.

8. A three-dimensional reconstruction device, characterized in that, The three-dimensional reconstruction device includes: The video acquisition module is configured to acquire videos including a person's body and their clothing; The processing module is configured to process the video using a trained neural network model to obtain a three-dimensional reconstruction model of a human body wearing corresponding clothing. The processing of the video using the trained neural network model includes: performing three-dimensional reconstruction of the human body based on the video; and performing clothing capture and rendering based on the video. The process of capturing and rendering clothing based on the video includes: obtaining the three-dimensional position of standard spatial points of clothing in the video frames of the video; obtaining a density field and a color field based on the three-dimensional position; extracting explicit geometric mesh information from the density field; and generating a clothing model in the three-dimensional reconstruction model based on the density field, the color field, and the explicit geometric mesh information.

9. A terminal, comprising: At least one memory and at least one processor; The at least one memory is used to store program code, and the at least one processor is used to call the program code stored in the at least one memory to execute the three-dimensional reconstruction method according to any one of claims 1 to 7.

10. A storage medium for storing program code for executing the three-dimensional reconstruction method according to any one of claims 1 to 7.