Dynamic digital human decoupling reconstruction method, electronic equipment and storage medium
Through the dynamic digital human decoupling reconstruction method, the hybrid symbol distance field and deformation field technology is used to realize efficient decoupling modeling of human bodies and clothing in virtual digital human technology and dynamic clothing details capture, solving the problems of low modeling accuracy and efficiency in the existing technology, and achieving efficient and low-cost dynamic human body modeling.
Patent Information
- Application Number
- CN202411965779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The existing virtual digital human technology lacks efficient decoupling modeling methods, and it is difficult to reconstruct the blocking area, and the modeling accuracy and efficiency are low, which cannot meet the needs of dynamic clothing changes and personalized editing.
Through the dynamic digital human decoupling reconstruction method, a hybrid symbol distance field (hmSDF) combined with explicit geometry and implicit geometry is used to realize independent representation of the human body and clothing, and capture the changes in dynamic clothing details through a non-rigid deformation field and a linear mixed skin deformation field.
It significantly improves the geometric accuracy of invisible areas, achieves dynamic clothing effects with high precision and rich details, supports clothing replacement and personalized editing in a variety of application scenarios, and quickly generates dynamic human models under limited hardware conditions.
Smart Images

Figure CN120070680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and graphics, and particularly to a method for decoupled reconstruction of dynamic digital humans, an electronic device, and a storage medium. Background Art
[0002] With the development of computer vision and graphics technologies in recent years, virtual digital humans based on artificial intelligence are in wide demand in applications such as virtual fitting, motion driving, and film and television production. Especially in the recently popular metaverse, similar to humans being the core in the real society, virtual digital human technology is also one of the core technologies in the metaverse. How to efficiently and high-quality represent a three-dimensional human body is a widely concerned issue. Among them, using low-dimensional vector parameterization to represent a three-dimensional human body is one of the core technologies in virtual digital humans.
[0003] However, existing virtual digital human technologies have technical problems such as a lack of an efficient decoupled modeling method, difficulty in reconstructing occluded regions, and low modeling accuracy and efficiency. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method for decoupled reconstruction of dynamic digital humans, an electronic device, and a storage medium that improve the efficiency and accuracy of decoupled reconstruction of virtual digital humans.
[0005] According to a first aspect of the present invention, there is provided a method for decoupled reconstruction of dynamic digital humans, including:
[0006] Performing image segmentation on the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result, and generating a human-clothing hybrid signed distance field by optimizing the initial human body geometry template representing the digital human;
[0007] Using the human-clothing hybrid signed distance field to separate the human body and clothing in the visible region and complete the human body in the invisible region of the initial human body geometry template, to obtain the completed human body geometry template;
[0008] Generating a decoupled human body geometry template using the completed human body geometry template, and deforming the decoupled human body geometry template using a non-rigid deformation field and a linear blend skinning deformation field to obtain the deformed human body geometry template;
[0009] Performing differentiable rendering optimization on the deformed human body geometry template using the current frame image segmentation result to obtain the decoupled reconstruction result of the digital human corresponding to the current frame image, and supervising the optimization process of the initial human body geometry template, the generation process of the decoupled human body geometry template, and the differentiable rendering optimization process using a predefined loss function;
[0010] Optimize the parameters of the non-rigid deformation field, and perform decoupled reconstruction of the digital human on each frame image in the monocular video to obtain the dynamic and continuous decoupled reconstruction result of the target person's digital human.
[0011] According to an embodiment of the present invention, the above-mentioned image segmentation of the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result includes:
[0012] Extract the current frame image from the monocular video of the target person, and use a visual image segmentation algorithm to segment the human body and clothing of the target person in the current frame image to obtain the human body segmentation area and clothing segmentation area of the current frame image.
[0013] According to an embodiment of the present invention, the above-mentioned generation of the human-body clothing hybrid signed distance field by optimizing the initial human body geometry template representing the digital human includes:
[0014] Construct a hybrid 3D tetrahedron based on implicit representation and explicit representation;
[0015] Initialize the hybrid 3D tetrahedron based on implicit representation and explicit representation to obtain the initial human body geometry template representing the digital human;
[0016] Optimize the initial human body geometry template to generate a human-body clothing hybrid signed distance field.
[0017] According to an embodiment of the present invention, the above-mentioned separation of the human body and clothing in the visible area and the complementation of the human body in the invisible area of the initial human body geometry template by using the human-body clothing hybrid signed distance field to obtain the complemented human body geometry template includes:
[0018] Use the human-body clothing hybrid signed distance field to separate the human body and clothing in the visible area of the initial human body geometry template to obtain the separated visible area of the human body;
[0019] Based on the separated visible area of the human body, use a statistical-based 3D human body model to complement the invisible area of the human body in the initial human body geometry to obtain the complemented human body geometry template.
[0020] According to an embodiment of the present invention, the above-mentioned generation of a decoupled human body geometry template by using the complemented human body geometry template and the deformation of the decoupled human body geometry template by using a non-rigid deformation field and a linear blend skinning deformation field to obtain the deformed human body geometry template includes:
[0021] Use the complemented human body geometry template to generate a decoupled human body geometry template, where the decoupled human body geometry template includes a decoupled human body template and a decoupled clothing template;
[0022] The decoupled human template and the decoupled clothing template are non-rigidly deformed using a non-rigid deformation field to obtain a non-rigidly deformed human geometry template;
[0023] The non-rigidly deformed human geometry template is deformed again using a linear blend skinning deformation field to achieve human dynamic pose changes and clothing detail deformations, obtaining a deformed human geometry template.
[0024] According to an embodiment of the present invention, the above-mentioned differentiable rendering optimization of the deformed human geometry template using the current frame image segmentation result to obtain the digital human decoupled reconstruction result corresponding to the current frame image includes:
[0025] By parsing the current frame image segmentation result, the RGB, normal map, 2D human mask, and 2D clothing mask of the current frame image are obtained;
[0026] The RGB differentiable rendering, normal map rendering, 2D human mask differentiable rendering, and 2D clothing differentiable rendering are performed on the deformed human geometry template using the RGB, normal map, 2D human mask, and 2D clothing mask of the current frame image to obtain the digital human decoupled reconstruction result of the target person in the current frame image.
[0027] According to an embodiment of the present invention, the above-mentioned predefined loss function includes a color error loss function, a mask matching error loss function, a normal perception matching error loss function, and a regularization loss function.
[0028] According to an embodiment of the present invention, the above-mentioned regularization loss function includes an eikonal regularization term, a regularization term for encouraging hole opening, a hole regularization term, a collision penalty regularization term, and a geometric regularization term.
[0029] A second aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above-mentioned one or more processors execute the above-mentioned one or more computer programs to implement the steps of the above-mentioned method.
[0030] A third aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above-mentioned computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.
[0031] The dynamic digital human decoupled reconstruction method provided by the present invention can separately perform decoupled modeling on the geometry and clothing of a three-dimensional dynamic human body, rather than simply generating an indivisible overall geometry model. By combining the advantages of explicit geometry and implicit geometry through a hybrid signed distance field (hmSDF), independent representations of clothing and the human body are achieved, thus supporting clothing replacement and personalized editing in various application scenarios. At the same time, the above method provided by the present invention can effectively complete the missing areas of human body geometry caused by clothing occlusion, ensuring the coherence and authenticity of the shape. Compared with the rough effect of traditional methods that only perform completion by simple speculation, the present invention significantly improves the geometric accuracy of invisible areas. Moreover, by combining linear blend skinning (LBS) with a non-rigid deformation field, the present invention can accurately capture the detailed changes of dynamic clothing under complex human postures, including wrinkling and stretching effects. Compared with the defect of existing methods that can only capture rough dynamic clothing changes, the present invention can generate a more realistic dynamic clothing effect. In addition, through differentiable rendering optimization combined with the supervision of monocular video frames, the present invention can quickly generate a dynamic human body model under limited hardware conditions. Compared with the low efficiency of traditional methods that rely on multiple cameras and complex template modeling, the present invention only requires a monocular video to complete the dynamic modeling of the entire video within a few hours, featuring high efficiency and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0033] Figure 1 is an application scenario diagram of the dynamic digital human decoupled reconstruction method according to an embodiment of the present invention;
[0034] Figure 2 is a flowchart of the dynamic digital human decoupled reconstruction method according to an embodiment of the present invention;
[0035] Figure 3 is a schematic diagram of the process of the dynamic digital human decoupled reconstruction method based on a monocular video according to an embodiment of the present invention;
[0036] Figure 4 is a schematic diagram of clothing mask extraction according to an embodiment of the present invention;
[0037] Figure 5 is a schematic diagram of clothing editing according to an embodiment of the present invention;
[0038] Figure 6 is a schematic diagram of human pose editing according to an embodiment of the present invention;
[0039] Figure 7 is a schematic diagram of the structure of the dynamic digital human decoupled reconstruction device according to an embodiment of the present invention;
[0040] Figure 8 It is a block diagram of an electronic device suitable for implementing the dynamic digital human decoupling reconstruction method according to an embodiment of the present invention. Detailed implementation manners
[0041] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0042] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0043] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0044] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0045] Virtual digital humans are widely used in fields such as virtual reality, film and television production, animation generation, virtual fitting, and the metaverse. With the explosive development of the metaverse, the research on virtual digital humans has also become a research hotspot in related technical fields.
[0046] The existing virtual digital human technology still has the following technical problems: (1) Lack of an efficient decoupled modeling method: Most of the existing monocular video human body reconstruction methods can only generate non-decoupled overall geometries, usually modeling the human body and clothing as a whole. This method cannot achieve independent representation of the human body and clothing, restricting the flexibility of the model in multiple application scenarios. Especially in scenarios such as clothing replacement, virtual fitting, and animation production that require high-precision and editable human body models, the existing methods cannot handle the dynamic changes of different clothing types, resulting in the inability to meet customization requirements or presenting unrealistic performances. In addition, the inability to decouple the representation of clothing and the human body makes it necessary to re-perform overall modeling every time the clothing or human body posture is modified, increasing the development cost and computing time. (2) Difficulty in reconstructing occluded areas: In monocular videos, the parts covered by clothing will cause local occlusion of the human body, making the geometric information of these areas invisible. Existing methods often rely on speculation or simplified model completion when dealing with occluded areas, resulting in the reconstructed shapes of these occluded areas lacking accuracy and coherence. Especially for clothing in dynamic videos, as the human body posture changes, the shape of the occluded area also changes, and it is difficult for existing methods to ensure the detail accuracy and geometric coherence of the model in this case, thus affecting the realism of the entire reconstruction effect. (3) Limited modeling accuracy and efficiency: Traditional human body reconstruction methods usually rely on multiple cameras and complex scanning templates to obtain higher modeling accuracy through multi-view collaboration. However, this method requires a large amount of equipment support and computing resources, and has low efficiency, unable to meet the real-time requirements in practical applications. On the other hand, although some simplified methods based on monocular videos have an advantage in modeling speed, they often ignore the capture of details, especially in the performance of clothing wrinkles, texture details, and human body surface morphology, often having large errors, resulting in overly rough or inaccurate reconstruction results and being unable to be used in high-quality animation production or virtual reality applications.
[0047] Embodiments of the present invention provide a dynamic digital human decoupled reconstruction method for solving at least one of the existing technical problems.
[0048] Figure 1 It is an application scenario diagram of the dynamic digital human decoupled reconstruction method according to an embodiment of the present invention.
[0049] As Figure 1 shown, the application scenario 100 according to this embodiment may include virtual reality, film and television production, animation generation, virtual fitting, and the metaverse, etc. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0050] Users can interact with the server 105 through the network 104 using the first terminal device 101, the second terminal device 102, and the third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0051] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and so on.
[0052] The server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0053] It should be noted that the dynamic digital human decoupling and reconstruction method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the dynamic digital human decoupling and reconstruction device provided by the embodiments of the present invention can generally be set in the server 105. The dynamic digital human decoupling and reconstruction method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the dynamic digital human decoupling and reconstruction device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0054] It should be understood, Figure 1 the numbers of terminal devices, networks, and servers in
[0055] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 6 scenario described below, the dynamic digital human decoupling and reconstruction method of the disclosed embodiments will be described in detail through
[0056] In the field of dynamic human body and clothing modeling, traditional methods often struggle to accurately represent both the human body and complex clothing dynamics simultaneously. Existing methods rely on pre-constructed templates, such as human body templates or templates for fixed clothing types. This approach has limited expressiveness for dynamic clothing, often being restricted by the resolution and topology of the templates, and it is difficult to capture richly detailed clothing folds and their dynamic changes. Therefore, the present invention proposes a dynamic human body modeling method based on the combination of explicit and implicit geometry. By combining the clear boundaries of explicit geometry with the high expressive power of implicit functions, it learns the dynamic prior knowledge of the human body and clothing from monocular videos and decouples the human body and clothing representations using the hybrid signed distance field (hmSDF). This method does not require predefined clothing templates and can generate high-precision, richly detailed dynamic clothing shapes with only a small number of network parameters. At the same time, it enables independent editing and optimization of the human body and clothing, significantly enhancing the modeling flexibility and realism.
[0057] It should be specifically noted that the information of the target person involved in the present invention (including but not limited to the image information, pose information, action information, etc. of the target person) and data (including but not limited to the data for analysis, stored data, displayed data, etc.) are all information and data authorized by the target person or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for the target person to choose to authorize or reject.
[0058] In the scenario of making automated decisions using the information of the target person, the methods, devices, and systems provided in the embodiments of the present invention all provide corresponding operation entrances for the target person to choose to agree or reject the results of the automated decisions; if the user chooses to reject, it enters the expert decision-making process. Here, the expression "automated decision" refers to the activity of automatically analyzing and evaluating an individual's behavior habits, interests, hobbies, or economic, health, credit status, etc. through a computer program and making decisions. Here, the expression "expert decision" refers to the activity of making decisions by personnel who are engaged in work in a specific field, have specialized experience, knowledge, and skills, and have reached a certain professional level.
[0059] Figure 2 It is a flowchart of the dynamic digital human decoupling and reconstruction method according to the embodiments of the present invention.
[0060] As Figure 2 shown, the above-mentioned dynamic digital human decoupling and reconstruction method includes operations S210 to S250.
[0061] In operation S210, image segmentation is performed on the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result, and a human-clothing hybrid signed distance field is generated by optimizing the initial human geometry template representing the digital human.
[0062] Frame images are extracted from the monocular video of the target person to obtain a set of frame images, i.e., a frame sequence. Image segmentation is performed on each frame image to obtain the human body region and clothing region of the target person. Note that at this time, the image segmentation result is used as the "ground truth label" for the rendering and optimization of the digital human geometry template in subsequent operations. The image segmentation result includes the RGB information, normal map information, 2D human mask, and 2D clothing mask of the image itself.
[0063] The above initial human geometry template is constructed based on DMTet, combining implicit representation and explicit representation.
[0064] In operation S220, the human-clothing hybrid signed distance field is used to separate the human body and clothing in the visible area and complete the human body in the invisible area of the initial human geometry template, obtaining the completed human geometry template.
[0065] In the above operation S220, the human body and clothing are first separated. Due to the presence of clothing, only part of the human body will exist in the visible area after separation. Therefore, it is necessary to complete the human body in the invisible area.
[0066] In operation S230, a decoupled human geometry template is generated using the completed human geometry template, and the decoupled human geometry template is deformed using the non-rigid deformation field and linear blend skinning deformation field to obtain the deformed human geometry template.
[0067] Through multiple deformations of the non-rigid deformation field and linear blend skinning deformation field, the details of the human body and clothing in the human geometry template are enriched.
[0068] Optionally, the above non-rigid deformation field is a multi-layer perceptron (MLP), which adopts a fully connected neural network, including an input layer, a fully connected layer, an activation layer, and skip connections.
[0069] In operation S240, the deformed human geometry template is optimized by differentiable rendering using the current frame image segmentation result to obtain the decoupled reconstruction result of the digital human corresponding to the current frame image, and the predefined loss function is used to supervise the optimization process of the initial human geometry template, the generation process of the decoupled human geometry template, and the differentiable rendering optimization process.
[0070] Throughout the process, the predefined loss function is used to supervise the decoupled reconstruction process, thereby optimizing the obtained decoupled reconstruction result of the digital human.
[0071] In operation S250, the parameters of the non-rigid deformation field are optimized, and digital human decoupled reconstruction is performed on each frame image in the monocular video to obtain the dynamic and continuous digital human decoupled reconstruction result of the target person.
[0072] For each frame image in the monocular video of the target person, operations S210 to S240 are performed, so as to obtain the digital human decoupled reconstruction result of each frame image. The digital human decoupled reconstruction results of the above-mentioned each frame image are output according to the frame image sequence, and the dynamic and continuous digital human decoupled reconstruction result of the target person, that is, the dynamic and continuous digital human model of the target person, is obtained.
[0073] The dynamic digital human decoupled reconstruction method provided by the present invention can separately perform decoupled modeling on the geometry and clothing of a three-dimensional dynamic human body, rather than simply generating an indivisible overall geometry model. By combining the advantages of explicit geometry and implicit geometry through the hybrid signed distance field (hmSDF), independent representations of clothing and the human body are achieved, thus supporting clothing replacement and personalized editing in a variety of application scenarios. At the same time, the above method provided by the present invention can effectively complement the missing areas of human body geometry caused by clothing occlusion, ensuring the coherence and authenticity of the shape. Compared with the rough effect of simply speculating for complementation in traditional methods, the present invention significantly improves the geometric accuracy of invisible areas. Moreover, by combining linear blend skinning (LBS) with the non-rigid deformation field, the present invention can accurately capture the detailed changes of dynamic clothing under complex human postures, including wrinkle and stretching effects. Compared with the defect that existing methods can only capture rough dynamic clothing changes, the present invention can generate a more realistic dynamic clothing effect. In addition, through the combination of differentiable rendering optimization and the supervision of monocular video frames, the present invention can quickly generate a dynamic human body model under limited hardware conditions. Compared with the low efficiency of traditional methods relying on multi-cameras and complex template modeling, the present invention only requires a monocular video to complete the dynamic modeling of the entire video within a few hours, featuring high efficiency and low cost.
[0074] The following further elaborates on the above-mentioned dynamic digital human decoupled reconstruction method provided by the present invention through specific embodiments in combination with the Figure 3 accompanying drawings.
[0075] Figure 3 is a schematic process diagram of a dynamic digital human decoupled reconstruction method based on a monocular video according to an embodiment of the present invention.
[0076] Extract a frame sequence from the monocular video, and obtain 2D masks of clothing and the human body through image segmentation. As Figure 3As shown, a human - clothing mixed signed distance field (hmSDF) is constructed to separate the visible regions of the human body and clothing and complete the invisible regions; a static template is generated based on the hmSDF, including a decoupled clothing template and a human body template; a linear blend skinning (LBS) deformation field is constructed, combined with a non - rigid deformation field, to achieve human body dynamic pose changes and clothing detail deformations; through differentiable rendering optimization, based on the supervision of RGB, normal maps, and 2D masks, a high - fidelity decoupled geometric model is generated; the non - rigid deformation field parameters are optimized to output a dynamic and continuous decoupled reconstruction result of the human body and clothing.
[0077] According to an embodiment of the present invention, the above - mentioned image segmentation of the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result includes: extracting the current frame image from the monocular video of the target person, and using a visual image segmentation algorithm to segment the human body and clothing in the current frame image to obtain the human body segmentation region and clothing segmentation region of the current frame image.
[0078] The following is a more detailed description of the input and segmentation of the above - mentioned monocular video provided by the present invention through specific embodiments and in combination with Figure 4 the above - mentioned input and segmentation of the monocular video provided by the present invention.
[0079] Figure 4 is a schematic diagram of clothing mask extraction according to an embodiment of the present invention.
[0080] Extract a frame sequence from the input monocular video and generate segmentation masks for clothing and the human body through existing 2D image segmentation methods.
[0081] Specifically, use an open - source image segmentation tool, such as SAM2 (Segment Anything 2), to parse each frame image into independent regions of the human body and clothing. The segmentation results are as Figure 4 shown, from left to right: the captured color image of the person in clothes, the complete clothing mask obtained from SAM2, the clothing mask obtained from SAM2, the mask obtained by presenting only the clothing geometry, and the mask of the effective clothing region after presenting the complete clothing geometry.
[0082] According to an embodiment of the present invention, the above - mentioned generation of the human - clothing mixed signed distance field by optimizing the initial human body geometry template representing the digital human includes: constructing a hybrid 3D tetrahedron based on implicit and explicit representations; initializing the hybrid 3D tetrahedron based on implicit and explicit representations to obtain the initial human body geometry template representing the digital human; and optimizing the initial human body geometry template to generate the human - clothing mixed signed distance field.
[0083] The above initial human body geometry template is generated based on DMTet, which is a deep learning version of the Marching Tetrahedra algorithm for high-resolution 3D shape synthesis. It combines implicit and explicit 3D representations and optimizes the reconstructed surface to produce fine geometric details. Through an end-to-end differentiable process, DMTet can generate 3D models from point clouds or voxel inputs.
[0084] According to an embodiment of the present invention, the above-mentioned separation of the human body and clothing in the visible area and the completion of the human body in the invisible area using the human-clothing hybrid signed distance field to obtain the completed human body geometry template includes: separating the human body and clothing in the visible area of the initial human body geometry template using the human-clothing hybrid signed distance field to obtain the separated visible area of the human body; based on the separated visible area of the human body, using a statistical-based 3D human body model to complete the invisible area of the human body in the initial human body geometry to obtain the completed human body geometry template.
[0085] Human body segmentation is achieved through hmSDF and clothing in the visible area, as defined in Equation (1):
[0086] (1),
[0087] where is the segmentation boundary. For occluded areas , the geometric parameters of the SMPL model are used to complete the filling.
[0088] According to an embodiment of the present invention, the above-mentioned generation of a decoupled human body geometry template using the completed human body geometry template and the deformation of the decoupled human body geometry template using a non-rigid deformation field and a linear blend skinning deformation field to obtain the deformed human body geometry template includes: generating a decoupled human body geometry template using the completed human body geometry template, where the decoupled human body geometry template includes a decoupled human body template and a decoupled clothing template; using the non-rigid deformation field to perform non-rigid deformation on the decoupled human body template and the decoupled clothing template respectively to obtain the non-rigid deformed human body geometry template; using the linear blend skinning deformation field to deform the non-rigid deformed human body geometry template again to achieve human body dynamic pose changes and clothing detail deformations to obtain the deformed human body geometry template.
[0089] The LBS deformation of the human body and the non-rigid deformation field of the clothing are independently modeled, and the dynamic consistency between the two is ensured during the fusion process. The human body deformation is described by linear blend skinning (LBS), as shown in Equation (2):
[0090] (2),
[0091] where is the skeleton weight, is the rigid transformation matrix of the skeleton.
[0092] The non-rigid deformation field of the garment is modeled by an MLP, as shown in Equation (3):
[0093] (3)
[0094] where is a point in the standard pose space, is the latent variable of the garment, are the network parameters.
[0095] According to an embodiment of the present invention, the above-mentioned differentiable rendering optimization of the deformed human body geometry template using the current frame image segmentation result to obtain the decoupled reconstruction result of the digital human corresponding to the current frame image includes: parsing the current frame image segmentation result to obtain the RGB, normal map, 2D human mask, and 2D garment mask of the current frame image; using the RGB, normal map, 2D human mask, and 2D garment mask of the current frame image to perform RGB differentiable rendering, normal map rendering, 2D human mask differentiable rendering, and 2D garment differentiable rendering on the deformed human body geometry template to obtain the decoupled reconstruction result of the digital human of the target person in the current frame image.
[0096] The purpose of differentiable rendering optimization is to combine the human body geometry template with the target person to generate a digital human model similar to the target person.
[0097] According to an embodiment of the present invention, the above-mentioned predefined loss function includes a color error loss function, a mask matching error loss function, a normal perception matching error loss function, and a regularization loss function.
[0098] According to an embodiment of the present invention, the above-mentioned regularization loss function includes an eikonal regularization term, a regularization term for encouraging cavity opening, a cavity regularization term, a collision penalty regularization term, and a geometric regularization term.
[0099] The loss function in the optimization process includes color error, mask matching error, normal perception matching error, and geometric regularization. As shown in Equation (4):
[0100] (4).
[0101] During the training process, the model is optimized by minimizing the difference between the rendering result and the input image. Specifically, it includes the following parts of losses:
[0102] (1) Color Loss
[0103] This loss term measures the reconstruction accuracy by calculating the L1 error between the rendered RGB image and the supervised image (ground truth image), as shown in Equation (5):
[0104] (5),
[0105] where represents the category to which the current pixel belongs. When is true, the pixel belongs to the human body; when is true, the pixel belongs to clothing. and can be true simultaneously, or only one of them can be true, depending on the mask used for supervision. is the rendered RGB image of the human body, is the rendered RGB image of clothing, is the ground truth RGB image of the human body, is the ground truth RGB image of clothing.
[0106] This loss term calculates the pixel-level error by comparing the rendered images of the human body and clothing with the ground truth images, ensuring that the color of the reconstructed image is as similar as possible to the ground truth image.
[0107] (2) Mask Loss
[0108] Although the irrelevant background has been removed in the RGB image, the Mask Loss can further constrain the precision of the boundaries, as shown in Equation (6):
[0109] (6),
[0110] where is the ground truth mask of the human body, is the ground truth mask of clothing. The Mask Loss further helps to optimize the precision of the boundary region, ensuring the faithful reproduction of the boundaries between clothing and the human body.
[0111] (3) Perceptual Normal Loss
[0112] By using the Sapiens dataset to obtain the normal map of the image as supervision, the effect of normal reconstruction is further enhanced, as shown in Equation (7):
[0113] (7),
[0114] where is the rendered normal, is the ground truth normal, Represents the first The activation function of the layer. This loss term aims to improve the reconstruction quality of the normal map by using perceptual loss to strengthen the match between the rendered normal and the real normal, thereby improving the perceptual quality of the normal.
[0115] (4) Regularization Loss
[0116] Regularization terms help control the behavior of the model during the optimization process to avoid overfitting or unreasonable results. This section contains the following regularization losses:
[0117] (4.1) Eikonal Loss
[0118] In order to ensure the rationality of the Signed Distance Field (SDF), when optimizing the SDF field, the gradient of the SDF value toward each tetrahedron vertex Add the Eikonal term, as shown in formula (8):
[0119] (8).
[0120] The Eikonal loss term ensures that the gradient of the signed distance field meets expectations during the optimization process, thus ensuring the correctness of the SDF value.
[0121] (4.2)Encourage Hole Opening
[0122] In order to identify the opening position by using only image information (especially when the viewing angle is limited), a regularization term is introduced to encourage the hole to open. As shown in formula (9):
[0123] (9),
[0124] in, Represents a vertex The signed distance of is the Huber loss, which is used to impose a smooth constraint on the opening.
[0125] This item helps the model identify reasonable opening positions through image information and prevents incorrect opening identification due to viewing angle limitations.
[0126] (4.3) Regularize Holes
[0127] In order to avoid the opening being too large, constraints are imposed on all points visible from the current view. As shown in formula (10):
[0128] (10),
[0129] wherein, is a positive number used to control the maximum range of the opening.
[0130] This regularization term ensures that during the optimization process, the opening of the signed distance field does not become too large, thus preventing unreasonable geometric deformations.
[0131] (4.4)Collision Penalty
[0132] To ensure that the clothing does not penetrate the human body surface, a collision penalty term is introduced. As shown in Equation (11):
[0133] (11),
[0134] wherein, represents the minimum distance from the clothing surface point to the human body surface; is the distance threshold, which is usually set to 0.005 to prevent calculation errors; is the weight of the collision penalty.
[0135] (4.5)Geometry Regularization
[0136] To ensure that the optimization process is constrained and generates smooth deformation results, the present invention introduces a geometry regularization term. The present invention adds the following two regularization terms: Normal Consistency Term: denoted as for maintaining the consistency of surface normals; Laplacian Term: denoted as for smoothing geometric deformations.
[0137] Figure 5 is the schematic diagram of clothing editing according to an embodiment of the present invention.
[0138] Figure 6 is the schematic diagram of human pose editing according to an embodiment of the present invention.
[0139] After obtaining the digital human decoupled reconstruction model of the target person by using the dynamic digital human decoupled reconstruction method provided by the present invention, the digital human decoupled reconstruction model can be edited, including clothing editing and human pose editing.
[0140] As Figure 5 shown, Figure 5 (a)shows the person about to have clothing editing, Figure 5 in (b), the upper three pictures show the clothing display for changing clothes, Figure 5 in (b), the lower three pictures show for Figure 5The result display after the human body changes clothes in (a).
[0141] Figure 6 It is a schematic diagram of human body pose editing. Figure 6 In (a), the dressed human body is in the standard pose. Figure 6 The three figures in (b) are the editing results in three poses respectively.
[0142] Through Figure 5 and 6 The clothing editing and human body pose editing shown, it can be verified that the dynamic digital human decoupling and reconstruction method provided by the present invention can efficiently edit the digital human model of the target person.
[0143] The above scheme of the embodiment of the present invention has the following advantages compared with the traditional three-dimensional human body parameterization representation method:
[0144] (1) Decoupled reconstruction of dynamic human body and clothing: The present invention can separately decouple and model the geometry and clothing of the three-dimensional dynamic human body, rather than simply generating an indivisible overall geometry model. By combining the advantages of explicit geometry and implicit geometry through the hybrid signed distance field (hmSDF), independent representations of clothing and human body are achieved, thus supporting clothing replacement and personalized editing in various application scenarios.
[0145] (2) Precise occlusion area completion: With the help of the SMPL model and geometric optimization technology, the present invention can effectively complete the missing human body geometry areas caused by clothing occlusion, ensuring the coherence and authenticity of the shape. Compared with the rough effect of simply speculating for completion by traditional methods, the present invention significantly improves the geometric accuracy of the invisible areas.
[0146] (3) Efficient capture of dynamic clothing details: By combining linear blend skinning (LBS) and non-rigid deformation fields, the present invention can accurately capture the detail changes of dynamic clothing under complex human body postures, including wrinkle and stretching effects. Compared with the defect that existing methods can only capture rough dynamic clothing changes, the present invention can generate more realistic dynamic clothing effects.
[0147] (4) Coexistence of high efficiency and high quality: The present invention can quickly generate a dynamic human body model under limited hardware conditions through differentiable rendering optimization combined with the supervision of monocular video frames. Compared with the low efficiency of traditional methods relying on multiple cameras and complex template modeling, the present invention only needs a monocular video to complete the dynamic modeling of the entire video within a few hours, and has the characteristics of high efficiency and low cost.
[0148] Corresponding to the foregoing embodiment of the three-dimensional dressed human body parameterization representation method based on the combination of explicit and implicit, the present invention also provides an embodiment of a three-dimensional human body parameterization representation device based on the three-dimensional dressed human body parameterization of the combination of explicit and implicit.
[0149] Figure 7 It is a schematic structural diagram of a dynamic digital human decoupled reconstruction device according to an embodiment of the present invention.
[0150] As Figure 7 shown, the above-mentioned dynamic digital human decoupled reconstruction device 700 includes an image segmentation and template initial optimization module 710, a separation and completion module 720, a multiple deformation module 730, a differentiable rendering and supervised training module 740, and a dynamic continuous decoupled reconstruction module 750.
[0151] The image segmentation and template initial optimization module 710 is used to perform image segmentation on the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result, and generate a human-clothing hybrid signed distance field by optimizing the initial human geometry template representing the digital human. In one embodiment, the image segmentation and template initial optimization module 710 can be used to perform the operation S210 described above, which will not be elaborated here.
[0152] The separation and completion module 720 is used to separate the human body and clothing in the visible area and complete the human body in the invisible area of the initial human geometry template by using the human-clothing hybrid signed distance field, so as to obtain the completed human geometry template. In one embodiment, the separation and completion module 720 can be used to perform the operation S220 described above, which will not be elaborated here.
[0153] The multiple deformation module 730 is used to generate a decoupled human geometry template by using the completed human geometry template, and deform the decoupled human geometry template by using a non-rigid deformation field and a linear blend skinning deformation field to obtain a deformed human geometry template. In one embodiment, the multiple deformation module 730 can be used to perform the operation S230 described above, which will not be elaborated here.
[0154] The differentiable rendering and supervised training module 740 is used to perform differentiable rendering optimization on the deformed human geometry template by using the current frame image segmentation result to obtain a digital human decoupled reconstruction result corresponding to the current frame image, and supervise the optimization process of the initial human geometry template, the generation process of the decoupled human geometry template, and the differentiable rendering optimization process by using a predefined loss function. In one embodiment, the differentiable rendering and supervised training module 740 can be used to perform the operation S240 described above, which will not be elaborated here.
[0155] The dynamic continuous decoupled reconstruction module 750 is used to optimize the parameters of the non-rigid deformation field and perform digital human decoupled reconstruction on each frame image in the monocular video to obtain a dynamic continuous digital human decoupled reconstruction result of the target person. In one embodiment, the dynamic continuous decoupled reconstruction module 750 can be used to perform the operation S250 described above, which will not be elaborated here.
[0156] According to an embodiment of the present disclosure, any plurality of modules among the image segmentation and template initial optimization module 710, the separation and completion module 720, the multiple deformation module 730, the differentiable rendering and supervised training module 740, and the dynamic continuous decoupled reconstruction module 750 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the image segmentation and template initial optimization module 710, the separation and completion module 720, the multiple deformation module 730, the differentiable rendering and supervised training module 740, and the dynamic continuous decoupled reconstruction module 750 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the image segmentation and template initial optimization module 710, the separation and completion module 720, the multiple deformation module 730, the differentiable rendering and supervised training module 740, and the dynamic continuous decoupled reconstruction module 750 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0157] Figure 8 is a block diagram of an electronic device suitable for implementing the dynamic digital human decoupled reconstruction method according to an embodiment of the present invention.
[0158] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which can perform various appropriate actions and processes according to the program stored in the read only memory (ROM) 802 or the program loaded from the storage section 808 into the random access memory (RAM) 803. The processor 801 can include, for example, a general microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 can also include on-board memory for caching purposes. The processor 801 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0159] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0160] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.
[0161] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0162] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803 described above.
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0164] Those skilled in the art can understand that the features described in various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0165] The embodiments of the present invention have been described above. However, these embodiments are only for illustrative purposes and not for limiting the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all these substitutions and modifications should fall within the scope of the present invention.
Claims
1. A dynamic digital human decoupling and reconstruction method, characterized in that: The method comprises: Perform image segmentation on the current frame image extracted from the monocular video of the target person to obtain the current frame image segmentation result, and generate a human-clothing hybrid signed distance field by optimizing the initial human body geometry template representing the digital human; Using the human body-clothing hybrid signed distance field, the initial human body geometry template is used to separate the human body and clothing in the visible area and to complete the human body in the invisible area, so as to obtain a completed human body geometry template; Generate a decoupled human body geometry template using the completed human body geometry template, and deform the decoupled human body geometry template using a non-rigid deformation field and a linear mixed skin deformation field to obtain a deformed human body geometry template; Using the segmentation result of the current frame image, the deformed human body geometry template is optimized by differentiable rendering to obtain a digital human decoupled reconstruction result corresponding to the current frame image, and using a predefined loss function to supervise the optimization process of the initial human body geometry template, the generation process of the decoupled human body geometry template, and the differentiable rendering optimization process; The parameters of the non-rigid deformation field are optimized, and digital human decoupling reconstruction is performed on each frame of the monocular video to obtain a dynamic and continuous digital human decoupling reconstruction result of the target person.
2. The method according to claim 1, characterized in that Perform image segmentation on the current frame image extracted from the monocular video of the target person, and obtain the current frame image segmentation results including: The current frame image is extracted from the monocular video of the target person, and the body and clothing of the target person in the current frame image are segmented using a visual image segmentation algorithm to obtain a body segmentation area and a clothing segmentation area of the current frame image.
3. The method according to claim 1, characterized in that The human-clothing hybrid signed distance field is generated by optimizing the initial human geometry template that represents the digital human, including: Construct a hybrid 3D tetrahedron based on implicit and explicit representations; Initializing the hybrid 3D tetrahedron based on implicit representation and explicit representation to obtain the initial human body geometry template representing the digital human; The initial human body geometry template is optimized to generate the human body-clothing hybrid signed distance field.
4. The method according to claim 1, characterized in that: The human body and clothing hybrid signed distance field is used to separate the human body and clothing in the visible area of the initial human body geometry template and to complete the human body in the invisible area, and the completed human body geometry template is obtained, including: Separating the human body and clothing in the visible area of the initial human body geometric template by using the human body-clothing mixed signed distance field to obtain a separated human body visible area; Based on the separated visible human body area, the invisible human body area in the initial human body geometry is completed using a statistically based three-dimensional human body model to obtain the completed human body geometry template.
5. The method according to claim 1, characterized in that The decoupled human body geometry template is generated by using the completed human body geometry template, and the decoupled human body geometry template is deformed by using a non-rigid deformation field and a linear mixed skin deformation field to obtain the deformed human body geometry template, including: Generate a decoupled human body geometry template using the completed human body geometry template, wherein the decoupled human body geometry template includes a decoupled human body template and a decoupled clothing template; Using the non-rigid deformation field, the decoupled human body template and the decoupled clothing template are respectively subjected to non-rigid deformation to obtain a human body geometric template after non-rigid deformation; The non-rigidly deformed human body geometry template is deformed again using the linear mixed skin deformation field to achieve human body dynamic posture change and clothing detail deformation, thereby obtaining the deformed human body geometry template.
6. The method according to claim 1, characterized in that Using the current frame image segmentation result to perform differentiable rendering optimization on the deformed human body geometric template to obtain a digital human decoupling reconstruction result corresponding to the current frame image includes: By parsing the segmentation result of the current frame image, RGB, normal map, 2D human body mask and 2D clothing mask of the current frame image are obtained; The deformed human body geometric template is subjected to RGB differentiable rendering, normal map rendering, 2D human body mask differentiable rendering and 2D clothing differentiable rendering by using the RGB, normal map, 2D human body mask and 2D clothing mask of the current frame image to obtain a digital human decoupling reconstruction result of the target person in the current frame image.
7. The method according to claim 1, characterized in that The predefined loss functions include a color error loss function, a mask matching error loss function, a normal-aware matching error loss function, and a regularization loss function.
8. The method according to claim 7, characterized in that The regularized loss function includes an einofunctional regularization term, a regularization term that encourages hole opening, a hole regularization term, a collision penalty regularization term, and a geometric regularization term.
9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.