A human body model relighting method for generative digital humans

By combining the hybrid representation of explicit triangular grids and implicit texture material fields, combined with ambient lighting information and shadow prediction, high-quality rendering of digital humans under different postures and lighting conditions is achieved, solving the shortcomings of lighting rendering in the existing technology, and improving the universality and real-timeness of digital human technology.

CN119762710BActive Publication Date: 2025-05-16ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510275667.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-16
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing digital human technology has significant shortcomings in lighting rendering, including high cost, low universality, single-view limitations, high computing costs, insufficient real-time and universality, making it difficult to adapt to complex lighting and dynamic scenes, and the model and lighting are difficult to decouple, affecting the visual texture.

Method used

A hybrid representation of an explicit triangle mesh and an implicit texture material field is used to reconstruct a digital human body model of any posture, and re-light rendering is performed by combining adjusted ambient lighting information, camera perspective and shadow information predicted through human body segments.

Benefits of technology

It realizes high-quality and realistic digital human rendering images under different postures and lighting conditions, overcomes the shortcomings of traditional methods in geometric recovery, graphics pipeline compatibility and lighting processing, and improves the universality and real-time nature of digital human technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762710B_ABST
    Figure CN119762710B_ABST
Patent Text Reader

Abstract

The present invention discloses a human body model relighting method for generative digital human, which belongs to the field of digital human relighting rendering technology, and comprises the following steps: using a hybrid representation method combining explicit triangular meshes and implicit texture material fields to reconstruct a human body model of a digital human in any posture, so that the human body model includes a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse reflection albedo map; combining the adjusted arbitrary ambient lighting information, the camera viewing angle, and the shadow information predicted by the human body parts, the human body model is relighted to obtain a digital human rendering image in any posture and any lighting condition. In this way, the relighting of the human body model in different postures and different lighting conditions can be achieved, and a high-realistic digital human video can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital human re-lighting rendering, and in particular relates to a human body model re-lighting method for a generative digital human. Background Art

[0002] At present, digital human technology is booming and has been widely penetrated into many fields such as intelligent customer service, virtual anchors, online education, game entertainment, etc., becoming a new driving force for the development of the industry.

[0003] However, digital human technology has significant shortcomings in lighting rendering. Traditional graphics rendering often relies on high-end equipment, which not only leads to high costs but also poor universality; backlighting rendering is limited by single or multiple perspectives. When faced with complex lighting and dynamic scenes, it is difficult to accurately restore objects and lighting-related information. Although neural network rendering has made some progress, due to its high computational cost and many assumptions, its performance in real-time and versatility is not satisfactory. For example, the application of neural radiation field methods in real-time scenes is subject to many restrictions.

[0004] In the process of digital human model construction and rendering, many outstanding problems are also exposed. In geometric modeling, the accuracy of human joints is low, and artifacts are prone to occur. For example, when dealing with joint bending based on the signed distance field method, problems such as surface unevenness and texture distortion will occur, which will cause the movement to appear unnatural; the shadow estimation in the material reconstruction process relies on ray tracing technology, which is inefficient and ineffective, and cannot adapt to the dynamic changes of the material, which has a negative impact on the visual texture; in addition, it is difficult to decouple the model from the lighting, and it is difficult to meet the lighting requirements brought by the body movement.

[0005] Moreover, with the increasingly fierce market competition, various industries have put forward higher and higher requirements for digital humans, expecting them to adapt to different lighting scenarios and show high quality and personalized characteristics. Obviously, the existing technical means are difficult to meet the above requirements, and technological innovation is imminent.

[0006] In summary, given the many limitations of existing technologies, there is an urgent need for innovative three-dimensional digital human re-illumination technology to fill the current gaps, meet market demand, and help digital human technology move to a higher stage of development. Summary of the invention

[0007] In view of the above, the purpose of the present invention is to provide a human body model re-lighting method for generative digital humans, which can achieve re-lighting of the human body model under different postures and different lighting conditions to obtain high-realistic digital human videos.

[0008] To achieve the above-mentioned purpose of the invention, an embodiment provides a human body model re-illumination method for a generative digital human, comprising the following steps:

[0009] A hybrid representation method combining explicit triangular mesh and implicit texture material field is used to reconstruct a human body model of a digital human in any posture, so that the human body model includes a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse albedo map;

[0010] Combining the adjusted arbitrary environment lighting information, camera perspective, and shadow information predicted by different parts of the human body, the human body model is re-illuminated to obtain a digital human rendering image in any posture and lighting conditions.

[0011] Preferably, a hybrid representation method combining explicit triangular meshes and implicit texture material fields is used to reconstruct a human body model of a digital human in any posture, including:

[0012] Constructing an explicit triangular network, specifically, extracting a character triangular mesh in a standard space from a signed distance function field through a mesh extraction module, performing character posture deformation on the character triangular mesh in the standard space based on a given arbitrary posture through a first linear blend skinning module, and obtaining a character triangular mesh in a given arbitrary posture in an observation space;

[0013] Based on the implicit texture material field, texture material information is added to the character triangular mesh. Specifically, the front position mapping and the back position mapping are performed on the given arbitrary posture parameters to obtain the vertex information of each face. The offset mapping of each vertex on each face is extracted based on the vertex information of each face through the feature extraction module. The offset mapping is used to extract the high-frequency texture detail information of the vertex. The offset mapping of the front and back sides of each vertex is obtained by query and input into the first multi-layer perceptron. At the same time, the same vertex position information is sampled from the character triangular mesh under the given arbitrary posture and input into the first multi-layer perceptron. Based on the output of the first multi-layer perceptron, the triangular mesh with high-frequency texture detail information, the roughness map, the normal map, and the diffuse albedo map are output to form a human body model of the digital human.

[0014] Preferably, the arbitrary posture parameters are derived from SMPL-H posture parameters.

[0015] Preferably, the mesh extraction module uses FlexiCubes algorithm to extract the character's triangular mesh, and the feature extraction module uses U-net to extract the offset mapping of each face vertex.

[0016] Preferably, the initial value of the signed distance function field is constructed in the following way:

[0017] Constructing a learning framework for a signed distance function field, including a second linear hybrid skinning module, a reversible deformation field, a second multi-layer perceptron, and a third multi-layer perceptron, wherein query vertices of a character model in an observation space are converted to query vertices in a specification space through the second linear hybrid skinning module, and under arbitrary posture conditions, query vertices in the specification space are deformed to query vertices in arbitrary postures in the specification space through the reversible deformation field, and the query vertices in arbitrary postures are predicted and outputted through the second perceptron, and the signed distance function field is also generated through the third multi-layer perceptron to generate a rendered image under a given camera perspective;

[0018] The first loss function is used to train the learning framework to optimize the signed distance function field. The optimized signed distance function field is used as the initial value of the signed distance function field when constructing the explicit triangulation network. The first loss function includes pixel value loss, geometric smoothness loss, residual displacement field loss, and foreground mask loss.

[0019] Pixel value loss The difference between the rendered pixel value of the rendered image and the true pixel value of the ground-truth image is adopted;

[0020] Geometric smoothness loss The difference in the gradient of the signed distance function field is used to smooth the generated geometry, and the calculation formula is ,in, Represents the query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the distance;

[0021] Residual displacement field loss Refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ;

[0022] Foreground mask loss Refers to the rendered image outline The contour of the real image The intersection-and-union ratio is expressed as: , where r represents the set of all camera rays in the forward rendering process;

[0023] Then the first loss function ,in, 、 、 、 They represent the loss weights of the corresponding losses respectively.

[0024] Preferably, the shadow information predicted by different parts of the human body includes:

[0025] Divide the human body into multiple local areas according to the parts, and determine the coordinate information of the vertices of each local area in the local coordinate system, wherein the coordinate information includes position information and the corresponding light direction;

[0026] A separate neural network is constructed for each local area to predict the shadow information of each local area. Specifically, the postures of adjacent vertices in the local area are used as conditional inputs of the neural network. The coordinate information of each vertex is also input to predict the illumination visibility of each vertex, thereby obtaining the illumination visibility of each area. The illumination visibility of each area is multiplied to obtain the illumination visibility of the entire human body as the shadow information of the human body.

[0027] The neural network corresponding to each local area also needs to undergo parameter optimization before being applied. Specifically, a binary cross entropy loss function is used to supervise the learning of the neural network.

[0028] Preferably, the re-lighting process adopts a physically based rendering module, which adopts a Microfacet BRDF model, and generates a digital human rendering image by rendering according to the input human body model, shadow information, camera viewing angle, and ambient lighting information.

[0029] Preferably, the Microfacet BRDF model involved in the re-illumination and the network involved in the reconstruction of the digital human body model are also subjected to parameter optimization. Specifically, during the optimization, the illumination probe is used as the ambient illumination information, and the second loss function used includes the photometric loss and the regularization loss.

[0030] Luminosity loss Measures the difference between the rendered image and the real image in terms of brightness, color, and texture, expressed as:

[0031] ;

[0032] in, Represents a rendered image, is the target image, represents the normal map generated by the first multi-layer perceptron, represents the normal map estimated from the target image, represents the tone mapping function, represents the perceived loss, Indicates the balance coefficient, symbol represents the L1 norm;

[0033] Regularization Loss It aims to guide the generation of virtual avatars that conform to physical laws, expressed as:

[0034] ;

[0035] in, represents the geometric smoothness loss, which uses the difference in the gradient of the signed distance function field to smooth the generated geometric shape, and the calculation formula is: ,in, Represents the query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the L1 norm;

[0036] Represents the residual displacement field loss, which refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ;

[0037] represents the map smoothing loss, expressed as: ,in, Indicates a small offset, and Represents the diffuse albedo map and the roughness map, respectively, as intermediate variable maps The value is or , Represents a vertex The value in the map k;

[0038] represents the lighting consistency loss, expressed as: ;

[0039] in, Represents the three channels of the image, c Indicates that the value is Any channel of one of them, Represents the channel in the ambient light map c The value of

[0040] , , ,as well as They all represent the weight parameters of the corresponding losses.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] A hybrid representation method combining explicit triangular meshes and implicit texture material fields is used to reconstruct the human body model of a digital human in any posture, so that the constructed human body model has its own texture information, which is more conducive to subsequent re-lighting rendering;

[0043] During re-lighting rendering, the ambient lighting information and camera perspective can be adjusted, and combined with a digital human model in any posture, high-quality and highly realistic digital human rendering images in any posture and lighting conditions can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0045] Figure 1 is a flow chart of a human body model re-illumination method for a generative digital human provided in an embodiment;

[0046] Figure 2 is a detailed flowchart of the human body model re-illumination provided in the embodiment;

[0047] Figure 3 It is a schematic diagram of a learning framework for constructing a signed distance function field provided in an embodiment. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0049] In view of the technical defects that the lighting adjustment of existing digital human rendered images is limited and the generated digital human rendered images do not meet the posture, lighting and quality requirements, the embodiments provide a human body model re-lighting method for generative digital humans, which can support lighting editing and adjustment, can drive the digital human with motion data, and can dynamically adjust the lighting during rendering.

[0050] The human body model relighting method for a generative digital human provided by an embodiment of the present invention is described as: given a camera matrix representing a camera viewing angle , and attitude parameters in SMP-H format ,in Indicates i The posture parameters of the vertices are given an ambient light map representing the ambient light information. , the goal is to generate a continuous sequence of digital human video frames , so that the characters in the generated frame sequence can be driven by the pose parameters while keeping the rendering effect realistic and natural under given lighting. This is a supervised conditional generation problem. Specifically, we need to learn a mapping function , where each generated frame Not only must the input posture parameters be maintained The guidance consistency must also be ensured under the lighting The rendering effect under different lighting conditions is realistic and natural, that is, the generated image can correctly reflect the direction, intensity and color of the light source, while retaining the physical consistency of the digital human material and details. Figure 1 As shown, the re-illumination method provided in the embodiment includes the following steps:

[0051] S1, a hybrid representation method combining explicit triangular mesh and implicit texture material field is used to reconstruct the human body model of a digital human in any posture, so that the human body model contains a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse albedo map.

[0052] Human body model reconstruction refers to the reconstruction of the human body's geometric shape, which is used to describe the body shape, surface details and dynamic changes of the person, such as the geometric features of the person's body contour, clothing folds, etc. in different postures. For the human body model reconstruction process, traditional methods often focus on a single representation method, such as explicit representation or implicit representation, which exposes many problems when dealing with complex human body geometric structures and dynamic scenes. Among them, although the method based on explicit representation can intuitively present the geometric structure, when faced with extremely complex geometric topological changes, it is easy to cause detail loss due to its reliance on fixed grid resolution. Although the implicit representation method performs well in terms of flexibility and continuity, its computational complexity is relatively high, especially when it is converted to an explicit grid. In order to solve the problems existing in the above two methods, the geometric reconstruction scheme combining explicit and implicit methods proposed in the embodiment of the present invention specifically combines explicit triangular meshes with implicit texture material fields, cleverly integrating the advantages of both.

[0053] Among them, explicit triangular meshes can clearly depict complex geometric structures and are well compatible with traditional graphics pipelines, and can be directly applied to ray tracing and rendering processes. Especially when dealing with dynamic postures of the human body, combined with linear blended skinning technology, it can achieve smooth dynamic changes in geometric shapes and ensure the visual coherence of the model in different postures.

[0054] At the same time, the implicit texture material field focuses on capturing high-frequency details. Whether it is complex geometric features such as clothing wrinkles or the bends of human joints, the implicit material field can accurately present them. It can naturally cope with complex surfaces and topological changes, and its natural optimization ability makes the model more outstanding in detail performance. This organic combination of explicit and implicit successfully overcomes the difficulties of existing methods in geometry recovery and graphics pipeline compatibility, and lays a solid and precise geometric foundation for subsequent lighting processing and material modeling, so that the model can present more realistic and delicate effects under different lighting conditions.

[0055] like Figure 2 As shown, constructing an explicit triangular network refers to extracting a character triangular mesh in a standard space from a signed distance function field (SDF field for short) through a mesh extraction module, and deforming the character triangular mesh in the standard space based on a given arbitrary posture through a first linear mixed skinning module to obtain a character triangular mesh in a given arbitrary posture in the observation space. In an embodiment, the mesh extraction module uses the FlexiCubes algorithm to extract the character triangular mesh. The FlexiCubes algorithm gives vertices the ability to freely adjust their positions within the grid unit to which they belong. This feature shows great advantages when processing complex geometric shape areas (such as human joints). It can accurately fit the surface shape according to the actual curvature changes of the joints, greatly enhances the ability to capture complex surface geometric features, effectively improves the mesh fitting accuracy to the reference geometry, and significantly reduces the local distortion caused by fixed structures.

[0056] The FlexiCubes algorithm supports dynamic adjustment of meshing modes and can automatically optimize the meshing strategy based on the direction of surface curvature changes. For example, in areas with rich geometric features (such as clothing wrinkles or turning points of human body contours), the FlexiCubes algorithm will automatically increase the sampling density to ensure that no subtle changes in details are missed; while in relatively flat areas (such as most surfaces of the human torso), the resolution is kept low, effectively balancing computational efficiency while ensuring geometric accuracy. This adaptive resolution strategy performs particularly well in dynamic postures, ensuring geometric consistency and successfully avoiding artifacts and mesh distortion problems common in traditional methods. The mesh generated by FlexiCubes is more precise and flexible in topology, and also performs well in manifold and watertightness, providing a more ideal geometric foundation for generating high-quality three-dimensional human models and further improving the model's performance in complex scenarios.

[0057] Research has found that in explicit three-dimensional mesh construction, randomized SDF fields can lead to inaccurate character mesh generation. To this end, the initial value of the signed distance function field in the embodiment is constructed in the following way:

[0058] A learning framework for signed distance function fields is constructed and trained to reconstruct the geometry and motion of the human body. A grid-based inverse skinning method is used to calculate rigid motion, and a reversible neural deformation field is used to model non-rigid deformation. By optimizing these mappings, volume rendering can accurately reproduce the structure (represented by the SDF field) and motion characteristics of the human body. Figure 3 As shown in Figure 1, the learning framework includes a second linear mixed skinning module, a reversible deformation field, a second multi-layer perceptron, and a third multi-layer perceptron. Among them, the second linear mixed skinning module is used to provide rigid motion calculation and realize the transformation of the character space. However, the rigid bone transformation alone cannot accurately describe the deformation in human motion, so a reversible deformation field is also introduced to represent non-rigid deformation. The neural network corresponding to the reversible deformation field needs to realize the bidirectional mapping between the observation space and the specification space and maintain the cyclic consistency of the transformation. In the forward transformation process, the neural network maps the points in the specification space to the observation space, while in the inverse transformation, it is necessary to ensure that the original coordinates can be restored. Since non-rigid deformation is closely related to human posture, this paper adds the posture parameters of SMPL-H (a three-dimensional human model that describes human morphology and posture) to the input of the neural deformation field to improve the rationality of the deformation. At the same time, the skin weight of the query point is introduced as an additional input condition to improve the accuracy of the deformation field.

[0059] Specifically, in the learning framework, the query vertices of the character model in the observation space are converted to query vertices in the standard space through the second linear blend skinning module. Under arbitrary posture conditions, the query vertices in the standard space are deformed to query vertices in arbitrary postures in the standard space through the reversible deformation field. The query vertices in arbitrary postures are predicted by the second perceptron to output the signed distance function field, and the signed distance function field is also generated by the third multi-layer perceptron to generate a rendered image under a given camera perspective. The learning framework is trained using the first loss function to optimize the signed distance function field, and the optimized signed distance function field is used as the initial value of the signed distance function field when constructing an explicit triangulation network, wherein the first loss function includes pixel value loss, geometric smoothness loss, residual displacement field loss, and foreground mask loss.

[0060] Pixel value loss Uses the rendered pixel value of the rendered image The actual pixel value of the true value image The difference between them ensures that the rendering result is as close to the real image as possible, expressed as: ,in, Represents the camera light during forward rendering, represents a collection of rays;

[0061] Geometric smoothness loss The difference in the gradient of the signed distance function field is used to smooth the generated geometry to ensure that the mesh does not have unnecessary sharp changes. The calculation formula is: ,in, Represents the query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the distance;

[0062] Residual displacement field loss Refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ;

[0063] Foreground mask loss Refers to the rendered image outline The contour of the real image The intersection-and-union ratio ensures that the contour of the rendered image is as consistent as possible with the real contour, which is expressed as: ,in, r Represents the camera ray during forward rendering;

[0064] Then the first loss function ,in, 、 、 、 They represent the loss weights of the corresponding losses respectively.

[0065] like Figure 2 As shown, when constructing a human body model, texture material information is added to the character's triangular mesh based on an implicit texture material field, and specifically, the front position mapping and the back position mapping are performed on a given arbitrary posture parameter (SMPL-H posture parameter can be used) to obtain the vertex information of each face, and the offset mapping of each vertex on each face is extracted based on the vertex information of each face through a feature extraction module, and the offset mapping is used to extract the high-frequency texture detail information of the vertex, and the offset mapping of the front and back sides of each vertex is obtained by querying and input into the first multi-layer perceptron, and at the same time, the same vertex position information is sampled from the character's triangular mesh under a given arbitrary posture and input into the first multi-layer perceptron, and the first multi-layer perceptron outputs a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse albedo map to form a human body model of a digital human.

[0066] In the embodiment, the feature extraction module uses UNet to process the standard SMPL-H template vertex information, obtains the position map through orthogonal projection, and outputs the vertex offset map after UNet extracts the features, simplifies the high-frequency feature generation process, and provides high-quality surface geometry information for physical rendering and material decomposition. The first multi-layer perceptron extracts the high-frequency vertex offset through bilinear interpolation, and directly inputs the vertex coordinates of the character triangular mesh from the observation space of the explicit triangular mesh part. After processing, the high-frequency detail geometry and material properties are predicted to ensure the geometry and material performance quality of the human body model.

[0067] In the human body model of a digital human with arbitrary postures reconstructed by hybrid representation provided by the present invention, the accuracy of the human body model is improved by explicit triangular mesh, implicit texture material field enhancement, and geometric detail optimization and fusion. The FlexiCubes algorithm is introduced into the explicit triangular mesh to give the vertex the ability to freely adjust the position, enhance the capture of complex surface geometric features, improve the fitting accuracy, and reduce local distortion. Its dynamic partitioning mode is optimized according to the change of surface curvature, and the adaptive resolution strategy balances geometric accuracy and computational efficiency, ensuring geometric consistency under dynamic postures, avoiding artifacts and mesh distortion, and generating a better mesh topology structure.

[0068] Implicit texture material field enhancement combines explicit triangular meshes with implicit texture fields to extract high-frequency features based on UNet. UNet extracts high-frequency texture details of vertices by inputting SMPL-H pose parameters and SDF fields, and generates texture details by combining the geometric shapes determined by explicit triangular meshes at complex joints, avoiding detail loss or distortion, and making the model look more realistic.

[0069] The optimization and fusion of geometric details are based on the mesh extracted by the SDF field. After UNet extracts the high-frequency detail map, the first multi-layer perceptron optimizes the fused vertex offsets based on the surrounding vertices and global features. For example, when processing clothing wrinkles, the explicit mesh and implicit material field work together to improve geometric consistency and rendering expressiveness.

[0070] S2, combining the adjusted arbitrary environment lighting information, the camera perspective, and the shadow information predicted by the human body parts, re-lights the human body model to obtain a digital human rendering image under arbitrary postures and arbitrary lighting conditions.

[0071] In the embodiment, the shadow information of human body parts is predicted to solve the problems of high computational overhead, low efficiency and poor effect in processing complex human body posture shadows in traditional shadow estimation methods. The specific process is:

[0072] The human body is divided into multiple (e.g., 15) local regions according to the parts, and the coordinate information of the vertices of each local region in the local coordinate system is determined, where the coordinate information includes the position information and the corresponding light direction. Specifically, given the query vertex position in the global observation space and light direction , they need to be transformed into the local coordinate system of each local area. The local coordinates are transformed by the bone transformation matrix The calculation formula is ,in, i Represents the vertex index. This conversion process is the basis for subsequent regionalization processing, which enables each local area to calculate the light visibility in its own coordinate system, thereby reducing the complexity of the problem.

[0073] A separate neural network is constructed for each local region to predict the light visibility of each local region as shadow information. The light visibility of a local region depends not only on the geometry of the region itself, but also on the posture of the adjacent joints, specifically the posture of the adjacent vertices in the local region. As the conditional input of the neural network, the coordinate information of each vertex is also input , to predict the lighting visibility of each vertex ,Right now , and then get the illumination visibility of each area, multiply the illumination visibility of each area, that is , M is the total number of vertices, and the illumination visibility of the entire human body is obtained As the shadow information of the human body.

[0074] The neural network corresponding to each local area also needs to be optimized before being applied, specifically using the binary cross entropy loss function To supervise the learning of neural networks, is the true value of illumination visibility. This training method enables the neural network to accurately learn the light occlusion relationship in the local area and improve the accuracy of shadow prediction.

[0075] Compared with traditional ray tracing methods, the present invention transforms the global occlusion relationship in complex fields into a local problem through partitioned local modeling, which greatly reduces the computational cost. Since the geometric changes in each local area are relatively smooth and regular, the neural network can capture local occlusion features more efficiently. In addition, this method can better adapt to the performance of dynamic human bodies under different lighting conditions, thereby generating more realistic and efficient shadow effects.

[0076] In the embodiment, when re-illuminating the human body model, Figure 2As shown, a physically based rendering module is used, which adopts a Microfacet BRDF model, and generates a digital human rendering image by rendering according to an input human body model, shadow information, camera viewing angle, and ambient lighting information.

[0077] The Microfacet BRDF model simulates the interactive behavior of light on the surface of the material through micro-geometry, and decomposes it into diffuse reflection and specular reflection. In the present invention, it is assumed that the metalness of all materials is 0, that is, they are all non-metallic materials, which simplifies the model's description of the material and enables more focus on the independent processing of diffuse reflection and specular reflection. At the same time, the fixed Fresnel coefficient is 0.04, which is used to describe the reflection characteristics of light on the surface of non-metallic materials, reducing the complexity of parameter optimization without significantly reducing the visual quality. In the process of light interaction, the diffuse reflection part is controlled by the diffuse reflectivity, which indicates the uniform scattering ability of the material surface to light, which is independent of the wavelength of the incident light and does not produce obvious highlights; the specular reflection part is determined by the specular roughness and the Fresnel coefficient, taking into account the influence of micro-roughness on highlights and the reflectivity changes of the incident light at different angles. Through this comprehensive modeling, the Microfacet BRDF model can truly simulate the interactive behavior of complex material light, and combined with the simplified assumptions in the present invention, effectively balance the physical accuracy and computational efficiency in material reconstruction.

[0078] The Microfacet BRDF model involved in the re-lighting and the network involved in reconstructing the digital human body model have also undergone parameter optimization. During the specific optimization, light probes are used as environmental lighting information to simulate the lighting distribution in the training data set. Light probes can represent the lighting distribution in the scene in a compact and efficient way, and return the light intensity distribution in the corresponding direction according to the sampling direction, which makes dynamic adjustment possible during the training process. After the model training is completed, in order to support the re-lighting operation of the actual scene, the light probe is replaced with a predefined environment map. The environment map is an HDR texture based on spherical expansion, which can accurately describe the real lighting environment, thereby simulating the performance of materials under different scene lighting conditions, so that digital humans can present realistic effects in various lighting environments.

[0079] When optimizing parameters, the second loss function also includes photometric loss and regularization loss. Measures the difference between the rendered image and the real image in terms of brightness, color, and texture, expressed as:

[0080] ;

[0081] in, Represents a rendered image, is the target image, represents the normal map generated by the first multi-layer perceptron, represents the normal map estimated from the target image, represents the tone mapping function, represents the perceived loss, Indicates the balance coefficient, symbol represents the L1 norm;

[0082] Regularization Loss It aims to guide the generation of virtual avatars that conform to physical laws, expressed as:

[0083] ;

[0084] in, represents the geometric smoothness loss, which uses the difference in the gradient of the signed distance function field to smooth the generated geometric shape, and the calculation formula is: ,in, Represents the query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the L1 norm;

[0085] Represents the residual displacement field loss, which refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ;

[0086] represents the map smoothing loss, expressed as: ,in, Indicates a small offset, and Represents the diffuse albedo map and the roughness map, respectively, as intermediate variable maps The value is or , Represents a vertex The value in the map k;

[0087] represents the lighting consistency loss, expressed as: ;

[0088] in, Represents the three channels of the image, c Indicates that the value is Any channel of one of them, Represents the channel in the ambient light map c The value of

[0089] , , ,as well as By adjusting these weight coefficients, a balance can be made between pixel accuracy, material simplicity, and shadow smoothness according to specific needs to adapt to different application scenarios and performance requirements.

[0090] After the parameter optimization is completed, Figure 2 The human body model re-lighting process shown in the figure allows users to flexibly edit the lighting of a specific person by providing arbitrary ambient lighting information and specifying the person's posture and camera position. The re-lighting process can dynamically adjust visual effects such as shadows and reflections according to the input lighting conditions, thereby generating a highly realistic digital human video that conforms to the input posture and is rendered under the target lighting conditions.

[0091] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A human body model re-illumination method for generative digital humans, characterized in that: The following steps are involved: A hybrid representation method combining explicit triangular mesh and implicit texture material field is used to reconstruct the human body model of a digital human in any posture, so that the human body model contains a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse albedo map, including: Constructing an explicit triangular network, specifically extracting a character triangular mesh in a standard space from a signed distance function field through a mesh extraction module, deforming the character triangular mesh in the standard space based on a given arbitrary posture through a first linear mixed skinning module, and obtaining a character triangular mesh in a given arbitrary posture in an observation space; adding texture material information to the character triangular mesh based on an implicit texture material field, specifically performing front position mapping and back position mapping on given arbitrary posture parameters to obtain vertex information of each face, extracting an offset mapping of each vertex based on the vertex information of each face through a feature extraction module, and the offset mapping is used to extract high-frequency texture detail information of the vertex, obtaining the offset mapping of the front and back faces of each vertex through a query and inputting it into a first multi-layer perceptron, and sampling the same vertex position information from the character triangular mesh in a given arbitrary posture and inputting it into the first multi-layer perceptron, and outputting a triangular mesh with high-frequency texture detail information, a roughness map, a normal map, and a diffuse albedo map based on the first multi-layer perceptron to form a human body model of a digital human; Combining the adjusted arbitrary environment lighting information, camera perspective, and shadow information predicted by different parts of the human body, the human body model is re-illuminated to obtain a digital human rendering image in any posture and lighting conditions.

2. The human body model re-illumination method for generative digital human according to claim 1, characterized in that: Arbitrary attitude parameters come from SMPL-H attitude parameters.

3. The human body model re-illumination method for generative digital human according to claim 1, characterized in that: The mesh extraction module uses the FlexiCubes algorithm to extract the character's triangular mesh, and the feature extraction module uses UNet to extract the offset mapping of each face vertex.

4. The human body model re-illumination method for generative digital human according to claim 1, characterized in that: The initial value of the signed distance function field is constructed as follows: Constructing a learning framework for a signed distance function field, including a second linear hybrid skinning module, a reversible deformation field, a second multi-layer perceptron, and a third multi-layer perceptron, wherein query vertices of a character model in an observation space are converted to query vertices in a specification space through the second linear hybrid skinning module, and under arbitrary posture conditions, query vertices in the specification space are deformed to query vertices in arbitrary postures in the specification space through the reversible deformation field, and the query vertices in arbitrary postures are predicted and outputted through the second perceptron, and the signed distance function field is also generated through the third multi-layer perceptron to generate a rendered image under a given camera perspective; The first loss function is used to train the learning framework to optimize the signed distance function field. The optimized signed distance function field is used as the initial value of the signed distance function field when constructing the explicit triangulation network. The first loss function includes pixel value loss, geometric smoothness loss, residual displacement field loss, and foreground mask loss. Pixel value loss The difference between the rendered pixel value of the rendered image and the true pixel value of the ground-truth image is adopted; Geometric smoothness loss The difference in the gradient of the signed distance function field is used to smooth the generated geometry, and the calculation formula is: ,in, Represents query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the distance; Residual displacement field loss Refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ; Foreground mask loss Refers to the rendered image outline The contour of the real image The intersection-over-union ratio is expressed as :,in, r Represents the camera light during forward rendering; Then the first loss function ,in, 、 、 、 They represent the loss weights of the corresponding losses respectively.

5. The human body model re-illumination method for generative digital human according to claim 1, characterized in that: Shadow information predicted by human body parts, including: Divide the human body into multiple local areas according to the parts, and determine the coordinate information of the vertices of each local area in the local coordinate system, wherein the coordinate information includes position information and the corresponding light direction; A separate neural network is constructed for each local area to predict the shadow information of each local area. Specifically, the postures of adjacent vertices in the local area are used as conditional inputs of the neural network. The coordinate information of each vertex is also input to predict the illumination visibility of each vertex, thereby obtaining the illumination visibility of each area. The illumination visibility of each area is multiplied to obtain the illumination visibility of the entire human body as the shadow information of the human body. The neural network corresponding to each local area also needs to undergo parameter optimization before being applied. Specifically, a binary cross entropy loss function is used to supervise the learning of the neural network.

6. The human body model re-illumination method for generative digital human according to claim 1, characterized in that: The re-lighting process adopts a physically based rendering module, which adopts a Microfacet BRDF model and generates a digital human rendering image by rendering according to the input human body model, shadow information, camera perspective, and ambient lighting information.

7. The human body model re-illumination method for generative digital humans according to claim 6, characterized in that: The Microfacet BRDF model involved in the re-illumination and the network involved in the reconstruction of the human body model of the digital human are also optimized. During the specific optimization, the illumination probe is used as the ambient illumination information, and the second loss function used includes the photometric loss and the regularization loss. Luminosity loss Measures the difference between the rendered image and the real image in terms of brightness, color, and texture, expressed as: ; in, Represents a rendered image, is the target image, represents the normal map generated by the first multi-layer perceptron, represents the normal map estimated from the target image, represents the tone mapping function, represents the perceived loss, Indicates the balance coefficient, symbol represents the L1 norm; Regularization Loss It aims to guide the generation of virtual avatars that conform to physical laws, expressed as: ; in, represents the geometric smoothness loss, which uses the difference in the gradient of the signed distance function field to smooth the generated geometric shape, and the calculation formula is: ,in, Represents the query vertex The signed distance function field The gradient of represents three-dimensional space, The symbol represents the square of the L1 norm; Represents the residual displacement field loss, which refers to querying vertices when performing spatial transformation The amplitude of the corresponding residual displacement field The regularization of is expressed as: ; represents the map smoothing loss, expressed as: ,in, Indicates a small offset, and Represents the diffuse albedo map and the roughness map, respectively, as intermediate variable maps The value is or , Represents a vertex In the map k The value of ; represents the lighting consistency loss, expressed as: ; in, Represents the three channels of the image, c Indicates that the value is Any channel of one of them, Represents the channel in the ambient light map c The value of , , ,as well as They all represent the weight parameters of the corresponding losses.

Citation Information

Patent Citations

  • Reconstruction method and device for relightable human body implicit model

    CN116051696A

  • Digital human reconstruction method with high-fidelity triangular mesh and material texture mapping

    CN117649490A