Method and system for estimating geometry, illumination and material of human body from multi-view image

Through the joint optimization method of neural radiation field rendering and inverse rendering, combined with lighting and material estimation in the geometric pre-training and joint optimization stages, the problem of complex materials and indirect lighting modeling of human bodies is solved, and a more realistic inverse rendering effect is achieved.

CN119941949AActive Publication Date: 2025-05-06WUHAN ZONGHENG TIANDI SPACE INFORMATION TECH CO LTD

Patent Information

Application Number
CN202411951382.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively estimate the geometry, lighting and materials of the human body, especially in terms of complex materials and indirect lighting modeling of human bodies.

Method used

The joint optimization method of neural radiation field rendering and inverse rendering is used to pre-train the human geometry through the geometry pre-training stage, and then the lighting and material are estimated in the joint optimization stage, and the geometry initialization results are fine-tuned. Use anisotropic material model and the exit radiation of indirect reflection points to model complex human materials and indirect light.

Benefits of technology

It solves the problem of inaccurate estimation of human body materials, reduces the cost of modeling indirect lighting and lighting visibility, and achieves a more realistic inverse rendering effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941949A_ABST
    Figure CN119941949A_ABST
Patent Text Reader

Abstract

The invention designs a method and a system for estimating the geometry, the illumination and the material of a human body from a multi-view image captured under unknown illumination on the basis of the thought of joint optimization of neural radiation field rendering and inverse rendering, and the method is mainly divided into a geometry pre-training stage and a joint optimization stage. Firstly, human body geometry is pre-trained to obtain a good initialization result, then illumination and materials are estimated, and the geometric initialization result is finely adjusted. When the material is estimated, an anisotropic material model is used for modeling a complex human body material. When the illumination is estimated, an additional training stage is not used for modeling indirect illumination, but the emergent radiance of an indirect reflection point is regarded as the indirect illumination, and the illumination visibility is estimated to distinguish direct illumination from the indirect illumination. According to the method, the problem of inaccurate estimation of the complex material of the human body is solved, geometry, illumination and the material can be effectively decoupled from the multi-view image of the human body, and the sense of reality of a body weight illumination result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of neural radiation field rendering and inverse rendering, and specifically relates to a method and system for estimating human body geometry, lighting and material from multi-view images. Background Art

[0002] Inverse rendering is a classic task in computer vision and computer graphics, and is very important for applications such as digitizing the real world, synthesizing new views, relighting, and material editing. Inverse rendering is ill-posed, meaning that different geometries, materials, and lighting conditions may result in similar appearances. With the rapid development of differentiable rendering and implicit neural representations, methods for inverse rendering of small object scenes have emerged. These methods usually only consider using simple color values ​​to represent the appearance, without using more realistic physically based rendering.

[0003] Recent methods represent geometry and materials as neural radiance fields. These methods use coordinate-based multi-layer perceptrons and recover geometry, materials, and lighting by reducing the rendering loss between the rendered image and the input image. Most of these methods work well for small objects because their geometry and materials are simple. However, these methods cannot be directly applied to inverse rendering of the human body.

[0004] There are two main challenges in human inverse rendering: ① The human material is relatively complex. For example, human skin is a layered material, and the geometry of human hair is difficult to reconstruct. Previous inverse rendering methods assumed that the estimated material is isotropic. However, human hair and clothing materials are anisotropic, and their physical properties are related to the viewing direction. ② Calculating indirect lighting and lighting visibility is crucial to correctly estimating materials. Most methods ignore indirect lighting or use additional stages and resources to train indirect lighting and lighting visibility. Summary of the invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for estimating human body geometry, lighting and material from multi-view images. The human body geometry is pre-trained in the geometry pre-training stage to obtain a good initialization result, and then the lighting and material are estimated in the joint optimization stage and the geometry initialization result is fine-tuned. This solves the problem of inaccurate human body material estimation, reduces the cost of modeling indirect lighting and lighting visibility, and achieves a more realistic inverse rendering effect.

[0006] According to one aspect of the present invention, a method for estimating human body geometry, lighting and material from multi-view images is provided, comprising:

[0007] Obtaining a three-dimensional human body model to be estimated and performing multi-view rendering;

[0008] The rendered data is input into a trained human body geometry, lighting and material estimation network, and the estimated human body geometry, lighting and material are output; wherein the training of the human body geometry, lighting and material estimation network includes:

[0009] Collecting 3D human body model data, and rendering multi-view human body images and corresponding camera pose parameters based on the collected 3D human body model data;

[0010] Perform geometric pre-training: Use the rendered data as input, use the signed distance function to implicitly represent the human body surface, use the spherical tracking algorithm to obtain surface points, and calculate the appearance color value corresponding to the surface point through the neural radiation field; select the loss function of geometric pre-training, and gradually optimize the signed distance function and neural radiation field by minimizing the loss function;

[0011] Perform joint optimization: use the pre-trained signed distance function and spherical tracking algorithm to obtain the surface points of the human body, input them into the designed visibility network and material network, and output the lighting visibility and material; use the lighting visibility and pre-trained neural radiation field to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function; calculate the color value of the surface point through the rendering equation; select the loss function for joint optimization, and jointly optimize the signed distance function, neural radiation field, visibility network, and material network by minimizing the loss function.

[0012] As a further technical solution, geometric pre-training is also performed, including:

[0013] The human body surface is implicitly represented as a set of zero values ​​of the signed distance function;

[0014] According to each input image and its corresponding camera pose, the pixel ray corresponding to each pixel point is obtained;

[0015] Sampling is performed on the pixel ray through the spherical tracing algorithm, the SDF value of each sampling point is calculated, and when the SDF value of the sampling point is within a preset threshold close to 0, the sampling point is regarded as the surface point corresponding to the pixel ray;

[0016] After obtaining the coordinates of the surface points, they are input into the designed neural radiation field to calculate the appearance color value.

[0017] As a further technical solution, the signed distance function and the neural radiation field are both represented by multi-layer perceptrons. During geometric pre-training, the weight parameters of the two multi-layer perceptrons are trained by narrowing the gap between the predicted values ​​and the true values, and the masked true values ​​are used for supervision.

[0018] As a further technical solution, joint optimization is also carried out, including:

[0019] When estimating illumination, the global illumination is decomposed into direct illumination and indirect illumination, the direct illumination is fitted using a spherical Gaussian function, the indirect illumination is calculated using the outgoing irradiance of the indirect bounce point, and visibility is used to indicate the visibility of the direct illumination to the surface points in the illumination direction;

[0020] When estimating materials, an anisotropic material model is used to calculate the bidirectional reflectance distribution function BRDF, and the rendering equation is used to calculate the color value based on physical rendering.

[0021] As a further technical solution, the jointly optimized loss function includes geometric pre-training loss, image reconstruction loss and smoothing loss.

[0022] According to one aspect of the present invention, a system for estimating human body geometry, lighting and material from multi-view images is provided, comprising:

[0023] A data acquisition module, used to acquire a three-dimensional human body model to be estimated and perform multi-view rendering;

[0024] A data estimation module is used to input the rendered data into a trained human body geometry, lighting and material estimation network, and output estimated human body geometry, lighting and material; wherein the training of the human body geometry, lighting and material estimation network includes:

[0025] Collecting 3D human body model data, and rendering multi-view human body images and corresponding camera pose parameters based on the collected 3D human body model data;

[0026] Perform geometric pre-training: Use the rendered data as input, use the signed distance function to implicitly represent the human body surface, use the spherical tracking algorithm to obtain surface points, and calculate the appearance color value corresponding to the surface point through the neural radiation field; select the loss function of geometric pre-training, and gradually optimize the signed distance function and neural radiation field by minimizing the loss function;

[0027] Perform joint optimization: use the pre-trained signed distance function and spherical tracking algorithm to obtain the surface points of the human body, input them into the designed visibility network and material network, and output the lighting visibility and material; use the lighting visibility and pre-trained neural radiation field to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function; calculate the color value of the surface point through the rendering equation; select the loss function for joint optimization, and jointly optimize the signed distance function, neural radiation field, visibility network, and material network by minimizing the loss function.

[0028] According to one aspect of the present invention specification, there is provided a device for estimating human body geometry, lighting and material from multi-view images, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method for estimating human body geometry, lighting and material from multi-view images.

[0029] According to one aspect of the present invention, a computer readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for estimating human body geometry, lighting and material from multi-view images are implemented.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention designs a method for estimating human body geometry, lighting and material from multi-view images captured under unknown lighting based on the idea of ​​joint optimization of neural radiation field rendering and inverse rendering, which is mainly divided into a geometric pre-training stage and a joint optimization stage. In the geometric pre-training stage, the present invention first pre-trains the human body geometry to obtain a good initialization result, and then estimates the lighting and material and fine-tunes the geometric initialization result in the joint optimization stage. When estimating the material, the present invention uses an anisotropic material model to model complex human body materials. In addition, when estimating the lighting, the present invention does not use an additional training stage to model indirect lighting, but instead regards the outgoing radiance of the indirect reflection point as indirect lighting, and estimates the lighting visibility to distinguish between direct lighting and indirect lighting. The present invention solves the problem of inaccurate estimation of complex human body materials, and can effectively decouple geometry, lighting and material from multi-view images of the human body, thereby improving the realism of the human body re-lighting results. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction is given below to the drawings used in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 It is a schematic flow chart of a method for estimating human body geometry, lighting and material from multi-view images provided by an embodiment of the present invention.

[0034] Figure 2 It is a schematic diagram of a framework for estimating human body geometry, lighting and material from multi-view images provided by an embodiment of the present invention.

[0035] Figure 3 Schematic diagram of different network structures provided by embodiments of the present invention. DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention are arbitrarily combined with each other to form a new technical solution. This combination is not restricted by the sequence of steps and / or the structural composition mode, but must be based on the ability of ordinary technicians in this field to achieve. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that this combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0037] The embodiment of the present invention provides a method for estimating human body geometry, illumination and material from multi-view images. First, a three-dimensional human body model to be estimated is obtained and multi-view rendering is performed; then, the rendered data is input into a trained human body geometry, illumination and material estimation network, and the estimated human body geometry, illumination and material are output. The human body geometry, illumination and material estimation network of the present invention mainly includes a geometric pre-training stage and a joint optimization stage.

[0038] See also Figures 1 to 3 As a preferred embodiment, the method for estimating human body geometry, illumination and material from multi-view images according to the embodiment of the present invention comprises the following steps:

[0039] Step 1: Collect 3D human models and construct a data set. Optionally, purchase and download publicly available 3D human models with complex materials such as different ages, different skin colors, and different ethnic costumes from the Internet.

[0040] Step 2: Import the 3D human body model data collected in step 1 into the rendering engine Blender and use different environment maps as lighting conditions. Write a python script to render multi-view human body images and corresponding camera pose parameters.

[0041] Step 3: Geometric pre-training phase: The data processed in step 2 (multi-view images and camera poses) are used as input, the human body surface is implicitly represented using the signed distance function, and the surface points are obtained using the spherical tracking algorithm. The appearance color values ​​corresponding to the surface points are then calculated using the neural radiation field.

[0042] Step 4: Use the image reconstruction loss as the loss function of the predicted value and the true value, and gradually optimize the signed distance function and the neural radiation field in step 3 by minimizing the loss function.

[0043] Step 5, joint optimization phase. It is mainly divided into the illumination calculation module and the material estimation module. Use the pre-trained signed distance function and spherical tracking algorithm in step 3 to obtain the surface points of the human body, and then input them into the designed visibility network and material network to output the illumination visibility and material. Then use the illumination visibility and the pre-trained neural radiation field in step 3 to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function BRDF. Finally, the color value of the surface point is calculated through the rendering equation.

[0044] Step 6: Use the loss function and smoothing loss in step 4 as the loss function in the joint optimization stage, and jointly optimize the signed distance function, neural radiation field in step 3, and the visibility network and material network in step 5 by minimizing the loss function.

[0045] The step 2 is specifically described as follows:

[0046] For each human body model, 100 images of different perspectives are randomly rendered on a sphere centered on the human body, of which 80 images are used as training sets and 20 images are used as test sets. The resolution of each image is 512 × 512. In addition, the existing mature algorithm is used to obtain the mask of each human body image.

[0047] The step 3 is described in detail as follows:

[0048] The human body surface is implicitly represented as a set of zero values ​​of the signed distance function (SDF). According to each input image and its corresponding camera pose, the pixel ray o+td corresponding to each pixel point can be obtained, where o represents the camera center, d represents the ray direction, t represents the position of the point on the pixel ray, and t=0 represents the starting point of the ray. 2048 pixel rays are sampled for each image.

[0049] Then, the spherical tracing algorithm is used to sample the pixel ray and calculate the SDF value of each sampling point. The specific process of the spherical tracing algorithm is to randomly initialize a t value as the sampling starting point, calculate the SDF value s of the sampling point, and if s is not within the preset threshold approaching 0, continue sampling, use t+s as the t value of the next sampling point, and repeat the operation. When the SDF value of the sampling point is within the preset threshold approaching 0, this point is approximately regarded as the surface point corresponding to the pixel ray. After obtaining the coordinates of the surface point, it is input into the designed neural radiation field together with the normal, viewpoint direction and eigenvector to calculate the appearance color value.

[0050] Signed distance function S and neural radiation field L oThey are all represented by multi-layer perceptrons, which are specifically defined as:

[0051] v sdf ,f=S(x;Θ)

[0052] C r =L o (x,n,f,d;Γ)

[0053] Where x represents the coordinates of the surface point, n represents the surface normal, f represents the feature vector, and d represents the viewpoint direction. Θ and Γ represent the weight parameters of two multi-layer perceptrons, respectively. In addition, the embodiment of the present invention uses the gradient of S(x;Θ) to calculate the surface normal value n, that is,

[0054] The step 4 is specifically described as follows:

[0055] The weight parameters Θ and Γ of the two multi-layer perceptrons in step 3 are trained by narrowing the gap between the calculated appearance color value and the true value. In addition, the embodiment of the present invention also uses the mask true value for supervision. The final loss function is defined as follows:

[0056] L geo =L color +λ1L mask +λ2L reg

[0057] Image reconstruction loss L color The definition is as follows:

[0058]

[0059] in, represents the color value calculated by the neural radiance field in step 3, C k represents the true color value in step 2, m represents the number of surface points, and R represents the L1 loss.

[0060] Mask loss L mask The definition is as follows:

[0061]

[0062] Among them, M p represents the mask value of the surface point p in step 2, S(p; Θ) represents the SDF value in step 3, BCE represents the cross entropy loss, α is a hyperparameter, and the preset value of α is 50.

[0063] We also use the regularization loss L reg To constrain the signed distance field, it is defined as follows:

[0064]

[0065] in, represents the gradient of the signed distance field at the surface point p in step 3.

[0066] The step 5 is mainly divided into a lighting calculation module and a material estimation module, which are described as follows:

[0067] For the illumination estimation module, the global illumination is decomposed into direct illumination and indirect illumination. The direct illumination is fitted by a spherical Gaussian function, which is defined as:

[0068] G(v;ξ,λ,μ)=μe λ(v·ξ-1)

[0069] Where v∈R 3 represents the input of the function, ξ∈S 2 represents the center direction of the spherical Gaussian, λ∈R + represents the decay rate of the spherical Gaussian, Indicates the illumination amplitude value at the center of the spherical Gaussian function. A total of M = 128 spherical Gaussian functions are used to represent the illumination direction ω i Direct light L d (ω i ), which is defined as follows:

[0070]

[0071] Indirect lighting is calculated by the outgoing irradiance of the indirect bounce point. Specifically, starting from the current surface point x, along the lighting direction ω i Moving forward, obtaining indirect bounce points through spherical tracking algorithm In the embodiment of the present invention, a total of 32 light directions are sampled, a t value in the light direction is randomly initialized as the sampling starting point, and the SDF value s of the sampling point is calculated. If s is not within the preset threshold close to 0, sampling is continued, and t+s is used as the t value of the next sampling point. When the SDF value of the sampling point is within the preset threshold close to 0, it means that with the surface point x as the starting point, it can collide with other surface points along the light direction, that is, the surface point x will receive other surface points along -ω i Indirect rays emitted, other surface points are considered indirect bounce points Indirect rebound point The outgoing irradiance is used as the indirect illumination of the surface point x. o Calculated The color value is Therefore, the indirect light L ind (x,ω i ) is defined as follows:

[0072]

[0073] in, Indirect rebound point The normal line of Indirect rebound point The feature vector of .

[0074] In addition, visibility is used to indicate the direct light in the light direction ω i For the visibility of surface point x, the visibility v is estimated using a multilayer perceptron, which is defined as follows:

[0075] v=Φ(β(x),d;θ)

[0076] Among them, β represents the position encoding, x represents the surface point, d represents the viewing direction, and θ represents the weight of the multi-layer perceptron. Finally, the global illumination L i (x,ω i ) by direct light L d (ω i ), indirect lighting L ind (x,ω i ) and visibility v are calculated using the following formula:

[0077] L i (x,ω i )=vL d (ω i )+(1-v)L ind (x,ω i )

[0078] For the material estimation module, an anisotropic material model is used to calculate the bidirectional reflectance distribution function BRDF, thereby characterizing the physical properties of the complex material of the human body, and making the rendered color value more realistic through a physically based rendering method. Physically based rendering is a shading model that complies with physical principles. The embodiment of the present invention adopts the cook-torrance model based on the microplane principle, which is widely used in the industry. The model divides the bidirectional reflectance distribution function BRDF into a diffuse reflection component and a specular reflection component. It is defined as follows:

[0079]

[0080] Among them, l represents the incident light direction, v represents the viewpoint direction, diffuse represents diffuse reflection, which is estimated by the material network, n represents the normal direction, D represents the normal distribution function, F represents the Fresnel equation, and G represents the geometric occlusion function. D, F, and G are all calculated using anisotropic material models. The specific calculation process is as follows:

[0081] Use the GTR function to calculate the normal distribution function D and the GGX function to calculate the geometric shading function G. The calculation formula is as follows

[0082]

[0083] Where t represents the tangent direction, b represents the binormal direction, h represents the half-angle vector of the incident light direction l and the viewpoint direction v, and the final G value is G(l)G(v). In addition, α x and α y Represents the roughness along the tangent direction and along the normal direction respectively. x and α y The calculation method is as follows:

[0084] α y =r 2 k

[0085] where k aniso Represents the degree of anisotropy parameter, set as a trainable parameter and optimized together with the material, r estimates the roughness. aniso =0, α x =α y , the anisotropic model becomes an isotropic model, so k aniso The range of is (0, 1). In addition, the incident light direction ω is also used o And normal n to calculate the tangent direction t and binormal direction b, the calculation formula is as follows:

[0086] t=ω o -(ω o ·n)n,b=n×t

[0087] The calculation formula of Fresnel equation F is as follows:

[0088] F(v,h)=F0+(1-F0)(1-(v·h)) 5

[0089] Where F0 represents the specular albedo, estimated by the material network, v is the view direction, and h represents the half-angle vector between the incident light direction l and the view direction v.

[0090] The materials used in the calculation include diffuse albedo a, specular albedo s, roughness r, subsurface scattering sss and glossiness sh. The material is also calculated using a multi-layer perceptron, which is defined as follows:

[0091] {a,r,sss,s,sh}=M(β(x);ψ)

[0092] Among them, β represents the position encoding, x represents the surface point, and ψ represents the weight of the multi-layer perceptron.

[0093] After obtaining the global illumination and BRDF, the color value L based on physical rendering is finally calculated through the rendering equation surf (x,ω i ,ω o ). The rendering equation is defined as follows:

[0094] L surf (x,ω i ,ω o )=∫ Ω L i (x,ω i )f r (x,ω i ,ω o )(ω i ·n)dω i

[0095] Where x represents a surface point, ω i represents the incident light direction, ω o Indicates the viewpoint direction, L i (x,ω i ) represents global illumination, f r (x,ω i ,ω o ) represents the bidirectional reflectance distribution function BRDF value, and n represents the normal direction.

[0096] The step 6 is specifically described as follows:

[0097] The loss function in the joint optimization phase includes the loss function L in step 4 geo , image reconstruction loss L rec and smoothing loss L smooth . It is defined as follows:

[0098] L total =L geo +0.1L rec +0.1L smooth

[0099] Image reconstruction loss L rec The definition is as follows:

[0100]

[0101] Where p represents the number of surface points, m represents the total number of sampled pixels, Represents the color value based on physical rendering, C p represents the true color value, and R represents the MAE loss.

[0102] Smoothing loss L smooth The definition is as follows:

[0103]

[0104] Among them, S represents the number of sampled pixels, x p Represents the number of surface points. The gradient of the input image Obtained by precalculation. Roughness gradient and the specular albedo gradient Calculated during the back propagation process. The loss function L total Jointly optimize the signed distance function S and the neural radiation field L in step 3 o As well as the visibility network Φ and material network M in step 5.

[0105] It should be noted that the above-mentioned pre-calculation refers to calculating the gradient value of the image pixels while reading the image during the data loading process, wherein reading the image and calculating the gradient value are both implemented by directly calling the corresponding functions in the OpenCV library.

[0106] The implementation basis of each embodiment of the present invention is to implement programmed processing through a device with a processor function. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this reality, on the basis of the above embodiments, an embodiment of the present invention provides a system for estimating human body geometry, lighting and material from multi-view images, which is used to execute a method for estimating human body geometry, lighting and material from multi-view images in the above method embodiment.

[0107] The system includes: a data acquisition module, which is used to acquire a 3D human body model to be estimated and perform multi-view rendering; a data estimation module, which is used to input the rendered data into a trained human body geometry, illumination and material estimation network, and output the estimated human body geometry, illumination and material; wherein the training of the human body geometry, illumination and material estimation network includes: collecting 3D human body model data, rendering a multi-view human body image and corresponding camera pose parameters according to the collected 3D human body model data; performing geometric pre-training: using the rendered data as input, implicitly representing the human body surface using a signed distance function, obtaining surface points using a spherical tracking algorithm, and obtaining the estimated human body geometry, illumination and material through a neural radiation field meter. Calculate the appearance color value corresponding to the surface point; select the loss function of geometric pre-training, and gradually optimize the signed distance function and neural radiation field by minimizing the loss function; perform joint optimization: use the pre-trained signed distance function and spherical tracking algorithm to obtain the surface points of the human body, input them into the designed visibility network and material network, and output the lighting visibility and material; use the lighting visibility and pre-trained neural radiation field to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function; calculate the color value of the surface point through the rendering equation; select the loss function for joint optimization, and jointly optimize the signed distance function, neural radiation field, visibility network and material network by minimizing the loss function.

[0108] An embodiment of the present invention provides a system for estimating human body geometry, lighting, and material from multi-view images. In response to the challenges of human body inverse rendering, the system adopts the above-mentioned several modules, pre-trains the human body geometry in the geometry pre-training stage to obtain a good initialization result, and then estimates the lighting and material in the joint optimization stage and fine-tunes the geometry initialization result, thereby solving problems such as inaccurate human body material estimation, reducing the cost of modeling indirect lighting and lighting visibility, and achieving a more realistic inverse rendering effect.

[0109] It should be noted that the system embodiments provided by the present invention are not only used to implement the methods in the above-mentioned method embodiments, but also used to implement the methods in other method embodiments provided by the present invention. The only difference lies in the setting of corresponding functional modules, and the principles thereof are basically the same as the principles of the above-mentioned system embodiments provided by the present invention. As long as technical personnel in this field refer to the specific technical solutions in other method embodiments on the basis of the above-mentioned system embodiments, obtain corresponding technical means and technical solutions composed of these technical means by combining technical features, and on the premise of ensuring the practicality of the technical solutions, they will improve the modules in the above-mentioned system embodiments to obtain corresponding system class embodiments for implementing the methods in other method class embodiments.

[0110] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention also provides a device for estimating human body geometry, lighting and material from multi-view images, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method for estimating human body geometry, lighting and material from multi-view images.

[0111] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention further provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for estimating human body geometry, lighting and material from multi-view images.

[0112] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0113] In summary of the above embodiments, the present invention designs a method and system for estimating human body geometry, lighting and material from multi-view images captured under unknown lighting based on the idea of ​​joint optimization of neural radiation field rendering and inverse rendering. The present invention is mainly divided into a geometric pre-training stage and a joint optimization stage. The geometric pre-training stage first pre-trains the human body geometry to obtain a good initialization result, and then estimates the lighting and material in the joint optimization stage and fine-tunes the geometric initialization result. When estimating the material, the present invention uses an anisotropic material model to model complex human body materials. In addition, when estimating the lighting, the present invention does not use an additional training stage to model indirect lighting, but instead regards the outgoing radiance of the indirect reflection point as indirect lighting, and estimates the lighting visibility to distinguish between direct lighting and indirect lighting. The present invention solves the problem of inaccurate estimation of complex human body materials, and can effectively decouple geometry, lighting and material from multi-view images of the human body, thereby improving the realism of the human body re-lighting results.

[0114] The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for estimating human geometry, lighting and material from multi-view images, characterized in that: include: Obtaining a three-dimensional human body model to be estimated and performing multi-view rendering; The rendered data is input into a trained human body geometry, lighting and material estimation network, and the estimated human body geometry, lighting and material are output; wherein the training of the human body geometry, lighting and material estimation network includes: Collecting 3D human body model data, and rendering multi-view human body images and corresponding camera pose parameters based on the collected 3D human body model data; Perform geometric pre-training: Use the rendered data as input, use the signed distance function to implicitly represent the human body surface, use the spherical tracking algorithm to obtain surface points, and calculate the appearance color value corresponding to the surface point through the neural radiation field; select the loss function of geometric pre-training, and gradually optimize the signed distance function and neural radiation field by minimizing the loss function; Perform joint optimization: use the pre-trained signed distance function and spherical tracking algorithm to obtain the surface points of the human body, input them into the designed visibility network and material network, and output the lighting visibility and material; use the lighting visibility and pre-trained neural radiation field to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function; calculate the color value of the surface point through the rendering equation; select the loss function for joint optimization, and jointly optimize the signed distance function, neural radiation field, visibility network, and material network by minimizing the loss function.

2. The method for estimating human body geometry, lighting and material from multi-view images according to claim 1, characterized in that: Perform geometry pre-training, also includes: The human body surface is implicitly represented as a set of zero values ​​of the signed distance function; According to each input image and its corresponding camera pose, the pixel ray corresponding to each pixel point is obtained; Sampling is performed on the pixel ray through the spherical tracing algorithm, the SDF value of each sampling point is calculated, and when the SDF value of the sampling point is within a preset threshold close to 0, the sampling point is regarded as the surface point corresponding to the pixel ray; After obtaining the surface point coordinates, they are input into the designed neural radiation field to calculate the appearance color value.

3. The method for estimating human body geometry, lighting and material from multi-view images according to claim 2, characterized in that: The signed distance function and the neural radiation field are both represented by multi-layer perceptrons. During geometric pre-training, the weight parameters of the two multi-layer perceptrons are trained by narrowing the gap between the predicted values ​​and the true values, and the masked true values ​​are used for supervision.

4. The method of estimating human body geometry, lighting and material from multi-view images according to claim 1, characterized in that: Joint optimization also includes: When estimating illumination, the global illumination is decomposed into direct illumination and indirect illumination, the direct illumination is fitted using a spherical Gaussian function, the indirect illumination is calculated using the outgoing irradiance of the indirect bounce point, and visibility is used to indicate the visibility of the direct illumination to the surface points in the illumination direction; When estimating materials, an anisotropic material model is used to calculate the bidirectional reflectance distribution function BRDF, and the rendering equation is used to calculate the color value based on physical rendering.

5. The method for estimating human body geometry, lighting and material from multi-view images according to claim 4, characterized in that: The loss function of the joint optimization includes geometric pre-training loss, image reconstruction loss and smoothing loss.

6. A system for estimating human geometry, lighting, and material from multi-view images, characterized in that: include: A data acquisition module, used to acquire a three-dimensional human body model to be estimated and perform multi-view rendering; A data estimation module is used to input the rendered data into a trained human body geometry, lighting and material estimation network, and output estimated human body geometry, lighting and material; wherein the training of the human body geometry, lighting and material estimation network includes: Collecting 3D human body model data, and rendering multi-view human body images and corresponding camera pose parameters based on the collected 3D human body model data; Perform geometric pre-training: Use the rendered data as input, use the signed distance function to implicitly represent the human body surface, use the spherical tracking algorithm to obtain surface points, and calculate the appearance color value corresponding to the surface point through the neural radiation field; select the loss function of geometric pre-training, and gradually optimize the signed distance function and neural radiation field by minimizing the loss function; Perform joint optimization: use the pre-trained signed distance function and spherical tracking algorithm to obtain the surface points of the human body, input them into the designed visibility network and material network, and output the lighting visibility and material; use the lighting visibility and pre-trained neural radiation field to calculate the global illumination, and use the material to calculate the bidirectional reflectance distribution function; calculate the color value of the surface point through the rendering equation; select the loss function for joint optimization, and jointly optimize the signed distance function, neural radiation field, visibility network, and material network by minimizing the loss function.

7. A device for estimating human geometry, lighting and material from multi-view images, characterized in that include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method for estimating human body geometry, lighting, and material from multi-view images as described in any one of claims 1 to 6.

8. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for estimating human body geometry, lighting and material from multi-view images according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent three-dimensional reconstruction method, device and equipment based on neural primary function, and medium

    CN114494611A

  • Three-dimensional scene geometry, material and illumination decoupling and editing system

    CN116977431A

  • Realistic real-time rendering method for intelligent workshop three-dimensional scene

    CN117351130A

  • Spatial variation indoor scene illumination estimation method based on neural radiation field

    CN117671126A

  • Neural rendering method based on multi-resolution network structure

    WO2023225891A1

Cited By

  • Method and system for dynamically rendering textures and materials of virtual clothes

    CN121527221A