3D Scene Inverse Rendering and Reconstruction Method Using Near-Field and Far-Field Light Sources

By using the inverse rendering method of near-field and long-range light sources in three-dimensional reconstruction technology, the rendering problem in different lighting environments is solved, the fineness and reconstruction speed of material estimation are improved, the utilization of near-field light sources is supported, and the shooting process is simplified.

CN118447147BActive Publication Date: 2025-06-13PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410556421.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-06-13
Estimated Expiration
2044-05-07

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology is difficult to accurately render three-dimensional objects under different lighting environments, and it is unable to effectively utilize near-field light sources, resulting in poorly estimated material and long reconstruction time.

Method used

The three-dimensional scene inverse rendering reconstruction method using near-field and long-distance light sources is adopted to input images through multiple perspectives, and the scene shape, material and radiation field under each light source conditions are respectively modeled and optimized, and the reconstruction process is accelerated by multi-resolution hash grid position coding.

Benefits of technology

It realizes accurate rendering of three-dimensional objects under any lighting conditions, improves the fineness and reconstruction speed of material estimation, supports the situation where near-field light sources are included in the input image, and simplifies the shooting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447147B_ABST
    Figure CN118447147B_ABST
Patent Text Reader

Abstract

A three-dimensional scene inverse rendering and reconstruction method using near-field and far-field light sources belongs to the field of three-dimensional reconstruction technology, and includes: obtaining appearance images of the same object under different viewpoints and different light source conditions as input images, and annotating the input images to obtain camera poses; modeling the scene shape, material, radiation field and light source; reconstructing the scene shape field and radiation field based on volume rendering; reconstructing the scene material field and light source based on surface rendering; constructing and exporting a three-dimensional model with materials. The present invention uses all far-field and near-field light source conditions including ambient light sources contained in the input images to perform disambiguation between the light source and the material, can more effectively utilize controllable light sources and different ambient light source conditions, and obtain a more refined and accurate object material reconstruction; at the same time, uses a more advanced multi-resolution hash grid position encoding as the representation method of the neural field, and performs a faster three-dimensional inverse rendering and reconstruction of the scene without reducing the quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional reconstruction, and particularly relates to a method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-distance light sources. Background Art

[0002] In today's digital age, the rapid development of multimedia technology has provided people with an unprecedented immersive experience. Especially in modern industries such as film production, game development, virtual reality (VR), and augmented reality (AR), high-quality three-dimensional scene creation and object rendering have become key technologies to enhance realism, immersion, and user experience. With the continuous popularization of these industries, the demand for filling virtual three-dimensional objects with realism and diversity in virtual world scenes is increasing. However, traditional three-dimensional object creation pipelines often involve multiple complex steps, including designing models, designing texture maps, retopologizing, baking maps, designing materials, etc. These steps are not only time-consuming but also require professional artists and software support. This cumbersome workflow greatly limits the creation efficiency of three-dimensional content, especially in application scenarios that require rapid generation of a large number of high-quality three-dimensional models, such as large-scale game environment design, rapid prototyping of movie scenes, etc.

[0003] Three-dimensional reconstruction technology can solve this problem. By using multi-view two-dimensional RGB images as input, three-dimensional reconstruction technology can create three-dimensional objects with less time and effort. The application value of this technology lies in its ability to greatly simplify the generation process of three-dimensional content, enabling people to quickly create and utilize three-dimensional models without professional three-dimensional modeling knowledge. Among many three-dimensional reconstruction technologies, Neural Radiance Fields (NeRF), as an emerging technology, has received extensive attention due to its ability to produce high-quality scene reconstruction effects. NeRF reconstructs a continuous and high-fidelity three-dimensional scene from a set of sparse views by modeling the radiation field of the scene (i.e., the result of the interaction between incident light and the scene surface and then being reflected). Although NeRF has achieved remarkable achievements in visual effects, since it models the scene under specific lighting conditions, it cannot meet the needs of the film or game industry to place three-dimensional objects in different lighting environments and render their variable appearances. What these industries urgently need is an inverse rendering three-dimensional reconstruction method that can separate environmental lighting from object material properties during the reconstruction process. The three-dimensional object with material properties obtained by this method can be accurately rendered under any lighting conditions and is more suitable for various application scenarios in the industry.

[0004] One of the core challenges of inverse rendering technology is to address the inherent ambiguity caused by limited viewing angles and lighting conditions. For example, a pixel that appears dark in a captured image could be because it is inherently dark in color, the lighting is dim, or the angle between the object's surface and the incident light is large. This ambiguity makes it extremely difficult to accurately distinguish the true shape and material of an object from limited observations. To address this challenge, a method can be adopted where the object is observed and photographed from multiple angles under different lighting conditions, and then three-dimensional reconstruction is performed through inverse rendering technology. By collecting data under various lighting conditions and leveraging the information increment brought about by lighting changes, this method helps to resolve the inherent ambiguity, enabling the three-dimensional model reconstructed from images taken from multiple angles and under multiple lighting conditions to not only be highly accurate but also separate the material properties of the object and the environmental lighting effects, achieving a more realistic and flexible three-dimensional scene rendering.

[0005] The literature "Ziang Cheng, Junxuan Li, and Hongdong Li. WildLight: In-the-wild inverse rendering with a flashlight. In Proc. of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023." discloses a method for three-dimensional reconstruction. As Figure 1 shown, first, multi-view images of an object are captured under a fixed environmental lighting. When capturing, an additional flashlight fixed to the camera is used. The flashlight is turned off when capturing a part of the images and turned on when capturing another part. Subsequently, the shape of the object and the neural radiance field when the flashlight is off (i.e., only under environmental lighting) are directly reconstructed using existing techniques. Using this reconstructed radiance field, the appearance under environmental lighting for this perspective can be subtracted from the pixel values of the images when the flashlight is on to obtain the appearance of the object with only the flashlight and no environmental lighting. Thereafter, based on the reflection model under the flashlight, the material of the object can be solved for. In addition to requiring a near-field flashlight, the method disclosed in the above literature requires the object to be captured under a fixed and non-purely dark environmental lighting to reconstruct the object shape, but it cannot estimate the object material using the appearance of the object captured under this environmental lighting, so a fine and accurate object material cannot be obtained. At the same time, the above literature adopts an inefficient neural field representation method, resulting in a very long reconstruction time.

[0006] The literature "Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Songfang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. TensoIR: Tensorial inverse rendering. In Proc. of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023." discloses a 3D reconstruction method: taking multi-view images of an object as input, this method adopts a scene representation based on tensor decomposition and a conductible renderer based on volume density to perform 3D inverse rendering reconstruction on the scene to obtain the shape, material, and light source of the scene. This method reconstructs the scene radiance field under each set of distant light sources and optimizes a shared shape and material field, so as to be able to utilize different distant light source conditions in the input images to improve the accuracy of estimation. Although the method disclosed in the above literature can disambiguate the light source and material using the appearance of the object under multiple distant environmental light sources, it cannot consider the situation where there are near-field light sources (such as flashlights) in the input, thus unable to make more effective use of controllable light sources. Summary of the Invention

[0007] The purpose of the present invention is to provide a 3D scene inverse rendering reconstruction method using near-field and distant light sources to solve the problems existing in the prior art.

[0008] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0009] The 3D scene inverse rendering reconstruction method using near-field and distant light sources of the present invention mainly includes:

[0010] Step S1: Obtain input images, where the input images are appearance images of the same object under different viewpoints and different light source conditions, and label the input images to obtain the camera poses;

[0011] Step S2: Model the scene shape, material, radiance field, and light source;

[0012] Step S3: Reconstruct the scene shape field and radiance field based on volume rendering;

[0013] Step S4: Reconstruct the scene material field and light source based on surface rendering;

[0014] Step S5: Construct and export a 3D model with materials.

[0015] Further, in step S1, assuming that each input image is illuminated by a certain far-field light source and several near-field light sources, the number N of far-field light sources that appear in all input images needs to be given in advance. far and the number N of near-field light sources near , and the far-field light source numbers corresponding to each input image and the type and switch state of each near-field light source when shooting the input image are marked.

[0016] Further, in step S1, the type of each near-field light source when shooting the input image is stationary and fixed to the camera.

[0017] Further, in step S2, each far-field light source i ∈ {1, 2, …, N far} is modeled as a spherical Gaussian function with 128 lobes, and each lobe j contains 6 parameters: axial lobe sharpness and lobe RGB amplitude Without considering occlusion, the incident irradiance of the far-field light source i from the direction shooting to any position in the scene is:

[0018]

[0019] where is the normalization coefficient.

[0020] Further, in step S2, each near-field light source i ∈ {1, 2, …, N near} is modeled as a point light source with variable position p i and anisotropic radiation; for the near-field light source of the stationary type, its position is initially unknown and is regarded as an optimizable parameter; for the near-field light source fixed to the camera, its position on each input image is set to be the same as the camera position; the anisotropic radiation is parameterized as the coefficients of the spherical harmonic function of order l The incident irradiance of the near-field light source i shooting to the three-dimensional position x in the scene is:

[0021]

[0022] where the incident direction of the near-field light source SH(ω; h i ) is to calculate the spherical harmonic function value with coefficient h i and direction ω.

[0023] Furthermore, in step S2, the scene shape, material, and radiation field are all represented as neural fields. For any three-dimensional position x, after performing position encoding on it using a multi-resolution hash grid respectively, it is input into a multi-layer perceptron to obtain the shape, material parameters, and neural features at the corresponding position. The viewing direction is encoded using spherical harmonics, and the encoded result and the light source embedding encoding are input into the radiation field multi-layer perceptron together to obtain the corresponding radiance.

[0024] Furthermore, in step S3, a random input image is selected, and a camera ray in the field of view of the input image is sampled, and several segmentation points are sampled on the camera ray. The SDF values obtained by querying the scene shape field at all the sampled points are converted into volume densities, and the radiance obtained by querying the scene radiation field is accumulated on the camera ray according to the volume density, and summed according to the light source situation to obtain the ray color. The L2 distance between the ray color and the color of the corresponding pixel in the input image, together with other regularization terms, is used as the loss function for iterative optimization, and the scene shape field and radiation field are reconstructed through reconstruction.

[0025] Furthermore, in step S4, a random input image is selected, and a camera ray in the field of view of the input image is sampled. The position and surface normal of the surface point where the camera ray first touches are obtained by querying the scene shape field, and the material parameters of the surface point where the camera ray first touches are obtained by querying the scene material field. When calculating the ray color under the direct illumination of a distant light source, multiple importance sampling is performed on the incident direction, and Monte Carlo integration is used for estimation. When calculating the ray color under the direct illumination of a near-field light source, the incident direction is directly taken as the near-field light source direction, and the reflection is calculated. The ray colors under each near-field light source and distant light source are summed according to the light source conditions in the selected input image to obtain the predicted ray color. The L2 distance between it and the color of the corresponding pixel, together with other regularization terms, is used as the loss function for iterative optimization, and the scene material and light source are reconstructed through reconstruction.

[0026] Furthermore, in step S5, the zero-value surface of the SDF field is extracted as a triangular mesh through an isosurface extraction algorithm and simplified. Then, the UV mapping technology is used to obtain the UV space coordinates of each vertex of the triangular mesh, and the triangular mesh is rasterized into a two-dimensional UV texture according to this coordinate. For each pixel of the UV texture, the corresponding three-dimensional coordinates are calculated, and the material map and normal map are obtained by querying the reconstructed material and shape fields, so as to obtain a three-dimensional model with materials.

[0027] The beneficial effects of the present invention are:

[0028] A 3D scene inverse rendering and reconstruction method using near-field and far-field light sources according to the present invention estimates the shape and material of an object using various near-field and far-field light source conditions of multi-view input images, including separately modeling each light source condition, and steps of reconstructing the same shape field and material field using the appearance under each light source. Compared with the prior art, the present invention improves the technical effect in the following aspects:

[0029] 1. The present invention uses all far-field and near-field light source conditions including ambient light sources contained in the input images to disambiguate between the light sources and the materials, so that the controllable light sources (such as flashlights) and different ambient light source conditions can be utilized more effectively, and a more refined and accurate object material reconstruction can be obtained.

[0030] 2. The present invention uses a more advanced multi-resolution hash grid position encoding as the representation of the neural field, and can perform a faster 3D inverse rendering and reconstruction of the scene without reducing the quality.

[0031] 3. The present invention supports near-field light sources in the input images, and there is no need to move the object to a new scene to obtain images under different light sources, making the shooting process more convenient. Description of the Drawings

[0032] Figure 1 It is a system block diagram in the literature "Ziang Cheng, Junxuan Li, and Hongdong Li. WildLight: In-the-wild inverse rendering with a flashlight. In Proc. of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023."

[0033] Figure 2 It is a system block diagram in the literature "Haian Jin, Isabella Liu, Peijia Xu, Xiaoshuai Zhang, Songfang Han, Sai Bi, Xiaowei Zhou, Zexiang Xu, and Hao Su. TensoIR: Tensorial inverse rendering. In Proc. of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023."

[0034] Figure 3Flowchart of a three-dimensional scene inverse rendering and reconstruction method using near-field and far-field light sources according to the present invention. Detailed implementation manners

[0035] The present invention will be further described in detail below with reference to the accompanying drawings.

[0036] A three-dimensional scene inverse rendering and reconstruction method using near-field and far-field light sources according to the present invention specifically includes the following steps: input image annotation; scene shape, material, radiation field, and light source reconstruction; scene shape field and radiation field reconstruction based on volume rendering; scene material field and light source reconstruction based on surface rendering; constructing and exporting a three-dimensional model with materials.

[0037] A three-dimensional scene inverse rendering and reconstruction method using near-field and far-field light sources according to the present invention, its specific operation process is as Figure 3 shown, specifically:

[0038] Step S1, input image annotation;

[0039] As Figure 3 shown in the lower left subfigure, the input of the present invention is multi-view input images under variable lighting conditions, that is, the appearances (i.e., RGB images) of the same object under different viewpoints and different light source conditions, as well as the camera poses calibrated for each input image. Assuming that each input image is illuminated by a certain far-field light source and several near-field light sources, the number N far of far-field light sources and the number N near of near-field light sources that appear in all input images need to be given in advance, and the far-field light source numbers corresponding to each input image and the type and on / off state of each near-field light source when shooting the input image are annotated. Among them, the types of near-field light sources when shooting the input image are mainly two types: stationary and fixed to the camera.

[0040] As a preferred embodiment, when taking images, the object can be photographed under different perspectives and different light source conditions, and the method of the present invention is used to reconstruct the rendered image of the object under the new light source. The photographing device is preferably an iPhone 12 Pro. The original image is exported using the "ProCam" application on it, and it is converted into a linear RGB image using a fixed camera image processing process. During shooting, the camera white balance, focal length, exposure time, and ISO are all fixed. When using the flash, the camera distance is about 0.5 m. The camera takes about 100 images around the object, with the flash on in half of the images and off in the other half. For the convenience of subsequent camera pose calibration, the object can be pasted on a wooden board covered with ARTags (when the object itself has rich features for calibration and the camera pose calibration will not fail, it can be not used). Then, the default parameters of the ready-made COLMAP method are used for image feature extraction and matching to obtain the camera pose of each image. When using ARTags, the corner points extracted by ARTags need to be used as matching point pairs to replace the point pairs obtained by the COLMAP method itself. In addition, the total number of light sources needs to be given, and the light source conditions (distant light source number, near-field light source on / off state) when each input image is taken need to be marked. The present invention can obtain more accurate object material reconstruction by using the different light source conditions included in the input images.

[0041] Step S2, scene shape, material, radiation field, and light source reconstruction;

[0042] As Figure 3 shown in the upper left subfigure, the present invention models the scene light source as a parametric light source model. Each distant light source i ∈ {1, 2, …, N far} is modeled as a spherical Gaussian function with 128 lobes, and each lobe j contains 6 parameters: axial lobe sharpness and lobe RGB amplitude Therefore, without considering occlusion, the incident radiance of the distant light source i from the direction shooting to any position in the scene is:

[0043]

[0044] where is the normalization coefficient.

[0045] Each near-field light source i ∈ {1, 2, …, N near} is modeled as a position p iA point light source with variable and anisotropic radiation. For a near-field light source of the fixed type, its position is initially unknown and is regarded as an optimizable parameter; for a near-field light source fixed to the camera, its position on each input image is set to be the same as the camera position. The anisotropic radiation is parameterized as the coefficients of the spherical harmonic functions of order l Thus, the incident radiance of the near-field light source i towards the three-dimensional position x in the scene is:

[0046]

[0047] where the incident direction of the near-field light source SH(ω; h i ) is to calculate the value of the spherical harmonic function at the direction ω with coefficients h i .

[0048] As shown in the lower subfigure of Figure 3 , the present invention represents the scene shape, material, and radiation field as neural fields. For any three-dimensional position x, after position encoding it using multi-resolution hash grids (enc geo and enc mat ), they are respectively input into multi-layer perceptrons (M geo and M mat ) to obtain the shape (SDF value s at the corresponding position, and its derivative with respect to the position x can obtain the surface normal n), material parameters β (parameters for parameterizing the GGX BRDF model), and neural features z. Since the radiance is related not only to the position but also to the viewing angle and the light source, it is necessary to additionally perform spherical harmonic encoding on the viewing direction v (enc dir ). The encoding result enc dir (v) and the light source embedding encoding ( and ) are jointly input into the radiation field multi-layer perceptrons (M far and M near ) to obtain the corresponding radiance That is and The present invention performs a faster three-dimensional inverse rendering reconstruction of the scene by using multi-resolution hash grid position encoding.

[0049] As a preferred embodiment, the parameterized light source model used in the light source model can be replaced with other similar models. For example, the spherical Gaussian function model for the far-field light source can be replaced with a spherical harmonic function model or replaced with an environment light map.

[0050] As a preferred embodiment, the parametric model used in the material model can be replaced with other similar models. For example, the parametric GGX BRDF model used can be replaced with the DisneyPrincipled BRDF model.

[0051] Step S3: Reconstruct the scene shape field and radiation field based on volume rendering;

[0052] Step S3 is used to minimize the difference between the volume-rendered image and the input image and reconstruct the scene shape field and radiation field. Specifically: First, randomly select an input image, sample a camera ray r(t) = o + td in the field of view of the input image (where o is the origin of the camera ray, d is the direction of the camera ray, and t > 0), and sample several segmentation points on the camera ray. Then, convert the SDF values obtained by querying the scene shape field at all sampling points into volume density α i , and based on the volume density α i accumulate the radiance obtained by querying the scene radiation field ( and ) along the camera ray, sum according to the light source situation, and obtain the ray color C rf (o, d). Take the L2 distance between this ray color C rf (o, d) and the color of the corresponding pixel in the input image, together with other regularization terms (including normal smoothness, etc.), as the loss function, and perform iterative optimization to reconstruct the scene shape field and radiation field.

[0053] Step S4: Reconstruct the scene material field and light source based on surface rendering;

[0054] Step S4 is mainly used to minimize the difference between the image obtained through Step S3 and the input image while fixing the scene shape field and radiation field obtained in the previous step, and reconstruct the scene material field and ambient light source. Specifically: Randomly select an input image, sample a camera ray r(t) = o + td in the field of view of the input image, obtain the surface point position x and surface normal n where the camera ray first touches through querying the scene shape field, and obtain the material parameter β of the surface point where the camera ray first touches through querying the scene material field. To calculate the ray color under the direct illumination of the distant light source i, perform multiple importance sampling on the incident direction ω s and use Monte Carlo integration, as follows:

[0055]

[0056] where V(x, ω s ) represents whether the ray emitted from the surface point position x in the incident direction ω s is blocked by the scene shape itself; Represents the line-of-sight color under the direct illumination of the distant light source i, L far (ω s ) represents the incident radiance of the distant light source i from the direction ω s , p(ω s ) represents the probability density function of sampling the incident direction ω s , f(x, ω s , -d; β) represents the ratio of the surface point x reflecting light from the incident direction ω s to the outgoing direction when the material parameter is β.

[0057] To calculate the line-of-sight color under the direct illumination of the near-field light source Since the light from the near-field light source only comes from a single point rather than the entire spherical surface, the incident direction is directly taken as the near-field light source direction, and the reflection can be calculated. The calculation of the line-of-sight color under indirect illumination is similar, just replace L far (ω s )V(x, ω s ) with c(x', ω s )(1 - V(x, ω s )) where x' is the intersection point of the light emitted from the surface point position x in the incident direction ω s with the scene itself, and c(x', ω s ) represents the outgoing radiance reflected from the point x' in the direction ω s .

[0058] By summing up the line-of-sight colors under the near-field light source and the distant light source according to the light source conditions in the selected input image, the predicted line-of-sight color can be obtained. The L2 distance between it and the color of the corresponding pixel, together with other regularization terms (including material smoothness, consistency between volume rendering and surface rendering colors, etc.), is used as the loss function for iterative optimization to reconstruct the scene material and light source.

[0059] Figure 3 In represents the line-of-sight color considering both direct and indirect illumination under the distant light source i based on surface rendering, represents the line-of-sight color considering both direct and indirect illumination under the near-field light source i based on surface rendering, C gt represents the color of the corresponding pixel of the line-of-sight in the input image, C pb represents the line-of-sight color obtained by summing according to the light source situation of the input image based on surface rendering, represents the line-of-sight color under indirect illumination under the near-field or distant light source based on surface rendering, represents the line-of-sight color under the direct illumination of the near-field light source i based on surface rendering, represents L far(ω)V(X, ω), represents c(X', ω)(1 - V(X, ω)), represents L near (ω)V(x, ω).

[0060] Step S5, construct and export a 3D model with materials;

[0061] After reconstructing the scene shape and material field through the above steps, the zero-value surface of the SDF field is extracted as a triangular mesh by the Marching Cubes algorithm and simplified. Then, the UV mapping technology is used to obtain the UV space coordinates of each vertex of the triangular mesh, and the triangular mesh is rasterized into a 2D UV texture according to these coordinates. For each pixel of the UV texture, after calculating the corresponding 3D coordinates and querying the reconstructed material and shape fields, the material map and normal map can be obtained, thus obtaining a 3D model with materials.

[0062] As a preferred embodiment, it can be run on a computer equipped with an NVIDIA graphics card above the 30 series with a video memory of more than 12G, taking the image processed in step S2, the camera pose, and the input image annotation as inputs, and using the three-dimensional scene inverse rendering and reconstruction method of the present invention using near-field and far-field light sources to reconstruct the scene to obtain a 3D model.

[0063] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A three-dimensional scene inverse rendering reconstruction method using near-field and far-field light sources, characterized in that: include: Step S1, obtaining an input image, wherein the input image is an appearance image of the same object under different viewing angles and different light sources, and annotating the input image to obtain a camera pose; Step S2, modeling scene shape, scene material, radiation field and light source; The light source includes a near-field light source and a long-distance light source; each long-distance light source i∈{1,2,…,N far } is modeled as a spherical Gaussian function with 128 lobes, each lobe j contains 6 parameters: axial Lobe sharpness and lobe RGB amplitude Without considering occlusion, the distant light source i is The incident radiance at any location in the scene is: in is the normalization coefficient; Step S3, reconstructing the scene shape field and radiation field based on volume rendering; Step S4, scene material field and light source reconstruction based on surface rendering; Step S5: construct and export a three-dimensional model with materials.

2. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and long-distance light sources according to claim 1, characterized in that: In step S1, assuming that each input image is illuminated by a distant light source and several near-field light sources, the number of distant light sources N that appear in all input images must be given in advance. far and the number of near-field light sources N near , and mark the long-distance light source number corresponding to each input image and the type and switch status of each near-field light source when shooting the input image.

3. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and long-distance light sources according to claim 2, characterized in that: In step S1, each near-field light source is of a type that is stationary or fixed to a camera when capturing an input image.

4. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-field light sources according to claim 1, characterized in that: In step S2, each near-field light source k∈{1,2,…,N near } is modeled as a position p k Point sources with variable, anisotropic radiation; for near-field sources of type stationary, their position is initially unknown and is treated as an optimizable parameter; for near-field sources of type fixed to the camera, their position on each input image is set to be the same as the camera position; anisotropic radiation is parameterized as the coefficients of spherical harmonics of order l The incident radiance of the near-field light source k at the three-dimensional position x in the scene is: The incident direction of the near-field light source SH(ω;h k ) is the calculation coefficient h k , the spherical harmonic function values ​​at the direction ω.

5. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-field light sources according to claim 1, characterized in that: In step S2, the scene shape, material and radiation field are all represented as neural fields; for any three-dimensional position x, it is position-encoded using a multi-resolution hash grid and then input into a multi-layer perceptron to obtain the shape, material parameters and neural features of the corresponding position; The viewing direction is encoded by spherical harmonics, and the encoding result and the light source embedding code are input into the radiation field multi-layer perceptron to obtain the corresponding radiance.

6. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-field light sources according to claim 1, characterized in that: In step S3, an input image is randomly selected, a camera line of sight in the field of view of the input image is sampled, and several segmentation points are sampled on the camera line of sight; the SDF values ​​obtained by querying the scene shape field at all sampling points are converted into volume density, and the radiance obtained by querying the scene radiation field is accumulated on the camera line of sight according to the volume density, and summed according to the light source situation to obtain the line of sight color; the L2 distance between the line of sight color and the color of the corresponding pixel in the input image is used as the loss function together with other regularization terms, and iterative optimization is performed to reconstruct the scene shape field and radiation field.

7. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-field light sources according to claim 1, characterized in that: In step S4, an input image is randomly selected, and a camera line of sight in the field of view of the input image is sampled. The position and surface normal of the surface point that the camera line of sight contacts for the first time are obtained by querying the scene shape field, and the material parameters of the surface point that the camera line of sight contacts for the first time are obtained by querying the scene material field; when calculating the line of sight color under direct illumination of a distant light source, multiple importance sampling is performed on the incident direction, and Monte Carlo integration is used for estimation; when calculating the line of sight color under direct illumination of a near-field light source, the incident direction is directly taken as the direction of the near-field light source, and the reflection is calculated; according to the light source conditions in the selected input image, the line of sight colors under each near-field light source and the distant light source are summed to obtain the predicted line of sight color, and the L2 distance between it and the color of the corresponding pixel, together with other regularization terms, is used as a loss function, and iterative optimization is performed to obtain the scene material and light source after reconstruction.

8. The method for inverse rendering and reconstruction of a three-dimensional scene using near-field and far-field light sources according to claim 1, characterized in that: In step S5, the zero-value surface of the SDF field is extracted as a triangular mesh by an isosurface extraction algorithm and simplified; then the UV space coordinates of each triangular mesh vertex are obtained by UV mapping technology, and the triangular mesh is gridded into a two-dimensional UV texture according to the coordinates; for each pixel of the UV texture, the three-dimensional coordinates corresponding to the pixel are calculated, and then the reconstructed material and shape field are queried to obtain the material map and normal map, thereby obtaining a three-dimensional model with material.

Citation Information

Patent Citations

  • Virtual object rendering method, electronic equipment and storage medium

    CN116797708A

  • Object inverse rendering method and device, equipment and storage medium

    CN117710569A