Nerve radiation field relighting method based on progressive mesh refinement

By introducing progressive grid refinement and frequency encoding technology into neural radiation field technology, combined with residual network structure, the problem of fuzzy detail reconstruction and high-frequency perceptual ambiguity in re-lighting tasks is solved, and the image quality and the accuracy of scene reconstruction are significantly improved.

CN119963715AActive Publication Date: 2025-05-09NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510033574.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing neural radiation field technology has problems with detail reconstruction blur and high-frequency perception ambiguity in re-lighting tasks, resulting in uneven light distribution or loss of details.

Method used

Using the neural radiation field re-lighting method based on progressive grid refinement, the three-dimensional space features are extracted and refined layer by layer by layer by layer by constructing progressive multi-resolution feature grids and frequency coding technology, and combining the residual network structure, the flexibility and efficiency of the network in processing high-frequency information are optimized.

Benefits of technology

It significantly improves the ability of neural radiation fields to reconstruct high-frequency details, generates smoother and clearer rendered images, improves the quality of scene reconstruction, and solves the problems of blurred detail reconstruction and high-frequency perception ambiguity under lighting changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963715A_ABST
    Figure CN119963715A_ABST
Patent Text Reader

Abstract

The invention provides a neural radiation field relighting method based on progressive mesh refinement, and relates to the technical field of computer graphics and new view synthesis, the method adopts a mesh structure with gradually progressive resolution, combines a frequency coding technology, and performs feature extraction on each layer of mesh through a multilayer perceptron, so as to improve the resolution of the mesh. And refining the three-dimensional structure and illumination information of the scene layer by layer. Specifically, a low-resolution grid is combined with low-frequency coding, a high-resolution grid is combined with high-frequency coding, the feature expression ability under different scales is effectively enhanced through residual connection, the re-lighting effect of a scene is optimized, and the details and definition of a rendered image are improved. According to the method, the training process is efficient, rapid training and reasoning can be achieved on the consumer-level graphics card, and the method has good expandability, is widely suitable for various environment perception tasks such as virtual reality, augmented reality and automatic driving and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer graphics and new perspective synthesis, and in particular to a neural radiation field relighting method based on progressive grid refinement. Background Art

[0002] In recent years, with the rapid development of computer vision and graphics technology, the fields of 3D reconstruction and image rendering have ushered in important innovations and optimizations. Driven by the needs of emerging applications, especially in the fields of virtual reality, augmented reality, and film and television production, the demand for high-quality 3D scene reconstruction and fine rendering is increasing. The birth of neural radiance field technology marks a breakthrough in deep learning-based 3D rendering methods. Neural radiance field uses neural networks to learn the 3D structure and lighting information of a scene from 2D image data. Taking 2D images from multiple perspectives as input, each point in the scene is modeled through a multi-layer perceptron (MLP). Unlike traditional 3D reconstruction methods, neural radiance field technology can generate high-fidelity images with photo-quality by parameterizing the radiance field and density field in 3D space and combining it with differentiable volume rendering technology. This technology not only improves the reconstruction accuracy of 3D scenes, but also shows significant advantages in tasks such as image generation and new perspective synthesis. It has broad application prospects in virtual reality, autonomous driving, interior and exterior design, and other fields.

[0003] As an innovative derivative technology, Relighting Neural Radiance Fields is dedicated to simulating the visual effects of scenes under different lighting conditions. Based on the combination of neural radiance field technology and lighting models for new perspective image rendering, Relighting Neural Radiance Fields can quickly generate high-quality images of scenes under new lighting conditions by learning the interactive relationship between the distribution of light sources and three-dimensional scenes. Although Relighting Neural Radiance Fields has made significant progress in lighting reconstruction, due to the complexity of lighting and the diversity of scene geometry, it is often unstable when dealing with high-frequency details and complex lighting interactions, and may experience uneven lighting distribution or loss of details. Summary of the invention

[0004] Aiming at the problems existing in existing neural radiation fields and technologies in relighting tasks, the present invention provides a neural radiation field relighting method based on progressive grid refinement;

[0005] A neural radiation field relighting method based on progressive grid refinement includes the following steps:

[0006] Step 1: Construct a data set and divide it into a training set and a test set;

[0007] Step 1-1: Collect pictures taken by the camera at different viewing angles to build a data set, obtain the pictures in the data set and the corresponding internal parameters, where the internal parameters include the camera viewing direction v, the camera position o when the picture was taken, and the light source position P l ;

[0008] Step 1-2: Divide the data set into training set and test set according to the set ratio;

[0009] Step 2: Launch a sampling ray from the camera position o in the data set and perform discrete point sampling;

[0010] Step 2-1: For each image in the data set, generate a sampling ray p(t) according to the camera viewing direction: p(t) = o + tv, where o is the origin of the sampling ray. That is, three-dimensional data, v is the direction of the sampling ray, That is, two-dimensional data, the sampling ray starts from the camera position, t represents the sampling distance, passes through the pixel points on the two-dimensional image, and extends to the three-dimensional scene under the view corresponding to this image;

[0011] Step 2-2: Perform random uniform sampling on the sampling ray. The sampling formula is as follows:

[0012] Among them, t i represents the sampling distance of the i-th sampling point, t n represents the proximal sampling point on the ray, t f represents the far sampling point on the ray, and N represents the number of sampling points;

[0013] Step 3: Construct a progressive multi-resolution feature grid for feature extraction;

[0014] Step 3-1: Multi-resolution representation of grids: Grids with multiple resolutions are constructed in a progressive manner, with each layer of grids corresponding to spatial detail information at different scales; specifically represented as G r ={G1,G2…G n},in G r represents the grid of the rth layer, N r is the resolution of the grid at this layer, and n is the number of grid layers;

[0015] Step 3-2: Feature extraction process: In the progressive grid, the features of each grid layer are refined layer by layer, and the multi-layer perceptron MLP is used to extract the features of the three-dimensional space from grids of different resolutions; for each layer of grid G r , the MLP network takes the sampling point position x and the sampling ray direction v obtained in step 2 as input, and obtains the feature representation of the input sampling point learned by the MLP, that is, the SDF value;

[0016] Step 3-3: Progressive training: Adopt a progressive training strategy. First, train on a low-resolution grid, and gradually introduce higher-resolution grid learning using a residual network for training and optimization.

[0017] Step 4: Construct position encodings of corresponding frequencies based on the progressive feature grid.

[0018] Step 4-1: Using the idea of progressive multi-resolution grid refinement and combining position encoding technology, generate encodings of different frequencies for each sampling point of the neural radiance field. Specifically, grids of different resolutions correspond to Fourier position encodings of different frequencies. For each grid point, use its spatial coordinates (X, Y, Z) to perform position frequency encoding using Fourier transform. Specifically, the position frequency encoding function is expressed as:

[0019] β(x): {sin(2 0 πx), cos(2 0 πx), sin(2 1 πx), cos(2 1 πx)... sin(2 L πx), cos(2 L πx)},

[0020] where x is the spatial coordinate of the grid point, X is the horizontal coordinate, representing the distance of the point along the horizontal direction relative to a certain reference plane (usually the horizontal plane). Y is the depth coordinate, representing the distance of the point along the front-back direction relative to the reference plane. Z is the vertical coordinate, representing the distance of the point along the up-down direction relative to the reference plane. L is the maximum frequency level of the encoding, corresponding to the resolution of the grid, and 2 L represents the calibration index of the high-frequency part;

[0021] Step 4-2: Progressive refinement encoding: Use the Fourier encoding of the set minimum frequency to capture large-scale scene features; while on high-resolution grids, the frequency gradually increases; assuming the grid has n levels, for the l-th level grid, the maximum frequency of its frequency encoding is 2 l-1 , l < n, that is, as the resolution increases, the frequency also increases exponentially, expressed as follows:

[0022] β l (x) = {sin(2 0 πx), cos(2 0 πx), sin(2 1 πx), cos(2 1 πx)... sin(2 l-1 πx), cos(2 l-1 πx)}

[0023] Step 5: Connect multiple multi-layer perceptrons (MLPs) using a residual network to achieve implicit expression of the scene;

[0024] The implicit expression is defined as a function mapping:

[0025] F: x, v→σ, c, where x is the spatial coordinate of the sampling point and v is the direction of the sampling ray. The network structure is composed of multiple multi-layer perceptrons MLP, including F1, F2, ... F n , the input of each MLP is the feature vector output by the previous MLP layer and the position encoding set of the corresponding frequency at this level;

[0026] Step 5-1: The input of the first-level MLP network F1 is the concatenation of the uncoded spatial coordinates x of the sampling point and the frequency-coded data β1(x). The output vector is the feature vector O1 learned by this layer, expressed as:

[0027] F1:x,β1(x)→O1,l=1

[0028] Step 5-2: The second layer and each subsequent layer of the MLP network input the output O from the previous layer and the position frequency code β(x) of the current layer; specifically, the MLP network of the lth layer converts the output O of the MLP of the l-1th layer into l-1 and the position frequency coding set β at this level l (x) as input and output the intermediate feature vector O learned at this level l , the formula is as follows:

[0029] F l :O l-1 ,β l (x)→O l ,l>1

[0030] Step 5-3: Connect each layer of network through the residual network to form the overall network architecture. Each position frequency encoding passes through the corresponding network layer to generate feature output; the residual network structure is used to connect the output of each layer of frequency, which is specifically expressed as follows:

[0031] O l =O l-1 +w l β l (x)+k, where w l is the weighting coefficient, k is a constant;

[0032] Step 5-4: Spectral encoding of the sampling ray direction v:

[0033]

[0034] Step 5-5: The feature vector output by the last layer of the network is O n , the output is passed through the linear layer to obtain the final learned SDF value sdf, and then the sampling ray direction information γ(v) obtained after frequency encoding is combined with the feature vector O n The concatenation is performed and sent to the additional fully connected ReLU layer, and finally the color value c is predicted by the sigmoid activation function;

[0035] Step 6: Estimation of the visibility of the sampling points to reconstruct the lighting and shadows in the relighting task; perform visibility estimation on each sampling point in the scene to determine whether the point will be directly illuminated by the light source under given lighting conditions;

[0036] Step 6-1: Calculate the sampling point p i With the light source point position P l Straight-line distance Z in three-dimensional space frag The formula is as follows:

[0037] Z frag =‖xP l ‖

[0038] Where x is the spatial coordinate of the sampling point, P l is the spatial coordinate of the light source point.

[0039] Step 6-2: Switch the camera position o to the light source point position P l , repeat steps 2 to 3 to learn the SDF value of the sampling point on the ray, and record the sampling point with SDF value 0 as q; connect the sampling point p i With the light source position P l , detect the SDF value of 0 on the straight line connecting the two points and record it as q i , calculate q i Point and light source position P l Straight-line distance Z in three-dimensional space shadow The formula is as follows:

[0040] Z shadow =‖x q -P l ‖

[0041] Among them, x q is the sampling point q i The spatial coordinates of l is the spatial coordinate of the light source point;

[0042] Step 6-3: Estimate the sampling point p i Visibility through Z frag and Z shadow To make visibility judgment; when Zfrag Greater than Z shadow When , it means that there is an occlusion on the line between the sampling point and the light source, so the sampling point is in the shadow area and the visibility is 0; otherwise, the visibility is 1; the formula is as follows:

[0043]

[0044] Among them, V p is the visibility of the sampling ray where the sampling point is located;

[0045] Step 7: Render the image using the volume rendering formula;

[0046] Step 7-1: Calculate the volume density σ corresponding to each sampling point on the sampling ray i :

[0047] Ω(x)=se -sx / (1+e -sx ) 2 ,σ i =Ω(F(p i ))

[0048] Among them, Ω(x) represents the density function, s is a hyperparameter, F is the network function for learning the SDF value, and p i Indicates the sampling point;

[0049] Step 7-2: Calculate the cumulative value of the volume density corresponding to all sampling points on the sampling ray, that is, the opacity T(t):

[0050]

[0051] Where t is the sampling distance and σ is the volume density of the sampling points.

[0052] Step 7-3: Calculate the density weight ω(t) corresponding to each sampling point on the sampling ray:

[0053] ω(t)=T(t)σ(t)

[0054] Step 7-4: Use the visibility V of the sampling ray p , the density weights ω(t) of all sampling points on the sampling ray, and the color values ​​c corresponding to the sampling points are calculated by cumulative summation to obtain the color value C(p) of the pixel point:

[0055] Where N is the number of sampling points, ω i Represents the density weight value of the sampling point, c i Represents the color learned by the sampling point.

[0056] Step 8: Calculate the overall loss;

[0057] Step 8-1: Calculate rendering loss;

[0058] The rendering loss is expressed as follows:

[0059]

[0060] Among them, C(p) represents the rendered RGB value, C gt (p) represents the true RGB value of the pixel, R(p) is the set of all rays, and p is the sampling ray.

[0061] Step 8-2: Calculate the SDF regularization loss;

[0062] The SDF regularization loss is expressed as follows:

[0063]

[0064] Where x is the sampling point on the sampling ray p.

[0065] Step 8-3: Calculate the overall loss;

[0066] The overall loss is the weighted sum of the rendering loss and the SDF regularization loss, expressed as follows:

[0067] l=l color +λl eikonal , where λ is a hyperparameter.

[0068] Step 9: Iterative training, iteratively train the MLP network, repeat steps 2 to 8, and use different training samples for optimization each time until the performance of the neural network reaches the predetermined convergence standard. Finally, save the trained model.

[0069] Step 10: Using the Volume Rendering Formula Generate new perspective images.

[0070] The beneficial effects of adopting the above technical solution are:

[0071] The present invention provides a neural radiation field relighting method based on progressive grid refinement. The present invention solves the problems of blurred detail reconstruction and high-frequency perception ambiguity in existing neural radiation field-based relighting methods. By adopting progressive grid refinement and frequency coding optimization technology, it is possible to generate smoother and clearer rendering images while retaining details, thereby significantly improving the quality of scene reconstruction. The proposed network architecture adopts an end-to-end training model, does not rely on supervision at all levels, has good results and a simple implementation method, and the modular design facilitates later upgrades and optimizations, and has strong deployability. This method has low requirements on hardware equipment and can be directly applied on consumer-grade graphics cards. It is widely used in tasks in the fields of augmented reality, autonomous driving, game design, interior and exterior design, etc. It specifically has the following beneficial effects:

[0072] 1. The present invention effectively improves the ability of neural radiation field to reconstruct high-frequency details through the combination of progressive grid refinement and frequency encoding. Through the dynamic adjustment of grids of different resolutions and corresponding frequencies, fine incremental details are provided for the neural radiation field, significantly improving the image quality in the relighting task.

[0073] 2. The present invention introduces a residual network structure, combined with gradually increasing frequency coding, which makes the network more flexible and efficient in processing high-frequency information. At the same time, this design optimizes the network training process, reduces training time, and effectively overcomes the problems caused by spectrum deviation.

[0074] 3. The present invention effectively solves the problems of detail reconstruction blur and high-frequency perception ambiguity in existing neural radiation field methods under conditions of changing lighting, and significantly improves image clarity and detail reconstruction accuracy.

[0075] 4. The present invention provides an end-to-end training framework with a simple and efficient structure and no need for complex hierarchical supervision. The method has good scalability and optimization potential, can be directly deployed on consumer-grade graphics cards, and is widely used in augmented reality, autonomous driving, interior and exterior design and other fields, with significant practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is an overall flow chart of the neural radiation field relighting method according to an embodiment of the present invention;

[0077] Figure 2 This is a network structure diagram of the neural radiation field according to an embodiment of the present invention. DETAILED DESCRIPTION

[0078] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0079] A neural radiation field relighting method based on progressive grid refinement, such as Figure 1 As shown, the following steps are included:

[0080] Step 1: Dataset input;

[0081] Step 1-1: Collect pictures taken by the camera at different viewing angles to build a data set, obtain the pictures in the data set and the corresponding internal parameters, where the internal parameters include the camera viewing direction v, the camera position o when the picture was taken, and the light source position P l ;

[0082] Step 1-2: Divide the dataset into training and test sets. The dataset contains data from 6 scenes, each scene contains 600 images with a resolution of 512×512 pixels. The dataset is divided into training and test sets, specifically 500 images for training and 100 images for testing. The training and test sets maintain the same distribution characteristics to ensure the generalization ability of the model.

[0083] Step 2: Sampling rays are emitted from the camera position o in each image in the dataset, and rays are generated according to the viewing direction of the sampling rays. The rays are emitted from the camera position, pass through the pixels of the two-dimensional image, and extend along the three-dimensional scene where each pixel is located. Each sampling ray is uniformly sampled, and the feature value of each sampling point is calculated.

[0084] Step 2-1: For each image in the data set, generate a sampling ray p(t) according to the camera viewing direction: p(t) = o + tv, where o is the origin of the sampling ray. That is, three-dimensional data, v is the direction of the sampling ray, That is, two-dimensional data, the sampling ray starts from the camera position, t represents the sampling distance, passes through the pixel points on the two-dimensional image, and extends to the three-dimensional scene under the view corresponding to this image;

[0085] Step 2-2: Perform random uniform sampling on the sampling ray. The sampling formula is as follows:

[0086] Among them, t i represents the sampling distance of the i-th sampling point, t n represents the proximal sampling point on the ray, t f represents the far sampling point on the ray, and N represents the number of sampling points;

[0087] Step 3: Construct a progressive multi-resolution feature grid for feature extraction. Perform initial training on a low-resolution grid, and then gradually introduce higher-resolution grids to refine geometric features, thereby improving the expressiveness of the model. Figure 2 shown.

[0088] Step 3-1: Multi-resolution representation of grids: Grids with multiple resolutions are constructed in a progressive manner, with each layer of grids corresponding to spatial detail information at different scales; low-resolution grids mainly capture macroscopic structures, while high-resolution grids are refined to more microscopic geometric features. Specifically represented as G r ={G1,G2…G n},in G r represents the grid of the rth layer, N r is the resolution of the grid layer, and n is the number of grid layers. Each grid layer refines the representation of the surface and details of the object, and higher-level grids can capture more detailed information.

[0089] Step 3-2: Feature extraction process: In the progressive grid, the features are extracted layer by layer for each grid layer, and each layer of the neural network extracts features based on the grid of the current resolution. Multilayer perceptron MLP is used to extract the features of the three-dimensional space from grids of different resolutions; for each layer of grid G r , the MLP network takes the sampling point position x and the sampling ray direction v obtained in step 2 as input, and obtains the feature representation of the input sampling point learned by the MLP, that is, the SDF value;

[0090] Step 3-3: Progressive training: Use a progressive training strategy, first train on low-resolution grids, gradually use the residual network to introduce higher-resolution grid learning, and perform more refined training and optimization. This method helps the network learn the rough structure of the scene in the early stage, and then gradually optimize the details in subsequent training, thereby improving the efficiency of training and the speed of convergence.

[0091] Step 4: Construct the position code of the corresponding frequency according to the progressive feature grid;

[0092] Step 4-1: Using the idea of ​​progressive multi-resolution grid refinement and combining it with position coding technology, generate different frequency codes for each sampling point of the neural radiation field. Specifically, grids of different resolutions correspond to Fourier position codes of different frequencies, which will be used to enhance the feature expression ability of the neural network at different scales. For each grid point, use its spatial coordinates (X, Y, Z) to perform position-frequency coding using Fourier transform. Specifically, the position-frequency coding function is expressed as:

[0093] β(x): {sin(2 0 πx),cos(2 0 πx),sin(2 1 πx),cos(2 1 πx)...sin(2 L πx),cos(2L πx)}

[0094] Where x is the spatial coordinate of the grid point, X is the horizontal coordinate, representing the distance of the point along the horizontal direction relative to a certain reference plane (usually the horizontal plane). Y is the depth coordinate, representing the distance of the point along the front-back direction relative to the reference plane. Z is the vertical coordinate, representing the distance of the point along the up-down direction relative to the reference plane. L is the maximum frequency level of the encoding, corresponding to the resolution of the grid, 2 L represents the calibration index of the high-frequency part, which can help the model better capture high-frequency details;

[0095] Step 4-2: Progressive refinement encoding: During the progressive grid refinement process, the resolution of the grid gradually increases with the increase of the level. To match the grid resolution, the frequency encoding also increases accordingly. On the low-resolution grid, we use the Fourier encoding with the set minimum frequency to capture the large-scale scene features; while on the high-resolution grid, the frequency gradually increases to capture more delicate high-frequency details. Assume the grid has n levels. For the l-th level grid, the maximum frequency of its frequency encoding is 2 l-1 , l < n, that is, as the resolution increases, the frequency also grows exponentially, expressed as follows:

[0096] β l (x) = {sin(2 0 πx), cos(2 0 πx), sin(2 1 πx), cos(2 1 πx)... sin(2 l-1 πx), cos(2 l-1 πx)}

[0097] Step 5: Connect multiple multi-layer perceptrons MLP using a residual network to achieve an implicit representation of the scene;

[0098] The implicit representation is defined as a function mapping:

[0099] F: x, v → σ, c, where x is the spatial coordinate of the sampling point, v is the direction of the sampling ray, The network structure is composed of multiple multi-layer perceptrons MLP, specifically including F1, F2,... F n , and the input of each MLP is the feature vector output by the previous layer MLP and the set of position encodings corresponding to the frequency of this layer;

[0100] Step 5-1: The input of the first-level MLP network F1 is the concatenation of the unencoded spatial coordinate x of the sampling point and the data β1(x) after frequency encoding, and the output vector is the feature vector O1 learned by this layer, expressed as:

[0101] F1:x,β1(x)→O1,l=1

[0102] Step 5-2: The second layer and each subsequent layer of the MLP network input the output O from the previous layer and the position frequency code β(x) of the current layer; specifically, the MLP network of the lth layer converts the output O of the MLP of the l-1th layer into l-1 and the position frequency coding set β at this level l (x) as input and output the intermediate feature vector O learned at this level l , the formula is as follows:

[0103] F l :O l-1 ,β l (x)→O l ,l>1

[0104] Step 5-3: Connect each layer of the network through the residual network to form the overall network architecture. This progressive residual structure can enhance the feature expression of each layer layer by layer. Each position frequency code passes through the corresponding network layer to generate feature output; in order to better integrate the information of different frequencies, the residual network structure is used to connect the output of each layer of frequency, which is specifically expressed as follows:

[0105] O l =O l-1 +w l β l (x)+k, where w l is a weighting coefficient used to control the contribution of each layer output, and k is a constant. This progressive residual structure enables the model to learn incremental information relative to the previous layer, thereby progressively improving the network representation capability.

[0106] Step 5-4: Spectral encode the sampling ray direction v using the existing frequency encoding formula:

[0107]

[0108] Step 5-5: The feature vector output by the last layer of the network is O n , the output is passed through the linear layer to obtain the final learned SDF value sdf, and then the sampling ray direction information γ(v) obtained after frequency encoding is combined with the feature vector O n The concatenation is performed and sent to the additional fully connected ReLU layer, and finally the color value c is predicted by the sigmoid activation function;

[0109] Step 6: Sampling point visibility estimation, reconstruction of lighting and shadows in the relighting task. In the relighting task, accurate lighting and shadow calculation is the key. Visibility estimation is performed for each sampling point in the scene to determine whether the point is directly illuminated by the light source under given lighting conditions.

[0110] Step 6-1: Calculate the sampling point p i With the light source position P l Straight-line distance Z in three-dimensional space frag The formula is as follows:

[0111] Z frag =‖xP l ‖

[0112] Where x is the spatial coordinate of the sampling point, P l is the spatial coordinate of the light source point.

[0113] Step 6-2: Switch the camera position o to the light source point position P l , repeat steps 2 to 3 to learn the SDF value of the sampling point on the ray, and record the sampling point with SDF value 0 as q; connect the sampling point p i With the light source position P l , detect the SDF value of 0 on the straight line connecting the two points and record it as q i , calculate q i Point and light source position P l Straight-line distance Z in three-dimensional space shadow The formula is as follows:

[0114] Z shadow =‖x q -P l ‖

[0115] Among them, x q is the sampling point q i The spatial coordinates of l is the spatial coordinate of the light source point;

[0116] Step 6-3: Estimate the sampling point p i Visibility through Z frag and Z shadow To make visibility judgment; when Z frag Greater than Z shadow When , it means that there is an occlusion on the line between the sampling point and the light source, so the sampling point is in the shadow area and the visibility is 0; otherwise, the visibility is 1; the formula is as follows:

[0117]

[0118] Among them, V pis the visibility of the sampling ray where the sampling point is located;

[0119] Step 7: Render the image using the volume rendering formula;

[0120] Step 7-1: Calculate the volume density σ corresponding to each sampling point on the sampling ray i :

[0121] Ω(x)=se -sx / (1+e -sx ) 2 ,σ i =Ω(F(p i ))

[0122] Among them, Ω(x) represents the density function, s is a hyperparameter, F is the network function for learning the SDF value, and p i Indicates the sampling point.

[0123] Step 7-2: Calculate the cumulative value of the volume density corresponding to all sampling points on the sampling ray, that is, the opacity T(t):

[0124]

[0125] Where t is the sampling distance and σ is the volume density of the sampling points.

[0126] Step 7-3: Calculate the density weight ω(t) corresponding to each sampling point on the sampling ray:

[0127] ω(t)=T(t)σ(t)

[0128] Step 7-4: Use the visibility V of the sampling ray p , the density weights ω(t) of all sampling points on the sampling ray, and the color values ​​c corresponding to the sampling points are calculated by cumulative summation to obtain the color value C(p) of the pixel point:

[0129] Where N is the number of sampling points, ω i Represents the density weight value of the sampling point, c i Represents the color learned by the sampling point.

[0130] Step 8: Calculate the overall loss; jointly train the network using the rendering loss and the SDF regularization loss.

[0131] Step 8-1: Calculate rendering loss;

[0132] Calculate the mean square error between the true RGB value of the pixel and the rendered RGB value, expressed as follows:

[0133]

[0134] Among them, C(p) represents the rendered RGB value, C gt (p) represents the true RGB value of the pixel, R(p) is the set of all rays, and p is the sampling ray.

[0135] Step 8-2: Calculate the SDF regularization loss;

[0136] The SDF regularization loss is used to ensure that the gradient of the learned SDF function is close to 1 near the surface of the object, thereby ensuring that the SDF can accurately represent the boundary of the object. It is expressed as follows:

[0137]

[0138] Where x is the sampling point on the sampling ray p.

[0139] Step 8-3: Calculate the overall loss;

[0140] The overall loss is the weighted sum of the rendering loss and the SDF regularization loss, expressed as follows:

[0141] l=l color +λl eikonal , where λ is a hyperparameter.

[0142] Step 9: During the experiment, we trained on Nvidia 2080Ti GPU and used PyTorch framework to implement neural network training. Each model was trained for 1000k iterations and optimized using Adam optimizer. We used PSNR (peak signal-to-noise ratio) as the evaluation metric. PSNR is mainly used to evaluate the detail restoration and structural consistency between the generated image and the real image.

[0143] Step 10: During the testing phase, we use the trained model to synthesize images from new perspectives. We tested it in 6 different scenes (Verdant Guardian, Blush Elephant, Blushing Girl, Tiny Guardian, Wise Owl, Colorful Bloom).

[0144] Comparison of PSNR↑ indicators on the Area light-Blender dataset

[0145]

[0146] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A neural radiation field relighting method based on progressive grid refinement, characterized in that: The following steps are involved: Step 1: Construct a data set and divide it into a training set and a test set; Step 2: Launch a sampling ray from the camera position o in the data set and perform discrete point sampling; Step 3: Construct a progressive multi-resolution feature grid for feature extraction; Step 4: Construct the position code of the corresponding frequency according to the progressive feature grid; Step 5: Connect multiple multi-layer perceptrons (MLPs) using a residual network to achieve implicit expression of the scene; The implicit expression is defined as a function mapping: F: x, v→σ, c, where x is the spatial coordinate of the sampling point and v is the direction of the sampling ray. The network structure is composed of multiple multi-layer perceptrons MLP, including F1, F2, ... F n , the input of each MLP is the feature vector output by the previous MLP layer and the position encoding set of the corresponding frequency at this level; Step 6: Estimation of the visibility of the sampling points to reconstruct the lighting and shadows in the relighting task; perform visibility estimation on each sampling point in the scene to determine whether the point will be directly illuminated by the light source under given lighting conditions; Step 7: Render the image using the volume rendering formula; Step 8: Calculate the overall loss; Step 9: Iterative training: iteratively train the MLP network, repeating steps 2 to 8, using different training samples for optimization each time, until the performance of the neural network reaches the predetermined convergence standard; finally, save the trained model; Step 10: Using the Volume Rendering Formula Generate new perspective images.

2. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1-1: Collect pictures taken by the camera at different viewing angles to build a data set, obtain the pictures in the data set and the corresponding internal parameters, where the internal parameters include the camera viewing direction v, the camera position o when the picture was taken, and the light source position P l ; Step 1-2: Divide the data set into training set and test set according to the set ratio.

3. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 2 includes the following steps: Step 2-1: For each image in the data set, generate a sampling ray p(t) according to the camera viewing direction: p(t) = o + tv, where o is the origin of the sampling ray. That is, three-dimensional data, v is the direction of the sampling ray, That is, two-dimensional data, the sampling ray starts from the camera position, t represents the sampling distance, passes through the pixel points on the two-dimensional image, and extends to the three-dimensional scene under the view corresponding to this image; Step 2-2: Perform random uniform sampling on the sampling ray. The sampling formula is as follows: Among them, t i represents the sampling distance of the i-th sampling point, t n represents the proximal sampling point on the ray, t f represents the far sampling point on the ray, and N represents the number of sampling points.

4. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 3 includes the following steps: Step 3-1: Multi-resolution representation of grids: Grids with multiple resolutions are constructed in a progressive manner, with each layer of grids corresponding to spatial detail information at different scales; specifically represented as G r ={G1,G2…G n },in G r represents the grid of the rth layer, N r is the resolution of the grid at this layer, and n is the number of grid layers; Step 3-2: Feature extraction process: In the progressive grid, the features of each grid layer are refined layer by layer, and the multi-layer perceptron MLP is used to extract the features of the three-dimensional space from grids of different resolutions; for each layer of grid G r , the MLP network takes the sampling point position x and the sampling ray direction v obtained in step 2 as input, and obtains the feature representation of the input sampling point learned by the MLP, that is, the SDF value; Step 3-3: Progressive training: Using a progressive training strategy, first train on a low-resolution grid, and gradually use the residual network to introduce higher-resolution grid learning for training and optimization.

5. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 4 includes the following steps: Step 4-1: Using the idea of ​​progressive multi-resolution grid refinement and combining it with position coding technology, generate different frequency codes for each sampling point of the neural radiation field; specifically, grids of different resolutions correspond to Fourier position codes of different frequencies. For each grid point, its spatial coordinates (X, Y, Z) are used to perform position frequency coding using Fourier transform. Specifically, the position frequency coding function is expressed as: β(x):{sin(2 0 πx),cos(2 0 πx),sin(2 1 πx),cos(2 1 πx)...sin(2 L πx),cos(2 L πx)}, Among them, x is the spatial coordinate of the grid point, X is the horizontal coordinate, which indicates the distance of the point relative to a reference plane (usually a horizontal plane) along the horizontal direction; Y is the depth coordinate, which indicates the distance of the point relative to the reference plane along the front-back direction; Z is the vertical coordinate, which indicates the distance of the point relative to the reference plane along the up-down direction; L is the maximum frequency level of the encoding, which corresponds to the resolution of the grid, 2 L Indicates the calibration index of the high frequency part; Step 4-2: Progressive refinement encoding: Use Fourier encoding with a set minimum frequency to capture large-scale scene features; while on high-resolution grids, the frequency gradually increases; assuming the grid has n levels, for the l-th level grid, the maximum frequency of its frequency encoding is 2 l-1 , l < n, that is, as the resolution increases, the frequency also grows exponentially, expressed as follows: β l (x)={sin(2 0 πx),cos(2 0 πx),sin(2 1 πx),cos(2 1 πx)...sin(2 l-1 πx),cos(2 l-1 πx)}。 6. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 5 includes the following steps: Step 5-1: The input of the first-level MLP network F1 is the concatenation of the uncoded spatial coordinates x of the sampling point and the frequency-coded data β1(x). The output vector is the feature vector O1 learned by this layer, expressed as: F1:x,β1(x)→O1,l=1 Step 5-2: The second layer and each subsequent layer of the MLP network input the output O from the previous layer and the position frequency code β(x) of the current layer; specifically, the MLP network of the lth layer converts the output O of the MLP of the l-1th layer into l-1 and the position frequency coding set β at this level l (x) as input and output the intermediate feature vector O learned at this level l , the formula is as follows: F l :O l-1 ,b l (x)→O l ,l>1 Step 5-3: Connect each layer of network through the residual network to form the overall network architecture. Each position frequency encoding passes through the corresponding network layer to generate feature output; the residual network structure is used to connect the output of each layer of frequency, which is specifically expressed as follows: O l =O l-1 +w l β l (x)+k, where w l is the weighting coefficient, k is a constant; Step 5-4: Spectral encoding of the sampling ray direction v: Step 5-5: The feature vector output by the last layer of the network is O n , the output is passed through the linear layer to obtain the final learned SDF value sdf, and then the sampling ray direction information γ(v) obtained after frequency encoding is combined with the feature vector O n The resulting images are concatenated and sent to an additional fully connected ReLU layer, and finally the color value c is predicted by the sigmoid activation function.

7. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 6 includes the following steps: Step 6-1: Calculate the sampling point p i With the light source position P l Straight-line distance Z in three-dimensional space frag The formula is as follows: Z frag =‖x-P l ‖ Where x is the spatial coordinate of the sampling point, P l is the spatial coordinate of the light source point; Step 6-2: Switch the camera position o to the light source point position P l , repeat steps 2 to 3 to learn the SDF value of the sampling point on the ray, and record the sampling point with SDF value 0 as q; connect the sampling point p i With the light source point position P l , detect the SDF value of 0 on the straight line connecting the two points and record it as q i , calculate q i Point and light source position P l Straight-line distance Z in three-dimensional space shadow The formula is as follows: Z shadow =‖x q -P l ‖ Among them, x q is the sampling point q i The spatial coordinates of l is the spatial coordinate of the light source point; Step 6-3: Estimate the sampling point p i Visibility through Z frag and Z shadow To make visibility judgment; when Z frag Greater than Z shadow When , it means that there is an occlusion on the line between the sampling point and the light source, so the sampling point is in the shadow area and the visibility is 0; otherwise, the visibility is 1; the formula is as follows: Among them, V p is the visibility of the sampling ray at the sampling point.

8. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 7 includes the following steps: Step 7-1: Calculate the volume density σ corresponding to each sampling point on the sampling ray i : Ω(x) = se -sx / (1+e -sx ) 2 ,σ i =Ω(F(p i )) Among them, Ω(x) represents the density function, s is a hyperparameter, F is the network function for learning the SDF value, and p i Indicates the sampling point; Step 7-2: Calculate the cumulative value of the volume density corresponding to all sampling points on the sampling ray, that is, the opacity T(t): Where t is the sampling distance, σ is the volume density of the sampling points; Step 7-3: Calculate the density weight ω(t) corresponding to each sampling point on the sampling ray: ω(t)=T(t)σ(t) Step 7-4: Use the visibility V of the sampling ray p , the density weights ω(t) of all sampling points on the sampling ray, and the color values ​​c corresponding to the sampling points are calculated by cumulative summation to obtain the color value C(p) of the pixel point: Where N is the number of sampling points, ω i Represents the density weight value of the sampling point, c i Represents the color learned by the sampling point.

9. The neural radiation field relighting method based on progressive grid refinement according to claim 1, characterized in that: Step 8 includes the following steps: Step 8-1: Calculate rendering loss; The rendering loss is expressed as follows: Among them, C(p) represents the rendered RGB value, C gt (p) represents the true RGB value of the pixel, R(p) is the set of all rays, and p is the sampling ray; Step 8-2: Calculate the SDF regularization loss; The SDF regularization loss is expressed as follows: Where x is the sampling point on the sampling ray p; Step 8-3: Calculate the overall loss; The overall loss is the weighted sum of the rendering loss and the SDF regularization loss, expressed as follows: l=l color +λl eikonal , where λ is a hyperparameter.

Citation Information

Patent Citations

  • Nerve radiation field training method based on grid representation and image rendering method

    CN115731340A

  • New view angle synthesis method based on spatial progressive neural radiation field

    CN118298092A