A general controllable spatial perception based radiation field enhancement method
By constructing a neural radiation field hybridization model and dynamically fusing the volume density and color weights of adjacent sampling points, the problem of insufficient spatial dependency modeling in new perspective synthesis of neural radiation fields is solved, thereby improving rendering efficiency and image quality.
Patent Information
- Application Number
- CN202411537195.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing neural radiation field-based methods suffer from insufficient spatial dependency modeling in novel perspective synthesis, resulting in low rendering efficiency and poor image quality.
A neural radiation field mixing model is constructed, and the volume density and color weight of adjacent sampling points are dynamically fused through cross-ray and same-ray cross-sampling point mixing modules to improve spatial perception capabilities.
While increasing inference time by 25%, it significantly improves the image quality of new view synthesis and enhances the expressiveness of scene structure.
Smart Images

Figure CN119399346B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a radiation field enhancement method, in particular to a general controllable space perception based radiation field enhancement method. BACKGROUND
[0002] The information provided in this section is merely background information related to the present disclosure and can not necessarily be prior art.
[0003] With the emergence of concepts such as the metaverse, virtual reality, and digital twins, there is an increasing demand for digital reconstruction of complex scenes to produce realistic and high-quality pictures. How to render a more realistic virtual scene from an image is a topic of focus in the field of computer graphics and computer vision. Among them, new view synthesis is a task of focus. Through limited two-dimensional views, a continuous three-dimensional scene is reconstructed to accurately restore the geometric structure, surface texture, and lighting of the scene, so as to generate realistic new views, which is of great significance to the fields of augmented reality, virtual reality, and video games.
[0004] Traditional image-based rendering synthesis scene representation methods require the geometry and properties of objects as a priori, and have large computational complexity and low rendering efficiency. With the rapid development of deep learning-driven image rendering technology, Neural Radiance Fields (NeRF) has pioneered the use of neural networks to represent scenes, enabling photo-realistic new view synthesis capabilities.
[0005] NeRF was first proposed by Mildenhall et al. in 2020 ECCV conference. This method represents a three-dimensional scene as a radiance field approximated by a multi-layer perception (MLP), describing the color and volume density of each sampling point in the scene at a specific observation direction. The method uses volume rendering to synthesize pictures at new viewing angles.
[0006] Previously, there have been many improvements to address the slow speed of NeRF, the jagged phenomenon of rendered pictures, and sampling efficiency, but all of these improvements are based on the basic assumption that spatial points are independent of each other and that light rays are independent of each other, which has the problem of insufficient modeling of spatial dependencies.
[0007] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0008] The present application relates to a radiation field enhancement method, in particular to a general controllable space perception based radiation field enhancement method. The technical problem to be solved by the present application is to provide a general controllable space perception based radiation field enhancement method to address the deficiencies of the prior art.
[0009] To solve the above technical problems, the application discloses a general controllable spatial perception based radiation field enhancement method, comprising the following steps:
[0010] Step 1, obtaining training data;
[0011] Step 2, generating light rays and sampling points according to the training data;
[0012] Step 3, constructing a neural radiation field hybrid model;
[0013] Step 4, inputting the coordinates of the sampling points and the directions of the light rays into the neural radiation field hybrid model constructed in step 3 to calculate the volume density and color of each sampling point;
[0014] Step 5, calculating the RGB color value of the pixel corresponding to each light ray;
[0015] Step 6, training the neural radiation field hybrid model constructed in step 3 according to the volume density and color of each sampling point obtained in steps 4 and 5 and the RGB color value corresponding to each light ray to obtain a trained neural radiation field hybrid model;
[0016] Step 7, using the trained neural radiation field hybrid model to render a two-dimensional picture under a new view angle according to the camera pose of the view angle to complete general controllable spatial perception based radiation field enhancement.
[0017] Further, the training data obtained in step 1 is obtained by using a camera to shoot two-dimensional pictures under different view angles and recording the camera pose corresponding to each picture, and the two-dimensional pictures and the corresponding camera pose are used as training data.
[0018] Further, the generation of light rays and sampling points in step 2 is to model the light rays using a neural radiation field, wherein each pixel in the two-dimensional picture emits a light ray, the starting point of the light ray is the camera position, the direction is a direction vector from the camera position to the pixel position in the two-dimensional image plane, and each pixel corresponds to a light ray; the starting point coordinates and the direction vector of the light ray are calculated according to the camera pose, and the coordinates of the sampling points are obtained by sampling on the light ray, which specifically comprises:
[0019] Step 2-1, calculating the starting point coordinates and the direction vector of the light ray corresponding to the pixel in the world coordinate system according to the camera pose and the pixel coordinates in the two-dimensional image plane, and the specific method comprises:
[0020] randomly selecting an image block on each two-dimensional picture, calculating the starting point coordinates and the direction vector of the light ray corresponding to the pixel in the world coordinate system through the pixels of the image block, and setting the length and width of the image block as p h and pw , then the tensor corresponding to the RGB color of the image block pixel , the starting point coordinate tensor of the light corresponding to the image block pixel , the direction vector tensor
[0021] Step 2-2, using a uniform sampling method to sample on the light to generate initial sampling points, and calculating the coordinates of the initial sampling points in the world coordinate system;
[0022] The uniform sampling method described in step 2-2 is as follows:
[0023]
[0024] , where t n is the nearest distance of the sampling point from the camera, t f is the farthest distance of the sampling point from the camera, t i is the distance of the i-th sampling point from the camera obtained by the uniform sampling method, N init is the number of initial sampling points, U represents uniform distribution, and ~ represents a probability distribution.
[0025] Further, the construction of the neural radiance field hybrid model in step 3 includes the following steps:
[0026] Step 3-1, constructing a neural radiance field hybrid model with two granularities, namely a coarse model and a fine model; wherein the coarse model includes: a sampling point coordinate feature encoder, a direction vector feature encoder, a volume density decoder, and a color decoder; the output of the coarse model is the volume density σ coarse and the color c coarse of the sampling point.
[0027] The fine model includes: a sampling point coordinate feature encoder, a direction vector feature encoder, a volume density decoder, a color decoder, and a hybrid module; wherein the hybrid module includes: a cross-light cross-sampling point hybrid module and a same-light cross-sampling point hybrid module; the output of the fine model is the volume density feature F σ , the color feature F c , the volume density σ fine , the color c fine , the cross-light cross-sampling point hybrid volume density σ inter_mixed , the cross-light cross-sampling point hybrid color c rnter_mixed , the cross-light and same-light hybrid volume density σ intra_inter_mixed , and the cross-light and same-light hybrid color c intra_inter_mixed .
[0028] Step 3-2, constructing a cross-ray cross-sampling point mixing module in the radiance field mixing model;
[0029] Step 3-3, constructing a same-ray cross-sampling point mixing module in the radiance field mixing model.
[0030] Further, the sampling point coordinate feature encoder in step 3-1 is composed of eight two-dimensional convolution layers with a 1*1 convolution kernel, which encodes the input sampling point coordinate feature to obtain the volume density feature;
[0031] The volume density decoder is composed of one two-dimensional convolution layer with a 1*1 convolution kernel, which decodes the volume density feature to obtain the volume density of each sampling point;
[0032] The direction vector feature encoder is composed of one two-dimensional convolution layer with a 1*1 convolution kernel, which encodes the volume density feature and the input direction vector feature to obtain the color feature after splicing;
[0033] The color decoder is composed of one two-dimensional convolution layer with a 1*1 convolution kernel, which decodes the color feature to obtain the color of each sampling point.
[0034] Further, the cross-ray cross-sampling point mixing module in step 3-2 is composed of a volume density mixing subnetwork and a color mixing subnetwork, wherein each subnetwork is composed of two two-dimensional convolution layers with a k*k convolution kernel;
[0035] The cross-ray cross-sampling point mixing module is used to predict the volume density weight and color weight of the adjacent sampling points on different rays and at the same depth as each sampling point on the sampling point, to obtain the volume density σ inter_mixed and the color c inter_mixed after cross-ray cross-sampling point mixing of each sampling point, the specific prediction method is as follows:
[0036] Step 3-2-1, input the volume density feature F σ of the sampling point obtained in the fine model into the volume density mixing subnetwork to predict the volume density weight W_inter σ of the adjacent sampling points on different rays and at the same depth as the sampling point on the sampling point; fine Weighted sum is performed on the volume density σ inter_mixed after cross-ray cross-sampling point mixing, the specific method is as follows:
[0037]
[0038] Wherein, k is the number of adjacent sampling points, the volume density weight of the i-th neighboring sample point to the sample point, the volume density of the i-th neighboring sample point;
[0039] Step 3-2-2, the color feature F of the sample point obtained in the fine model c input into the color mixing sub-network, to predict the color weight W_inter of the neighboring sample point to the sample point at the same depth on the same light line c , the color c of the sample point is obtained by weighted summation fine , the specific method is as follows: inter_mixed
[0040]
[0041] wherein k is the number of neighboring sample points, the volume color weight of the i-th neighboring sample point to the sample point, the color of the i-th neighboring sample point.
[0042] Further, the same light line cross sample point mixing module in step 3-3 is composed of a volume density mixing sub-network and a color mixing sub-network, wherein each sub-network is composed of two one-dimensional convolution kernel layers with k convolution kernels;
[0043] The same light line cross sample point mixing module is used to predict the volume density weight and the color weight of the neighboring sample point to the sample point on the same light line, to obtain the volume density σ of each sample point after cross light line and same light line mixing intra_inter_mixed and the color c after cross light line and same light line mixing intra_inter_mixed , the specific prediction method is as follows:
[0044] Step 3-3-1, the volume density feature F of the sample point obtained in the fine model σ input into the volume density mixing sub-network, to predict the volume density weight W_intra of the neighboring sample point to the sample point on the same light line σ , the volume density σ of step 3-2 after cross light line cross sample point mixing inter_mixed is weighted and summed to obtain the volume density σ after cross light line and same light line mixing intra_inter_mixed , the specific method is as follows:
[0045]
[0046] wherein k is the number of neighboring sample points, the volume density weight of the i-th neighboring sample point to the sample point, the color of the i-th neighboring sample point after cross-ray cross-sample point mixing;
[0047] Step 3-3-2, input the color feature F of the sample point obtained in the fine model into the color mixing sub-network to obtain the color weight W_intra of the neighboring sample point on the same ray to the sample point c Step 3-3-3, input the color weight W_intra into the color mixing sub-network to obtain the color of the neighboring sample point on the same ray to the sample point c Step 3-3-4, input the color c of the sample point after cross-ray cross-sample point mixing in step 3-2 into the color mixing sub-network to obtain the color c of the sample point after cross-ray cross-sample point mixing and cross-ray mixing inter_mixed Step 3-3-5, perform weighted summation to obtain the color c of the sample point after cross-ray cross-sample point mixing and cross-ray mixing intra_inter_mixed The specific method is as follows:
[0048]
[0049] wherein k is the number of neighboring sample points, is the volume color weight of the i-th neighboring sample point to the sample point, is the color of the i-th neighboring sample point after cross-ray cross-sample point mixing.
[0050] Further, the calculation of the volume density and color of each sample point in step 4 includes the following steps:
[0051] Step 4-1, encode the coordinates and direction vectors of the sample point to obtain coordinate features and direction vector features, and the specific encoding method is as follows:
[0052] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp))
[0053] wherein p is the coordinates of the sample point or the direction vector of the ray on which the sample point is located, L is the encoding frequency, and γ(p) represents the encoded encoding features, wherein the coordinate encoding obtains the coordinate features, and the direction vector encoding obtains the direction vector features; the coordinate feature tensor is wherein D x = 3 × (2 × L x + 1); the direction vector feature tensor is V ∈ wherein D v = 3 × (2 × L v + 1), and N is the number of sample points on each ray;
[0054] Step 4-2, input the coordinate feature tensor and the direction vector feature tensor of the initial sample point in step 2-2 into the coarse model in step 3 to obtain the volume density of the initial sample point in the coarse model and color N init is the number of initial sampling points;
[0055] Step 4-3, using the hierarchical sampling method, based on the volume density σ coarse of the sampling points in the coarse model, the fine sampling points are resampled, specifically as follows:
[0056] Step 4-3-1, the volume density σ coarse of the sampling points on each ray in the coarse model is normalized and taken as the probability p i that the sampling point is located on the surface of the object in the current scene, and the specific method is as follows:
[0057]
[0058] where N init is the number of initial sampling points on each ray;
[0059] Step 4-3-2, according to the probability p i distribution that each sampling point on each ray is located on the surface of the object in the current scene, the hierarchical sampling points are generated by inverse transform sampling;
[0060] Step 4-3-3, the initial sampling points in step 2-2 and the hierarchical sampling points in step 4-3-2 are merged into the final fine sampling points;
[0061] Step 4-4, the coordinate and direction vector of the fine sampling points are encoded using the encoding method in step 4-1 to obtain the coordinate feature tensor and the direction vector feature tensor and input into the fine model constructed in step 3 to obtain the volume density feature color feature volume density color volume density mixed across rays and across sampling points color mixed across rays and across sampling points volume density mixed across rays and across sampling points color mixed across rays and across sampling points N fine is the number of fine sampling points, and the specific steps are as follows:
[0062] Step 4-4-1, through the coordinate feature encoder in the fine model, the volume density feature F σ in the fine model is obtained;
[0063] Step 4-4-2, obtain the color feature F in the fine model through the direction vector feature encoder in the fine model c ;
[0064] Step 4-4-3, obtain the volume density σ in the fine model through the volume density decoder in the fine model fine ;
[0065] Step 4-4-4, obtain the color c in the fine model through the color decoder in the fine model fine ;
[0066] Step 4-4-5, input the above volume density feature F inter_mixed , color feature F inter_mixed , volume density σ σ and color c c into the cross-ray cross-sampling point mixing module to obtain the cross-ray cross-sampling point mixed volume density σ inter_mixed and color c imter_mixed ;
[0067] Step 4-4-6, input the above volume density feature F intra_inter_mixed , color feature F intra_inter_mixed , cross-ray cross-sampling point mixed volume density σ i and cross-ray cross-sampling point mixed color c i into the same-ray cross-sampling point mixing module to obtain the cross-ray and same-ray mixed volume density σ j and color c j+1 .
[0068] Further, the method for calculating the RGB color value of the pixel corresponding to each ray in step 5 includes:
[0069]
[0070] wherein, represents the final RGB color value of the pixel corresponding to the ray, σ j and c i are the volume density and color of the i-th sampling point, δ coarse =t coarse -t fine is the distance between adjacent sampling points on the same ray, and T fine is the cumulative transmittance from the camera position to the i-th sampling point.
[0071] Further, the method for training the neural radiance field mixing model constructed in step 3 in step 6, i.e., updating the parameters of the neural radiance field mixing model using the Adam optimizer by calculating the loss, wherein the method for calculating the loss includes:
[0072] Step 6-1, for the volume density σ and color c obtained by the coarse model in step 4-2 coarse and color c coarse , calculate the pixel RGB value rendered by the radiance field in the coarse model by the method in step 5
[0073] Step 6-2, for the volume density σ and color c obtained by the fine model in step 4-2 fine and color c fine , calculate the pixel RGB value rendered by the radiance field in the fine model by the method in step 5
[0074] Step 6-3, for the volume density σ and color c mixed by the cross-ray cross-sampling point mixing module in the fine model in step 4-4 inter_mixed and color c inter_mixed , calculate the pixel RGB value rendered by the radiance field in the fine model after mixing by the cross-ray cross-sampling point mixing module by the method in step 5
[0075] Step 6-4, for the volume density σ and color c mixed by the same-ray cross-sampling point mixing module in step 4-5 intra_nter_mixed and color c intra_inter_mixed , calculate the pixel RGB value rendered by the radiance field in the fine model after mixing by the same-ray cross-sampling point mixing module by the method in step 5
[0076] Step 6-5, calculate the loss, as follows:
[0077] L RF = L coarse + L fine
[0078]
[0079] Wherein, L RF is the total loss of the neural radiance field mixing model, R is the set of rays, L coarse and L fine are the losses of the coarse model and the fine model respectively, is the true RGB value of the corresponding pixel of the image block pixel.
[0080] Advantages:
[0081] 1) The method proposes a radiance field mixing strategy, which eliminates the implicit assumption of independent spatial sampling points based on the neural radiance field method, and enhances its ability to structure the scene structure.
[0082] 2) The method can greatly improve the quality of new view synthesis based on neural radiance field method on the basis of increasing 25% reasoning time. BRIEF DESCRIPTION OF DRAWINGS
[0083] The above and / or other aspects of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:
[0084] Figure 1 Flowchart for the spatial perception based radiance field enhancement method.
[0085] Figure 2 Flowchart for calculating the density and color of the sampling points.
[0086] Figure 3 Process diagram for calculating the initial density and color of the sampling points of the coarse model.
[0087] Figure 4 Process diagram for calculating the density and color of the fine sampling points of the fine model. DETAILED DESCRIPTION
[0088] The present application provides a general controllable spatial perception based radiance field enhancement method to improve the rendering quality when synthesizing new view images, aiming at the shortcomings of the existing neural radiance field method. The specific technical solutions are as follows: The present application discloses a general controllable spatial perception based radiance field enhancement method. The core is to use dynamic kernel to locally fuse adjacent sampling points of the radiance field, improve the spatial perception ability of the features, and thus improve the rendering quality of the synthesized pictures under new view angles. As shown in the figure, it specifically includes the following steps: Figure 1
[0089] Step 1: Read the training data; read the two-dimensional pictures under different view angles and the camera poses corresponding to each picture;
[0090] Step 2: Generate light and sampling points; the modeling method of neural radiance field for light is as follows: one light is emitted for each pixel to simulate the process of light reaching the scene from the observer (or camera), the starting point of the light is usually the camera position, and the direction is the direction vector from the camera position to the specific pixel on the image plane, so each pixel corresponds to a light. Calculate the starting point coordinates and direction vector of the light through the camera pose, and sample the light to get the coordinates of the sampling points. The process includes the following steps:
[0091] Step 2-1: According to the camera pose and the pixel coordinates in the two-dimensional image plane, calculate the starting point coordinates and direction vector of the light corresponding to the pixel in the world coordinate system, the specific method includes:
[0092] Randomly select an image block on each two-dimensional picture, calculate the start coordinate and direction vector of the light ray corresponding to the pixel in the world coordinate system through the pixel of the image block; let the length and width of the image block be p h and p w , then the tensor corresponding to the RGB color of the image block pixel The start coordinate tensor of the light ray corresponding to the image block pixel The direction vector tensor
[0093] Step 2-2: Use uniform sampling to sample on the light ray to generate initial sampling points, and calculate the coordinates of the initial sampling points in the world coordinate system;
[0094] In step 2-2, the uniform sampling is calculated by the following formula:
[0095]
[0096] Where t n is the nearest distance of the sampling point from the camera, t f is the farthest distance of the sampling point from the camera, t i is the distance of the i-th sampling point from the camera obtained by the uniform sampling method, N init is the number of initial sampling points, U represents uniform distribution, and ~ represents subject to a certain probability distribution.
[0097] Step 3: Construct a neural radiance field hybrid model, including the following steps:
[0098] Step 3-1, construct a neural radiance field hybrid model with two granularities, namely a coarse model and a fine model; wherein the coarse model includes a coordinate feature encoder composed of eight two-dimensional convolution layers with a convolution kernel of 1*1, a volume density decoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1, a direction vector feature encoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1, and a color decoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1,
[0099] Step 3-2, construct a fine model, including a coordinate feature encoder composed of eight two-dimensional convolution layers with a convolution kernel of 1*1, a volume density decoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1, a direction vector feature encoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1, a color decoder composed of one two-dimensional convolution layer with a convolution kernel of 1*1, and a hybrid module; wherein the hybrid module includes: a cross-ray cross-sampling-point hybrid module and a same-ray cross-sampling-point hybrid module;
[0100] Step 3-3, constructing a cross-ray cross-sampling point mixing module composed of a volume density mixing sub-network and a color mixing sub-network, wherein the volume density sub-network is composed of two two-dimensional convolution layers with a k*k convolution kernel, which are used to predict the volume density weight W_inter of the adjacent sampling points on different rays to the sampling point at the same depth as the sampling point σ ; the color sub-network is composed of two two-dimensional convolution layers with a k*k convolution kernel, which are used to predict the color weight W_inter of the adjacent sampling points on different rays to the sampling point at the same depth as the sampling point c ;
[0101] Step 3-4, constructing an intra-ray cross-sampling point mixing module composed of a volume density mixing sub-network and a color mixing sub-network, wherein the volume density sub-network is composed of two one-dimensional convolution layers with a k convolution kernel, which are used to predict the volume density weight W_intra of the adjacent sampling points on the same ray to the sampling point σ ; the color sub-network is composed of two one-dimensional convolution layers with a k convolution kernel, which are used to predict the color weight W_intra of the adjacent sampling points on the same ray to the sampling point c ;
[0102] Step 4: input the coordinates and direction vectors of the initial sampling points into the neural radiance field mixing model in step 3 to calculate the volume density and color of each sampling point, as shown in Figure 2 , which is divided into the following steps:
[0103] Step 4-1: obtain the initial sampling points according to the uniform sampling in step 2-2, and encode the coordinates and direction vectors of the sampling points according to the coordinate information and direction vector information of the sampling points using the following encoding formula:
[0104] γ(p)=(sin(2 0 πp),cos(2 0 πp),…,sin(2 L-1 πp),cos(2 L-1 πp))
[0105] wherein p is the coordinate of the sampling point or the direction vector of the ray on which the sampling point is located, L is the encoding frequency, and γ(p) represents the encoded encoding feature, wherein the coordinate encoding obtains a coordinate feature, and the direction vector encoding obtains a direction vector feature; the coordinate feature tensor is wherein D x =3×(2×L x +1); the direction vector feature tensor is wherein D v =3×(2×L v +1), and N is the number of sampling points on each ray;
[0106] Step 4-2: Convert the coordinate feature tensor from Step 4-1 With direction vector feature tensor Input the coarse model from the neural radiation field mixing model in step 3 to obtain the volume density σ of the initial sampling points. coarse and color c coarse This step is as follows Figure 3 As shown;
[0107] Step 4-3: Use a hierarchical sampling method based on the volume density σ of the sampling points in the coarse model. coarse Resampling yields finer sampling points, achieving the purpose of importance sampling, as detailed below:
[0108] Step 4-3-1, calculate the volume density σ of the initial sampling points on each optical path in step 4-2. coarse The specific method for normalization is as follows:
[0109]
[0110] Where, p i N represents the probability that the i-th sampling point in the current scene is located on the surface of an object. init This represents the initial number of sampling points on each ray;
[0111] Step 4-3-2: Based on the probability distribution in step 4-3-1, generate hierarchical sampling points through inverse transformation sampling;
[0112] Step 4-3-3: Merge the initial sampling points in Step 2-2 with the hierarchical sampling points in Step 4-3-2 into the final fine sampling points;
[0113] Step 4-4: Encode the coordinates and direction vectors of the fine sampling points using the encoding method in Step 4-1 to obtain the coordinate feature tensor. With direction vector feature tensor N fine The number of fine sampling points is given, and the fine model constructed in step 3-2 is input, where the volume density feature F is obtained through the coordinate feature encoder. σ The direction vector feature encoder obtains the color feature F. c The volume density decoder obtains the volume density σ fine The volume density decoder obtains color c. Fine This step is as follows Figure 4 As shown;
[0114] Step 4-5: Convert the volume density feature F from step 4-4 into... σ With color feature F cThe cross-ray cross-sample point mixing module in step 3-3 is inputted. For each sample point, the module predicts the body density weight W_inter σ and color weight W_inter c of the neighboring sample points on different rays to the sample point at the same depth as the sample point. Finally, the body density σ fine and color c fine obtained in step 4-4 are respectively weighted and summed to obtain the cross-ray cross-sample point mixed body density σ inter_mixed and color c inter_mixed , as shown in FIG. 4-5. The formula for weighted sum is as follows: Figure 4
[0115]
[0116] where k is the number of neighboring sample points, is the body density weight of the i-th neighboring sample point to the sample point, is the body density of the i-th neighboring sample point, is the body color weight of the i-th neighboring sample point to the sample point, is the color of the i-th neighboring sample point.
[0117] Step 4-6: The body density feature F σ and color feature F c in step 4-4 are inputted into the same-ray cross-sample point mixing module in step 3-4. For each sample point, the module predicts the body density weight W_intra σ and color weight W_intra c of the neighboring sample points on the same ray to the sample point. Finally, the cross-ray cross-sample point mixed body density σ inter_mixed and color c inter_mixed obtained in step 4-4 are respectively weighted and summed to obtain the cross-ray and same-ray mixed body density σ intra_inter_mixed and color c intra_inter_mixed , as shown in FIG. 4-7. The formula for weighted sum is as follows: Figure 4
[0118]
[0119] where k is the number of neighboring sample points, is the body density weight of the i-th neighboring sample point to the sample point, is the cross-ray cross-sample point mixed body density of the i-th neighboring sample point, is the body color weight of the i-th neighboring sample point to the sample point, Ci is the color of the i-th sample point.
[0120] Step 5: Calculate the RGB color value of each pixel corresponding to the light ray through the volume rendering equation.
[0121] The volume rendering equation formula is as follows:
[0122]
[0123] wherein, represents the final RGB color value of the pixel corresponding to the light ray, σ i and c i are the volume density and color of the i-th sample point, δ j = t j+1 -t j is the distance between adjacent sample points on the same light ray, T i is the cumulative transmittance from the camera position to the i-th sample point.
[0124] Step 6: Train the neural radiance field mixing model in step 3; update the model parameters using the Adam optimizer by calculating the loss, wherein the calculation of the loss can be divided into the following steps:
[0125] Step 6-1: For the volume density σ coarse and color c coarse obtained by the coarse model in step 4-2, calculate the pixel RGB value rendered by the radiance field in the coarse model through the method in step 5.
[0126] Step 6-2: For the volume density σ fine and color c fine obtained by the fine model in step 4-4, calculate the pixel RGB value rendered by the radiance field in the fine model through the method in step 5.
[0127] Step 6-3: For the volume density σ inter_mixed and color c inter_mixed mixed by the cross-ray cross-sample point mixing module in step 4-5, calculate the pixel RGB value rendered by the fine model after mixing through the cross-ray cross-sample point mixing module through the method in step 5.
[0128] Step 6-4: For the volume density σ intra_nter_mixed and color c intra_inter_mixed mixed by the same-ray cross-sample point mixing module in step 4-6, calculate the pixel RGB value rendered by the fine model after mixing through the same-ray cross-sample point mixing module through the method in step 5.
[0129] Step 6-5: For the neural radiance field hybrid model defined in step 3, the loss is calculated by the following formula:
[0130] L RF =L coarse +L fine
[0131]
[0132] Wherein, L RF is the total loss of the neural radiance field hybrid model, R is the set of light rays, L coatse and L fine are the losses of the coarse model and the fine model respectively, is the true RGB value of the corresponding pixel of the image block pixel. Step 7: According to the camera pose of the new view angle, render the two-dimensional picture under this view angle.
[0133] Step 7 only involves the forward rendering process of the model, given the camera pose under a specific view angle and the length and width of the synthesized picture, the final synthesized RGB picture is obtained according to steps 2-7.
[0134] Neural radiance field is a new view image generation method with excellent effect that has emerged in recent years. This technology is based on deep learning, which implicitly models a three-dimensional scene through a neural network, and realizes high-fidelity image synthesis through volume rendering technology. Traditional neural radiance field methods often base on the assumption that spatial sampling points and light rays are independent, which limits the performance of synthesized images in color coherence and spatial details. Therefore, the present application proposes a general controllable radiance field enhancement method based on spatial perception, which reorganizes the sampling points and dynamically generates combination weights for the reorganized elements, which can efficiently extract and fuse feature information. The present application can be generally deployed in most neural radiance field-based new view generation methods, and can greatly improve the quality of synthesized images without increasing excessive memory and computing overhead.
[0135] Embodiment:
[0136] Step 1, obtaining training data;
[0137] Step 2, generating light rays and sampling points according to the training data;
[0138] Step 3, constructing a neural radiance field hybrid model;
[0139] Step 4, inputting the coordinates of the sampling points and the directions of the light rays into the neural radiance field hybrid model constructed in step 3 to calculate the volume density and color of each sampling point;
[0140] Step 5, calculate the RGB color value of the pixel corresponding to each light ray;
[0141] Step 6, train the neural radiance field mixing model constructed in step 3 according to the volume density and color of each sampling point obtained in steps 4 and 5, and the RGB color value corresponding to each light ray, to obtain a trained neural radiance field mixing model;
[0142] Step 7, using the trained neural radiance field mixing model, render a two-dimensional picture under a new view angle according to the camera pose of the view angle, to complete the general controllable spatial perception based radiance field enhancement.
[0143] In step 1, taking the classroom scene in the open source dataset DONeRF as an example, the dataset contains two-dimensional pictures of a classroom with fine texture features under different view angles, and the camera pose corresponding to each two-dimensional picture. The length and width of the two-dimensional picture of the classroom scene are both 800 pixels.
[0144] In step 2, generate light rays and sampling points, calculate the starting point coordinates and direction vectors of the light rays according to the camera pose, and sample the coordinates of the sampling points on the light rays, which specifically includes:
[0145] Step 2-1: randomly select an image block on each two-dimensional picture, calculate the starting point coordinates and direction vectors of the light rays corresponding to the pixels in the image block in the world coordinate system; in the dataset of the classroom scene, the length and width of each two-dimensional picture are both 800 pixels, and an image block is randomly selected on each picture; in the experiment, the length and width of the image block are both 40 pixels, so the RGB color corresponding to the pixels in the image block is a tensor C gt ∈R 3×40×40 , the starting point coordinates tensor O of the light rays corresponding to the pixels in the image block ∈R 3×40×40 , and the direction vector tensor Dir ∈R 3×40×40 .
[0146] Step 2-2, use a uniform sampling method to sample the initial sampling points on the light rays and calculate the coordinates of the initial sampling points in the world coordinate system;
[0147] The uniform sampling method in step 2-2 is as follows:
[0148]
[0149] Where t n is the nearest distance of the sampling point from the camera, t f is the farthest distance of the sampling point from the camera, t i is the distance of the i-th sampling point from the camera obtained by the uniform sampling method, N initFor the number of sampling points, U represents uniform distribution, ~ represents subject to a certain probability distribution; in the data set of the classroom scene, t n is 0.64, t f is 8.80, N init is 64.
[0150] Step 3, construct a neural radiance field hybrid model.
[0151] Step 4, calculate the volume density and color of each sampling point, including the following steps:
[0152] Step 4-1, coordinate and direction vector encoding is performed on the 64 initial sampling points generated in step 2-2 to obtain coordinate features and direction vector features, and the specific encoding method is as follows:
[0153] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp))
[0154] Where p is the coordinate of the sampling point or the direction vector of the light ray it is in, L is the encoding frequency, and γ(p) represents the encoded encoding features. In the experiment, the coordinate encoding frequency is 10, the direction vector encoding frequency is 4, the coordinate encoding obtains the coordinate feature tensor X init ∈R 64 × 63 × 40 × 40 , and the direction vector encoding obtains the direction vector feature tensor V init ∈R 64 × 27 × 40 × 40 ;
[0155] Step 4-2, input the coordinate feature tensor X init and the direction vector feature tensor V init of the initial sampling points in step 4-1 into the coarse model in step 3 to obtain the volume density σ coarse ∈R 64×1×40×40 and the color c coarse ∈R 64×3×40×40 of the initial sampling points in the coarse model;
[0156] Step 4-3, using the hierarchical sampling method, based on the volume density σ coarse of the sampling points in the coarse model, the hierarchical sampling points are sampled again and combined with the initial sampling points to obtain the final fine sampling points;
[0157] Step 4-4, the coordinates and direction vector of the fine sampling points are encoded using the encoding method in step 4-1, and the body density σ of the fine sampling points is input into the fine model to obtain the body density σ of the fine sampling points fine ∈R 192×1×40×40 , color c fine ∈R 192×3×40×40 , the body density σ mixed across light rays and sampling points inter_mixed R 192×1×40×40 , the color c mixed across light rays and sampling points inter_mixed ∈R 192 ×3×40×40 , the body density σ mixed across light rays and the same light rays intra_inter_mixed R 192×1×40×40 , the color c mixed across light rays and the same light rays intra_inter_mixed ∈R 192×3×40×40 , wherein the number of fine sampling points is 192;
[0158] Step 5, calculate the RGB color value corresponding to each light ray, the specific method comprising:
[0159]
[0160] wherein, represents the final RGB color value of the pixel corresponding to the light ray, σ i and c i are the body density and color of the i-th sampling point, δ j is the distance between adjacent sampling points on the same light ray, T i is the cumulative transmittance from the camera position to the i-th sampling point, and N is the number of sampling points on the light ray;
[0161] Step 6, train the neural radiance field mixing model constructed in step 3, i.e., update the parameters of the neural radiance field mixing model using the Adam optimizer by calculating the loss, wherein the method for calculating the loss comprises:
[0162] Step 6-1, for the body density σ coarse and the color c coarse obtained by the coarse model in step 4-3-2, calculate the pixel RGB value rendered by the radiance field in the coarse model through the method in step 5
[0163] Step 6-2, for the body density σ fine obtained by the fine model in step 4-3-3 and the color c fine obtained by the fine model in step 4-3-4, calculate the pixel RGB value rendered by the radiance field in the fine model through the method in step 5
[0164] Step 6-3, for the volume density sigma mixed by the cross-ray-cross-sampling point mixing module in the fine model in step 4-3-5 inter_mixed and color c inter_mixed , the pixel RGB value rendered by the method in step 5 after the mixed volume density sigma in the fine model is mixed by the cross-ray-cross-sampling point mixing module
[0165] Step 6-4, for the volume density sigma mixed by the same-ray-cross-sampling point mixing module in step 4-3-6 intra_nter_mixed and color c intra_inter_mixed , the pixel RGB value rendered by the method in step 5 after the mixed volume density sigma in the fine model is mixed by the same-ray-cross-sampling point mixing module
[0166] Step 6-5, calculate the loss, as follows:
[0167] L RF = L coarse + L fine
[0168]
[0169] Wherein, L RF is the total loss of the neural radiance field mixing model, R is a set of rays, L coarse and L fine are the losses of the coarse model and the fine model respectively, C gt ∈R 3 × 40 × 40 is the true RGB value of the pixel corresponding to the image block pixel;
[0170] The general controllable spatial perception based radiance field enhancement method provided by the application solves the limitation of the basic assumption that the sampling points are independent of each other in the neural radiance field method, which leads to the lack of fine texture in the synthesized picture. The mixing strategy provided by the application significantly improves the quality of the synthesized picture under the premise of only increasing about 25% of the rendering time, and the application provides a general radiance field enhancement module which can be integrated into various neural radiance field based methods to further improve the quality of the synthesized picture.
[0171] Table 1 experimental results of the classroom scene data set
[0172]
[0173] In the table, the evaluation indexes of the quality of the synthesized picture include a peak signal-to-noise ratio, a structural similarity index and a perceived image block similarity, the larger the value is, the higher the quality of the picture is, the smaller the value is, the higher the quality of the picture is. It can be obtained from the data in the table that the present application is superior to the original neural radiance field method in the three indexes. In addition, the time for rendering a picture by the method provided by the present application is increased from 28.235 seconds to 35.088 seconds by the original neural radiance field method, and the picture quality can be significantly improved under the premise of only increasing 24.24% of the time.
[0174] The present application provides a general controllable space perception based radiance field enhancement method, and there are many methods and approaches to realize the technical scheme, and the above description is only the preferred embodiment of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principle of the present application, some improvements and refinements can be made, and these improvements and refinements should be regarded as the protection scope of the present application. The components not explicitly described in the embodiment can be realized by the existing technology.
Claims
1. A general and controllable method for enhancing radiation fields based on spatial sensing, characterized in that, Includes the following steps: Step 1, Obtain training data; Step 2: Generate light rays and sampling points based on the training data; Step 3: Construct a neural radiation field hybrid model; Step 4: Input the coordinates of the sampling points and the direction of the light rays into the neural radiation field mixing model constructed in Step 3, and calculate the volume density and color of each sampling point; Step 5: Calculate the RGB color value of the pixel corresponding to each ray; Step 6: Based on the volume density and color of each sampling point obtained in Steps 4 and 5, and the RGB color value corresponding to each ray, train the neural radiation field mixing model constructed in Step 3 to obtain the trained neural radiation field mixing model. Step 7: Using the trained neural radiation field mixing model, render a 2D image from the new camera pose to complete a general and controllable radiation field enhancement based on spatial awareness. The construction of the neural radiation field hybrid model in step 3 includes the following steps: Step 3-1: Construct two granularity models under the neural radiation field hybrid model: a coarse model and a fine model. The coarse model includes: a sampling point coordinate feature encoder, a direction vector feature encoder, a volume density decoder, and a color decoder. The output of the coarse model is the volume density σ of the sampling points. coarse and color c coarse ; The detailed model includes: a sampling point coordinate feature encoder, a direction vector feature encoder, a volume density decoder, a color decoder, and a mixing module; wherein, the mixing module includes: a cross-ray cross-sampling point mixing module and a same-ray cross-sampling point mixing module; the output of the detailed model is the volume density feature F of the sampling points. σ Color feature F c Volume density σ fine Color c fine The volume density σ after mixing across light rays and sampling points inter_mixed The color c after mixing across light rays and sampling points inter_mixed The volume density σ after both trans-ray and in-ray are mixed intra_inter_mixed The color c after the trans-ray and the same-ray are mixed intra_inter_mixed ; Step 3-2: Construct the cross-ray, cross-sampling-point mixing module in the radiation field mixing model; Step 3-3: Construct the same ray cross-sampling point mixing module in the radiation field mixing model; The sampling point coordinate feature encoder described in step 3-1 consists of eight two-dimensional convolutional layers with 1*1 kernels, which encode the input sampling point coordinate features to obtain volume density features. The volume density decoder consists of a single two-dimensional convolutional layer with a 1*1 kernel, which decodes the volume density features to obtain the volume density of each sampling point. The orientation vector feature encoder consists of a two-dimensional convolutional layer with a 1*1 kernel. After concatenating the volume density features with the input orientation vector features, the orientation vector feature encoder is used to encode the color features again. The color decoder consists of a single 2D convolutional layer with a 1*1 kernel, which decodes the color features to obtain the color of each sampling point. The cross-ray cross-sampling point mixing module described in step 3-2 consists of a volume density mixing sub-network and a color mixing sub-network, wherein each sub-network consists of two two-dimensional convolutional layers with a kernel of k*k. The cross-ray cross-sampling-point blending module is used to predict the volume density weights and color weights of each sampling point from its neighboring sampling points at the same depth on different rays, thus obtaining the volume density σ of each sampling point after cross-ray cross-sampling. inter_mixed The color c after mixing across light rays and sampling points inter_mixed The specific prediction method is as follows: Step 3-2-1, the volume density feature F obtained from the sampling points in the fine model σ The volume density mixing subnetwork is input to predict the volume density weights W_inter of the sampling point for each adjacent sampling point at the same depth on different light rays. σ For the volume density σ fine By performing a weighted summation, the volume density σ after mixing across rays and sampling points is obtained. inter_mixed The specific method is as follows: Where k is the number of adjacent sampling points. The volume density weight of the i-th adjacent sampling point on that sampling point. The volume density of the i-th adjacent sampling point; Step 3-2-2: Calculate the color features F obtained from the sampling points in the fine model. c The color mixing subnetwork is input to predict the color weights W_inter of the sampling point for each neighboring sampling point at the same depth on different light rays. c For the color c fine By performing a weighted summation, we obtain the color c after mixing across rays and sampling points. inter_mixed The specific method is as follows: Where k is the number of adjacent sampling points. The volume color weight of the i-th adjacent sampling point to that sampling point. The color of the i-th adjacent sampling point; The same light cross-sampling point mixing module described in step 3-3 consists of a volume density mixing sub-network and a color mixing sub-network, wherein each sub-network consists of two one-dimensional convolutional layers with a kernel of k. The same-ray cross-sampling-point mixing module is used to predict the volume density weight and color weight of each sampling point from its adjacent sampling points on the same ray, and to obtain the volume density σ of each sampling point after both cross-ray and same-ray mixing. cintra_inter_mixed The color c after the trans-ray and the same-ray are mixed intra_inter_mixed The specific prediction method is as follows: Step 3-3-1, the volume density feature F obtained from the sampling points in the fine model σ The volume density mixing subnetwork is input to predict the volume density weight W_intra of the adjacent sampling points on the same ray for that sampling point. σ The volume density σ after mixing across light rays and sampling points in step 3-2 inter_mixed By performing a weighted summation, we obtain the volume density σ after mixing both the trans-ray and in-ray rays. intra_inter_mixed The specific method is as follows: Where k is the number of adjacent sampling points. The volume density weight of the i-th adjacent sampling point on that sampling point. The volume density of the i-th adjacent sampling point after cross-ray and cross-sampling point mixing; Step 3-3-2, the color features F obtained from the sampling points in the fine model c The color mixing subnetwork is input to predict the color weight W_intra of the adjacent sampling points on the same ray for that sampling point. c The color c after mixing across light rays and sampling points in step 3-2 inter_mixed By performing a weighted summation, we obtain the color c after mixing both the cross-ray and the same-ray. intra_inter_mixed The specific method is as follows: Where k is the number of adjacent sampling points. The volume color weight of the i-th adjacent sampling point to that sampling point. The color of the i-th adjacent sampling point is the result of cross-ray and cross-sampling point mixing.
2. The universal and controllable spatially-aware radiation field enhancement method according to claim 1, characterized in that, The step 1 involves acquiring training data by taking pictures with a camera to obtain two-dimensional images from different perspectives and recording the camera pose corresponding to each image. The two-dimensional images and their corresponding camera poses are then used as training data.
3. The universal and controllable spatially-aware radiation field enhancement method according to claim 2, characterized in that, Step 2, which involves generating rays and sampling points, uses neural radiation fields to model the rays. Each pixel in the 2D image emits a ray, with the starting point of the ray being the camera position and the direction being a direction vector from the camera position to the pixel position in the 2D image plane. Each pixel corresponds to one ray. The starting coordinates and direction vector of the ray are calculated based on the camera pose, and the coordinates of the sampling points are obtained by sampling along the ray. Specifically, this includes: Step 2-1: Based on the camera pose and pixel coordinates in the 2D image plane, calculate and generate the starting coordinates and direction vector of the ray corresponding to that pixel in the world coordinate system. Specific methods include: Randomly select an image patch from each 2D image, and calculate the origin coordinates and direction vector of the ray corresponding to each pixel in the world coordinate system; let the length and width of the image patch be p. h and p w Then the tensor corresponding to the RGB color of the image block pixels The tensor of the origin coordinates of the rays corresponding to the pixels of the image patch Direction vector tensor Step 2-2: Use the uniform sampling method to sample the light rays to generate initial sampling points, and calculate the coordinates of the initial sampling points in the world coordinate system; The uniform sampling method described in step 2-2 is as follows: Among them, t n t represents the closest distance between the sampling point and the camera. f t represents the farthest distance between the sampling point and the camera. i Let N be the distance from the camera to the i-th sampling point obtained by the uniform sampling method. init Let U be the number of initial sampling points, U represents a uniform distribution, and ~ represents a certain probability distribution.
4. The universal and controllable space-aware radiation field enhancement method according to claim 3, characterized in that, Step 4, which involves calculating the volume density and color of each sampling point, includes the following steps: Step 4-1: Encode the sampling points using coordinate and direction vectors to obtain coordinate features and direction vector features. The specific encoding method is as follows: γ(p)=(sin(2 0 πp),cos(2 0 πp),…,sin(2 L-1 πp),cos(2 L-1 πp)) Where p is the coordinate of the sampling point or the direction vector of the ray it contains, L is the encoding frequency, and γ(p) represents the encoded feature, where coordinate features are obtained after coordinate encoding, and direction vector features are obtained after direction vector encoding; the coordinate feature tensor is... Where D x =3×(2×L) x +1); the characteristic tensor of the direction vector is Where D v =3×(2×L) v +1), where N is the number of sampling points on each ray; Step 4-2, convert the coordinate feature tensor of the initial sampling points mentioned in Step 2-2. With direction vector feature tensor Input the coarse model from step 3 to obtain the volume density of the initial sampling points in the coarse model. and color N init This is the initial number of sampling points; Step 4-3: Use a hierarchical sampling method based on the volume density σ of the sampling points in the coarse model. coarse Resampling yields finer sampling points, as detailed below: Step 4-3-1: Calculate the volume density σ of the sampling points on each ray in the coarse model. coarse Normalize it and use it as the probability p that the sampling point is located on the surface of an object in the current scene. i The specific method is as follows: Where, N init This represents the initial number of sampling points on each ray; Step 4-3-2: Based on the probability p of each sampling point on each ray being located on the surface of an object in the current scene. i The distribution is used to generate hierarchical sampling points through inverse transformation sampling. Step 4-3-3: Merge the initial sampling points described in step 2-2 and the hierarchical sampling points described in step 4-3-2 into the final fine sampling points; Step 4-4: Encode the coordinates and direction vectors of the fine sampling points using the encoding method in Step 4-1 to obtain the coordinate feature tensor. With direction vector feature tensor Then input the fine model constructed in step 3 to obtain the volume density features of the fine sampling points. Color characteristics Volume density color Volume density after cross-ray cross-sampling point mixing Color after mixing across light rays and sampling points Volume density after mixing of trans-rays and in-rays The color after mixing with both trans-rays and same-rays N fine To determine the number of fine sampling points, the specific steps are as follows: Step 4-4-1: Obtain the volume density feature F in the fine model using the coordinate feature encoder in the fine model. σ ; Step 4-4-2: Obtain the color feature F in the fine model using the direction vector feature encoder in the fine model. c ; Step 4-4-3: Obtain the volume density σ in the fine model using the volume density decoder in the fine model. fine ; Step 4-4-4: Obtain the color c in the fine model using the color decoder in the fine model. fine ; Step 4-4-5, apply the above volume density feature F σ Color feature F c Volume density σ fine and color c fine Input the cross-ray cross-sampling point mixing module to obtain the volume density σ after cross-ray cross-sampling point mixing. inter_mixed and color c inter_mixed ; Step 4-4-6, apply the above volume density feature F σ Color feature F c The volume density σ after mixing across light rays and sampling points inter_mixed The color c after mixing across light rays and sampling points inter_mixed Input the same ray across sampling point mixing module to obtain the volume density σ after mixing both the cross-ray and same ray. intra_inter_mixed With color c intra_inter_mixed .
5. A universal and controllable spatially-aware radiation field enhancement method according to claim 4, characterized in that, The specific method for calculating the RGB color value of the pixel corresponding to each ray in step 5 includes: in, σ represents the final RGB color value of the pixel corresponding to the light ray. i and c i Let δ be the volume density and color of the i-th sampling point. j =t j+1 -t j T is the distance between adjacent sampling points on the same ray. i Let be the cumulative penetration rate from the camera position to the i-th sampling point.
6. The universal and controllable spatially-aware radiation field enhancement method according to claim 5, characterized in that, Step 6 describes training the neural radiation field mixture model constructed in step 3, i.e., updating the parameters of the neural radiation field mixture model using the Adam optimizer by calculating the loss. The methods for calculating the loss include: Step 6-1, for the volume density σ obtained from the coarse model in Step 4-2 coarse and color c coarse The pixel RGB values obtained by rendering the radiation field in the coarse model are calculated using the method in step 5. Step 6-2, for the volume density σ obtained from the fine model in step 4-2 fine and color c fine The pixel RGB values obtained by rendering the radiation field in the fine model are calculated using the method in step 5. Step 6-3, for the volume density σ of the fine model in step 4-4, which is mixed through the cross-ray, cross-sampling point mixing module. inter_mixed and color c inter_mixed The pixel RGB values obtained after blending and rendering through the cross-ray and cross-sampling point blending module in the fine model are calculated using the method in step 5. Step 6-4, for the volume density σ mixed by the same light beam across sampling points mixing module in step 4-5 intra_nter_mixed and color c intra_inter_mixed The pixel RGB values obtained after rendering by blending through the same light cross-sampling point blending module in the fine model are calculated using the method in step 5. Step 6-5, calculate the loss, as follows: L RF L coarse +L fine Among them, L RF This is the total loss of the neural radiation field mixture model, where R is the set of light rays, and L is the total loss. coarse and L fine These are the losses for the coarse model and the fine model, respectively. It is the actual RGB value of the pixel corresponding to the pixel of the image block.
Citation Information
Patent Citations
Scene new view generation method based on local space aggregation neural radiation field
CN116993826A
Neural radiation field-based structure three-dimensional model updating method, apparatus and device, and medium
CN118365817A