NeRF multi-model scene construction method and device based on twisted light and medium

By using density voxel grid optimization hierarchical sampling and distorted ray methods in NeRF, the problems of low training and rendering efficiency and poor flexibility in multi-model combination are solved, and efficient and flexible scene construction and editing are achieved.

CN119919563AActive Publication Date: 2025-05-02NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510397116.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-02
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the prior art, the training and rendering efficiency of neural radiation field (NeRF) is low, and the flexibility of multi-model combination is poor, making it difficult to meet the needs of real-time or high-efficiency scene editing.

Method used

The NeRF multi-model scenario construction method based on distorted light is adopted, and the hierarchical sampling strategy of the neural radiation field is optimized through the density voxel grid, and the spherical harmonic neural radiation field model based on the density voxel grid is trained, and the spherical harmonic coefficients and volume density of the contributing sampling points are baked into the hash table to reduce the dependence on the neural network.

Benefits of technology

It significantly reduces the training time of NeRF, improves the training and rendering efficiency of neural radiation fields, realizes the rendering speed at the user interaction level, and realizes the free combination of multiple models through the distorted light method, improving the efficiency and flexibility of scene construction and editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919563A_ABST
    Figure CN119919563A_ABST
Patent Text Reader

Abstract

The invention discloses a NeRF multi-model scene construction method and device based on distorted light, and a medium, and the method comprises the following steps: obtaining a training data set, carrying out the preprocessing of the data set, and calculating the pose of an image; constructing a GPU environment, setting training parameters, and loading a scene configuration file and a two-dimensional image with a camera pose; optimizing a neural radiation field stratified sampling strategy by using a density voxel grid, efficiently training a spherical harmonic neural radiation field (NeRF-SH) model, and recording network output by using a hash table; combining and rendering the plurality of neural radiation field models by using a light distortion method to obtain a combined scene image; models are increased or decreased, various parameters are adjusted, and scene construction is achieved through rendering. According to the method, the stratified sampling strategy in the neural radiation field is optimized, the neural radiation field is baked into the hash table, the training and rendering efficiency of the neural radiation field is remarkably improved, multiple models are combined by twisting light and combining the maximum combination strategy, and construction and editing of a scene are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a NeRF multi-model scene construction method, device and medium based on twisted light. Background Art

[0002] In the fields of computer vision, virtual reality, augmented reality, and film and television production, high-quality scene construction is one of the key technologies to achieve realistic visual effects. Whether it is building a virtual environment, generating special effects, or reconstructing a scene, an efficient and accurate scene construction method is required. Traditional scene construction techniques usually rely on manual modeling or physically based rendering methods. Although these methods can generate high-quality visual effects, the process is time-consuming and labor-intensive, and it is difficult to handle the details of complex scenes. Therefore, it is of great practical significance to develop an efficient and automated scene construction method.

[0003] Traditional scene construction methods can be divided into two categories: methods based on manual modeling and methods based on image reconstruction. The methods based on manual modeling rely on designers to manually create three-dimensional models and generate scenes through techniques such as texture mapping and lighting simulation. Although this method can generate high-quality visual effects, its process is extremely time-consuming and requires high technical skills from designers. In addition, manual modeling is difficult to handle details in complex scenes (such as natural landscapes, complex object surfaces, etc.), which limits its scope of application. The methods based on image reconstruction reconstruct three-dimensional scenes through multi-view images, such as technologies based on structured light, laser scanning or stereo vision. Although these methods have achieved automation to a certain extent, their hardware cost is high and they are sensitive to environmental factors such as lighting conditions and object materials, making it difficult to achieve high-precision reconstruction in complex scenes. In summary, traditional scene construction methods have problems such as low efficiency, high cost, and poor adaptability, and it is difficult to meet the needs of modern computer vision and graphics fields for efficient and high-quality scene construction.

[0004] In recent years, the emergence of Neural Radiance Fields (Nerf: Representing scenes as neural radiance fields for view synthesis. ECCV, 65(1): 99–106, 2021.) technology has brought revolutionary breakthroughs in the field of scene construction. Neural Radiance Fields (NeRF) is a scene representation method based on neural networks. It learns the 3D geometry and appearance information of scenes from multi-view 2D images and can generate high-quality new-view images. Compared with traditional scene construction methods, NeRF has the advantages of high-quality rendering, high degree of automation and high flexibility. NeRF can generate realistic new-view images and has good processing capabilities for the details of complex scenes (such as transparent objects, reflective surfaces, etc.). NeRF only needs to input multi-view 2D images, without manual modeling or complex hardware equipment, which significantly reduces the cost and difficulty of scene construction. In addition, NeRF can be applied to a variety of scenes, including indoor and outdoor environments, natural landscapes, dynamic scenes, etc., and has broad application prospects. Therefore, the emergence of NeRF technology provides a new solution for efficient and high-quality scene construction.

[0005] Although neural radiance field technology has achieved remarkable results in scene representation and new perspective synthesis, the high computational complexity of its training and rendering process limits its practical application. In order to improve the efficiency of NeRF, researchers have proposed a variety of acceleration methods. Yu et al. (Yu A, Li R, Tancik M, et al. Plenoctrees for real-time rendering of neural radiance fields. In ICCV, pages 5752-5761, 2021.) proposed PlenOctree, which pre-samples and stores the trained NeRF-SH into an octree structure, eliminating the need to run the neural network during the test phase, thereby achieving several orders of magnitude acceleration in rendering. In addition, Hedman et al. (Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, et al. Baking neural radiance fields for real-time view synthesis. In ICCV, pages 5875-5884, 2021.) proposed Baking-NeRF, which uses sparse voxel grids to store neural radiation field information, eliminating the need to perform complex neural network calculations during rendering. Although these methods maintain high rendering quality while accelerating NeRF rendering, they still face long training times during the training phase.

[0006] The patent document with the publication number CN116958367A discloses a method for quickly combining and rendering complex neural scenes, the method comprising: including: according to the collected 2D images of multiple objects under different perspectives, training the neural radiation field model corresponding to each object and the neural depth field model corresponding to the neural radiation field model, preparing digital resources; according to the neural radiation field model and the neural depth field model corresponding to each object contained in the current scene, quickly rendering the result of the combined virtual scene. Using this method, the user can realize the progressive object insertion and interactive-level object operation functions, thereby providing a fast result preview for the target virtual scene and assisting the construction of the neural scene. However, this method has the following defects: first, it adopts the basic NeRF training method, and the basic NeRF usually takes dozens of hours or even days to train, and the training efficiency is extremely low; second, after the NeRF training is completed, the method needs to further train the neural depth field model (NeRD) corresponding to each object according to the rendering results of NeRF. This process not only adds additional computational overhead, but also further prolongs the time of overall scene construction, which is difficult to meet the needs of real-time or efficient scene editing, and the additional training of auxiliary fields will greatly affect the flexibility of scene combination.

[0007] The patent document with publication number CN116363299A discloses an explicit scene combination method based on neural radiation field, which includes: collecting multi-viewpoint data of source scene and target scene; using the collected multi-viewpoint data to pre-train the source neural radiation field and the target neural radiation field; using the trained source neural radiation field to generate multi-viewpoint data of arbitrary perspective; jointly training the occlusion field with the pre-trained source neural radiation field; using the combined rendering equation to combine the source neural radiation field and the target neural radiation field and render the image of the combined scene; using the trained occlusion field to realize scene editing. The invention uses the occlusion field to combine two scenes represented by neural radiation fields in a self-supervised mode, and renders a realistic picture of the combined scene with natural object edge transitions at any perspective, and also has the function of scene editing. However, this method has the following defects: First, this method and the method for quickly combining and rendering complex neural scenes disclosed in the patent document with publication number CN116958367A both use the basic NeRF training method, and the training efficiency is extremely low; second, after the source neural radiation field training is completed, the method also needs to use the multi-viewpoint images generated by the source neural radiation field to supervise the joint training of the occlusion field and the source neural radiation field, which will further extend the time of the overall scene construction and affect the flexibility of the scene combination; finally, this method can only combine two NeRF models, which limits the performance of this method in scene construction. Summary of the invention

[0008] The purpose of the present invention is to address the shortcomings and limitations of the prior art and thus provide a NeRF multi-model scene construction method based on twisted rays to solve the problems of low efficiency in neural radiation field training and rendering and poor flexibility in multi-model combination in the prior art.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is: a NeRF multi-model scene construction method based on twisted light, comprising the following steps: S1, obtain real scene and virtual scene data sets, preprocess the data sets, and use motion recovery structure algorithm to calculate image pose; S2, building a GPU environment for training and rendering neural radiance fields, setting training parameters, and loading a scene configuration file and a set of two-dimensional images with camera poses processed in step S1; S3. Use the density voxel grid to optimize the neural radiation field layered sampling strategy. Use the two-dimensional image and camera pose loaded in step S2 to train the spherical harmonic neural radiation field model based on the density voxel grid, and bake the spherical harmonic coefficients and volume density of the contributing sampling points into the hash table in the last round (epoch) of training. After the training is completed, the hash table can be used to accelerate the neural radiation field rendering; use step S2 to load different data sets, repeat step S3, and train to obtain multiple target combination models; S4, selecting the neural radiation field model trained in step S3, configuring transformation parameters and cropping parameters for each model, and rendering to obtain a combined scene image by distorting light and combining it with a maximum combination strategy; S5. Modify the transformation parameters and cropping parameters of each model in step S4 to achieve scene construction and editing.

[0010] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0011] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0012] Compared with the prior art, the NeRF multi-model scene construction method based on twisted rays of the present invention only needs to train the neural radiation field itself, without the need to additionally train other auxiliary fields (such as neural depth fields or occlusion fields), thereby significantly reducing the training time. By optimizing the layered sampling strategy of the neural radiation field through a density voxel grid, the network calculation amount is greatly reduced, and the training and rendering efficiency of the neural radiation field is improved. Finally, the spherical harmonic coefficients and volume density of the contributing sampling points are baked into a hash table to get rid of the dependence on the neural network during rendering, and achieve user interaction-level rendering speed. In addition, the present invention innovatively proposes a NeRF multi-model combination method and a maximum combination strategy based on twisted rays to achieve the free combination of multiple models. The present invention solves the problems of slow training and slow rendering of neural radiation fields from the root, and provides a flexible and efficient solution for the construction of complex scenes through the method of twisted rays. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is the network structure of NeRF-SH. Figure 3 A flow chart of efficient training of neural radiation fields guided by a density voxel grid of the present invention; Figure 4 A flow chart for efficient rendering of neural radiation fields guided by a density voxel grid of the present invention; Figure 5 This is a schematic diagram of the maximum combination strategy principle of the present invention; Figure 6 It is a flow chart of the multi-model free combination method based on twisted light of the present invention; Figure 7 This is a diagram showing the effect of combining and transforming the model postures in an embodiment of the present invention; Figure 8 This is a diagram showing the effect of combining models and adjusting the viewing angle in an embodiment of the present invention. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0015] like Figure 1 As shown, the present invention provides a NeRF multi-model scene construction method based on twisted light, comprising the following steps: S1. Obtain real scene and virtual scene datasets. For real scene datasets, you can shoot scene videos with your mobile phone and use the frame extraction method to obtain image sets. Then you need to preprocess and remove low-quality images caused by motion blur and images that cannot match feature points, and use the Structure from Motion (SfM) algorithm to calculate the camera pose of each image. For virtual scene datasets, you need to import the model to be rendered into the 3D modeling software, place a virtual camera in the scene, and determine its position, orientation, focal length and other parameters. Next, set the camera to sample uniformly on a sphere with a fixed radius, generate two-dimensional images frame by frame from the set camera trajectory, and record the camera pose corresponding to each image.

[0016] S2. Build a GPU environment for training and rendering neural radiation fields, set training parameters, and load a scene configuration file and a set of two-dimensional images with camera poses. The scene configuration file must specify key parameters such as the dataset path, data type, and downsampling ratio. The remaining training parameters (such as learning rate, batch size, number of iterations, network structure, density voxel grid size, etc.) can be configured separately according to needs, or use the default configuration. Load the written scene configuration file into the training program to complete the environment initialization.

[0017] S3, use the density voxel grid to optimize the neural radiation field layered sampling strategy, remove a large number of invalid sampling points and non-contributing sampling points in the dense uniform sampling of light, these sampling points do not need to be calculated by the neural network. And use the two-dimensional image and camera pose loaded in step S2 to train the spherical harmonic neural radiation field model based on the density voxel grid. In the last epoch of training, start to bake the spherical harmonic coefficients and volume density of the contributing sampling points into the hash table. After the training is completed, the hash table can be used to accelerate the neural radiation field rendering; use step S2 to load different data sets, repeat step S3, and train to obtain multiple target combination models.

[0018] S4. Select the neural radiation field model trained in step S3 as the object in the combined scene. The combined scene can contain one or more models. Configure transformation parameters (such as translation, rotation, scaling, etc.) and cropping parameters (normal vectors and translation constants of 6 cropping planes) for each model to determine its position and posture in the scene. Construct a virtual camera to observe the combined scene, including setting the intrinsic parameter matrix (focal length, principal point coordinates) and extrinsic parameter matrix (camera position, orientation) of the virtual camera. For each model to be combined, distort the light emitted by the virtual camera according to its transformation parameters and combine the maximum combination strategy to fuse the rendering results of multiple models into the neural radiation field of the combined scene. This method can ensure the correct blocking relationship between models in the combined scene and generate a two-dimensional image of the NeRF multi-model combined scene from any perspective.

[0019] S5. By adding or removing models in the target scene and modifying the transformation parameters of each model (such as translation, rotation, scaling, etc.) and then rendering, the scene construction and editing functions can be realized. For example, adjust the translation parameters of the model to change the position of the model in the recombined scene, or modify the rotation parameters of the model to change its orientation. After the model parameters are modified, it is only necessary to re-query the properties of the sampling points on the light lines distorted by the model. No operation is required for other models, and after the model parameters are adjusted, there is no need to train anything else. This method significantly improves the efficiency and flexibility of scene construction and editing.

[0020] The steps, principles and effects of the method of the present invention are described in detail below.

[0021] (1) Figure 2 The network structure of NeRF-SH is shown in Figure 1. is the three-dimensional coordinate in space The result after positional encoding (PE), the positional encoding function is:

[0022] express Applied in In the network structure of NeRF-SH, , and the three-dimensional coordinates The result after position encoding is the same as The concatenated 63-dimensional vector is used as the network input.

[0023] Compared with the network structure of classic NeRF, the network structure of NeRF-SH no longer requires the input of light direction, and the network no longer directly predicts the color RGB value of the sampling point, but predicts the spherical harmonic function coefficients at that position, referred to as spherical harmonic coefficients:

[0024] in, Represents the coordinate position of the neural network in the neural radiation field processing, They represent the degree and order of the spherical harmonics, and Represent the network predicted position The spherical harmonic coefficients and volume density obtained at . Spherical harmonics are a set of standard orthogonal basis functions defined on the sphere. These basis functions can be efficiently represented and reconstructed on the sphere by linear combination. Therefore, spherical harmonics can be calculated according to any viewing direction to restore the color that depends on the viewing direction in NeRF:

[0025] in, Indicates the viewing direction, represents the spherical harmonic coefficients, is the Sigmoid function, Represents the spherical harmonics in the direction The value on .

[0026] The main reason for the low efficiency of neural radiation field training and rendering is that NeRF requires dense sampling on each ray to capture complex geometry and appearance details, and each sampling point needs to be queried in high dimensions through the network, resulting in a very large amount of computation for NeRF.

[0027] To solve this problem, the present invention determines which sampling points in the space are valid and which are invalid, and avoids inputting invalid sampling points into the network, thereby saving the computational consumption occupied by these sampling points.

[0028] To determine whether a sampling point is a valid sampling point, it is necessary to determine the contribution of the sampling point to the final integrated color of the light. According to the volume rendering technology:

[0029] Where N is the number of sampling points on the ray, Represents light Previous sampling points, is the cumulative transmittance, is the volume density, is the distance between adjacent sampling points, is the color. Furthermore, the final integrated color of the light can also be expressed as the weighted sum of the colors of the sampling points on the light:

[0030] in, is the contribution of the sampling point to the final integrated color of the light. It can be seen that when hour, , indicating that this sampling point has no contribution to the final integrated color of the light, and this sampling point can only occupy computing resources in vain through the network. The sampling points are valid sampling points. The sampling point is an invalid sampling point. In addition, if the contribution of the sampling point If it is too small, the contribution of the sampling point to the final integrated color of the light is approximately zero, so The sampling points of are non-contributing sampling points. The sampling points of are contributing sampling points, among which is the contribution threshold.

[0031] In order for NeRF to directly determine the type of sampling points during training and rendering, a density voxel grid is used to record the volume density of the sampling points.

[0032] Furthermore, the specific steps of spherical harmonic neural radiation field training based on density voxel grid optimization layered sampling are as follows: Convert the input multi-view image data and the corresponding camera pose into color and light (including starting point and direction ) is input into the neural radiation field model.

[0033] Load a batch of colors and lights, and sample densely and evenly across the lights points, denoted as , from the size of The density of the voxel grid Search this The volume density corresponding to the sampling points :

[0034] A point in three-dimensional space Voxels in a voxel grid with density The corresponding relationship is as follows:

[0035] in express The coordinates in the world coordinate system, Represents the bounds of the scene in world coordinates, represents the size of the density voxel grid in any dimension, express The coordinates of the corresponding position in the density voxel grid in 3 dimensions.

[0036] remember The sampling point is the effective sampling point , The sampling point is an invalid sampling point.

[0037] Set the spherical harmonic coefficients of invalid sampling points to 0, the volume density to 0, and the valid sampling points to Input to the coarse network Calculate the spherical harmonic coefficients of valid sampling points and body density :

[0038] According to the spherical harmonic coefficients of valid sampling points and invalid sampling points, the spherical harmonic coefficients of P can be obtained: , calculated in the direction of the light Down, Color:

[0039] according to Calculate sampling points Contribution to the final integrated color of the light:

[0040] in, Indicates the light Sampling points for light The final integral color contribution, .

[0041] Coarse network prediction rays The colors are:

[0042] in Respectively represent the first The color of the sampling point and its response to light The contribution of the final integrated color.

[0043] The volume density of effective sampling points calculated by the coarse network Input to the density voxel grid, update the density voxel grid:

[0044] in, Indicates sampling point In a density voxel grid The volume density stored at the location in , is a smoothing factor that updates the volume density recorded in the density voxel grid using an exponential moving average (EMA) to reduce noise and fluctuations in volume density updates.

[0045] remember The sampling points are contributing sampling points , The sampling points of are non-contributory sampling points. Set the spherical harmonic coefficients of the non-contributory sampling points to 0, the volume density to 0, and the contributing sampling points Input to fine network Calculate the spherical harmonic coefficients of contributing sampling points and body density :

[0046] According to the spherical harmonic coefficients of contributing sampling points and non-contributing sampling points, the spherical harmonic coefficients of P can be obtained , calculated in the direction of the light Down, Color:

[0047] Thin network prediction light The colors are:

[0048] A loss function based on image color is used to optimize the neural radiance field:

[0049] in, represents the set of rays for each batch, Represents light The true color, Represents the rough network prediction ray Color, Represents the predicted light rays of the fine network After designing the loss function, the network is trained using stochastic gradient descent and the network parameters are updated.

[0050] When training reaches the last epoch, the spherical harmonic coefficients and volume density of the contributing sampling points are baked into the hash table as follows: An additional new field is opened in the density voxel grid to store the hash value of the sampling point, and all of them are initialized to -1, indicating that the sampling point is not stored in the hash table.

[0051] Calculate contributing sampling points Positions in the density voxel grid and read from them The hash value of .

[0052] If there are contributing sampling points The hash value of Can be found in the hash table, then use Update the hash table with spherical harmonic coefficients and volume density: If the query is not found in the hash table, Insert the spherical harmonic coefficients and volume density into the hash table and record The hash value of is its position in the hash table and updates A hash field corresponding to the position in the density voxel grid. In the specific implementation, the GPU training environment is built, the training parameters are set, and the data configuration file and the model configuration file are loaded, including: the GPU is selected as 3090, the video memory is 24G; the number of network training epochs is set to 20 rounds, and each epoch processes all the pixels in the training set; the density voxel grid size D is set to 256, the volume density field is initialized to 30, the hash field is initialized to -1, and the number of uniform sampling points is Set to 198, contribution threshold is set to 0.0001, the learning rate is set to 0.01, the bath size is set to 2048, and the learning rate reduction factor is set to , after the training is completed, an efficient neural radiation field guided by a dense voxel grid can be obtained.

[0053] Furthermore, after each epoch of training is completed, the model file, training results and verification results can be obtained. The verification results include loss, peak signal-to-noise ratio (PSNR), structural similarity (SSIM), learning perceptual image block similarity (LPIPS), training time, validation set rendering time and other indicator parameters. Among them, PSNR, SSIM and LPIPS are important parameters for evaluating the rendering quality of the model and can be used to verify the feasibility of the present invention. However, please note that all sampling points in the first epoch are considered to be valid sampling points, and only after the last epoch of training is completed, the model with the hash table constructed and stored is an efficient neural radiation field model. Figure 3 Shown is a flow chart of the efficient training of neural radiation fields guided by the density voxel grid of the present invention.

[0054] Efficient rendering of neural radiance fields guided by a dense voxel grid Figure 4 As shown, the rendering process is similar to the training process, except that the reliance on the neural network is eliminated.

[0055] (2) The volume density of sampling points in space can be calculated through the network , and then the cumulative transmittance can be calculated Among them, the volume density can be interpreted as the ray ending at The differential probability at can be interpreted as the ray from the starting point To location The probability that the ray is not terminated. Therefore, the ray at position The probability of termination is:

[0056] in, Indicates that the ray ends at The probability of express small changes in .

[0057] So use represents the probability density of a ray being terminated at position t.

[0058] If different models are to be combined, the blocking relationship between the different models needs to be determined. To solve this problem, the present invention proposes a maximum combination strategy, such as Figure 5 Shown: Rays emitted by a virtual camera observing the combined scene After different transformations, we get the light passing through different models. ) uniformly sampled points, the corresponding sampling points can be found on different rays. A sampling point on , respectively calculate the termination probability density of the corresponding sampling points on the corresponding light of different models. By selecting the sampling point with the largest termination probability density , you can know the light Passing through When all corresponding sampling points of , then the light Upsampling point The color and volume density should be set to In this way, the correct occlusion relationship is achieved and the scene combination is completed.

[0059] Furthermore, the specific steps of the multi-model combination method based on twisted light are as follows: choose A trained NeRF model to be combined with the target scene, and configure the transformation parameters for each model, including translation, rotation, and scaling factors. Configure 6 clipping plane parameters as required, including a normal vector for each plane. and a translation constant .

[0060] Construct a virtual camera for observing the combined scene, including setting the virtual camera's intrinsic parameter matrix (focal length, principal point coordinates) and extrinsic parameter matrix (camera position, orientation).

[0061] According to the intrinsic parameter matrix and extrinsic parameter matrix of the virtual camera, the light emitted by the virtual camera passing through each virtual image pixel is constructed ,include and , respectively represent the starting point and direction of the light. choose NeRF models to be combined, for the NeRF model, Shi Yidi Rotation transformation of a NeRF model configuration:

[0062] in, Indicates The rotation matrix of the model distorts the light and changes the Rendering view of a NeRF model.

[0063] The starting point of the light obtained by the transformation and light direction , uniformly sample N points on the ray , then Apply translation and scaling transformations:

[0064] in, Indicates The scaling matrix of the model, Indicates The translation matrix of the model controls the The size and position of each model in the combined scene.

[0065] In order to adapt to complex scenes, each NeRF model uses 6 clipping planes to control the area to be retained. The retained area of ​​the clipping plane is defined as follows: Clipping plane The normal vector is , the translation constant is , for a point in space ,like , then point is within the retained area of ​​the clipping plane; on the contrary, if , then point The culling region lies within this clipping plane. The sample point in the hold space is calculated using the following formula:

[0066] in Indicates the sampling point On clipping plane The sampling points of the reserved area.

[0067] In the The density voxel grid of the model is queried The volume density , filter out Invalid sampling points, calculate Contribution to the final integrated color of the light ,as well as Contributing sampling points The position in the density voxel grid and reads the corresponding hash field from it as .

[0068] From A hash table of models Search Spherical harmonic coefficients and volume density of:

[0069] Set the volume density of invalid sampling points, non-contributing sampling points, and sampling points located in the non-retained area of ​​the six clipping planes to 0, and the spherical harmonic coefficients to 0, and we get The spherical harmonic coefficients and body density .

[0070] In the direction of light Next, calculate the color value of the sampling point:

[0071] calculate The corresponding cumulative transmittance :

[0072] in, , , Respectively indicate observation The model emits the The first line The cumulative transmittance, volume density and distance between adjacent sampling points of each sampling point can be obtained. The cumulative transmittance .

[0073] calculate The termination probability density :

[0074] Perform the above operations on each NeRF model and we can get , as well as .

[0075] Rays emitted by the virtual camera Upper uniform sampling point, because in and All are uniformly sampled point, so The sampling points in can all correspond to The sampling points in At the same time, it passes through the corresponding sampling points of different models. Use the maximum combination strategy to construct the combined scene radiation field: Used to describe the position of light The termination probability density at ; if Middle The sampling point on the ray , the light parameters are , exist The corresponding sampling points are , the light parameters are ,and The probability density of ray termination at is ,from Select the one with the largest termination probability density; if Maximum, then Middle The light in It is more likely to be terminated at The corresponding sampling points in the NeRF model Therefore, The color and volume density are color and volume density.

[0076] According to the maximum combination strategy, Upsampling points are assigned colors and body density :

[0077]

[0078] in Respectively Middle The first line The color and volume density of the sampling points, Respectively Middle The first line The color and volume density of the sampling points, .

[0079] Use volume rendering technology to render the light emitted by the virtual camera , we can get a two-dimensional image of the multi-model combination scene. Figure 6 Shown is a flow chart of the multi-model free combination method based on twisted light of the present invention.

[0080] Furthermore, as the number of combined objects increases, the time required to render the entire scene will inevitably increase. For this reason, the present invention only renders the transformed model each time, and re-renders the entire scene only when rendering for the first time and changing the viewing angle, so as to reduce computing overhead and speed up the rendering efficiency of the combined scene.

[0081] like Figure 7 As shown in the figure, the result of rendering the combination of two NeRF models is shown, and the models can be freely manipulated to make various complex transformations; Figure 8 As shown in the figure, the rendering results of multiple NeRF models are displayed. In addition to model transformation, the viewing angle can also be adjusted.

[0082] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.

[0083] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0084] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the NeRF multi-model scene construction methods based on twisted rays in the above-mentioned embodiments.

[0085] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.

[0086] The embodiment of the present application also provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. Memory, used to store computer programs; The processor is used to implement the NeRF multi-model scene construction method based on twisted rays when executing the program stored in the memory.

[0087] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.

[0088] The communication interface is used for communication between the above electronic device and other devices.

[0089] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0090] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0091] It should also be noted that electronic devices also include terminal devices, which can also be called terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Terminal devices can be mobile phones, smart TVs, wearable devices, tablet computers (Pad), computers with wireless transceiver functions, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of the present application do not limit the specific technology and specific device form adopted by the terminal devices.

[0092] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0093] The specific implementation methods described above provide a detailed description of the technical solutions of the present invention. It should be understood that the above is only the most preferred implementation scheme of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A NeRF multi-model scene construction method based on twisted rays, characterized in that: The following steps are involved: S1, obtain real scene and virtual scene data sets, preprocess the data sets, and use motion recovery structure algorithm to calculate image pose; S2, building a GPU environment for training and rendering neural radiance fields, setting training parameters, and loading a scene configuration file and a set of two-dimensional images with camera poses processed in step S1; S3. Use the density voxel grid to optimize the neural radiation field layered sampling strategy. Use the two-dimensional image and camera pose loaded in step S2 to train the spherical harmonic neural radiation field model based on the density voxel grid, and bake the spherical harmonic coefficients and volume density of the contributing sampling points into the hash table in the last round of training. After the training, the hash table can be used to accelerate the neural radiation field rendering; use step S2 to load different data sets, repeat step S3, and train to obtain multiple target combination models; S4, selecting the neural radiation field model trained in step S3, configuring transformation parameters and cropping parameters for each model, and rendering to obtain a combined scene image by distorting light and combining it with a maximum combination strategy; S5. Modify the transformation parameters and cropping parameters of each model in step S4 to achieve scene construction and editing.

2. A NeRF multi-model scene construction method based on twisted light according to claim 1, characterized in that: In step S3, the size of The density voxel grid of the neural radiation field is optimized by layered sampling strategy, specifically: (1-1) Densely and uniformly sample N points on ray r, and query the volume density of these sampled points from the density voxel grid , the sampling points are divided into two categories: is the effective sampling point, are invalid sampling points, among which is the contribution threshold; (1-2) Use the coarse network to calculate the effective sampling points, and use the calculated volume density to update the density voxel grid. Calculate the contribution of the effective sampling points to the final integrated color of the light according to the queried volume density. , further divide the effective sampling points into two categories: are contributing sampling points, are sampling points with no contribution; (1-3) A fine network is used to calculate the contributing sampling points.

3. A NeRF multi-model scene construction method based on twisted light according to claim 1, characterized in that: In step S3, the spherical harmonic coefficients and volume density of the contributing sampling points are baked into a hash table. Specifically, in the last round of training, the spherical harmonic coefficients and volume density of the contributing sampling points are baked into a hash table until the training is completed. The baking steps are as follows: (2-1) Create a new field in the density voxel grid to store the hash value of the sampling point; (2-2) Calculate contributing sampling points Position in the density voxel grid : in express The coordinates in the world coordinate system, Represents the bounds of the scene in world coordinates, represents the size of the density voxel grid in any dimension, express The coordinates of the corresponding position in the density voxel grid in 3 dimensions; (2-3) Get contributing sampling points from the hash field of the density voxel grid The hash value of ; (2-4) If there are contributing sampling points The hash value of It can be found in the hash table, then use Update the hash table with spherical harmonic coefficients and volume density: If the query is not found in the hash table, Insert the spherical harmonic coefficients and volume density into the hash table and record The hash value of is its position in the hash table and updates the density voxel grid The hash field.

4. The NeRF multi-model scene construction method based on twisted light according to claim 1, characterized in that: In step S4, the specific steps of the method for distorting light are: (3-1)Select A trained target scene should contain a neural radiance field model, and the transformation parameters and cropping parameters configured for each model; (3-2) The light emitted by the virtual camera observing the combined scene Rays distorted by the transformation parameters configured for each NeRF model ; (3-3) Upper uniform sampling; Query the volume density of the sampling point in the density voxel grid, and calculate its contribution to the final integrated color of the light, filter out the contributing sampling points, query the spherical harmonic coefficients and volume density of the contributing sampling points in the hash table of each model, set the spherical harmonic coefficients and volume density of the remaining sampling points to 0, and calculate the termination probability density and spherical harmonic coefficients of all sampling points in the light The corresponding color below; (3-4) Let the light At the same time, through each model The corresponding points on the light are integrated into the light by using the maximum combination strategy to integrate the color and volume density of the corresponding sampling points on different light rays. The color and volume density of the corresponding sampling points are shown above, and the maximum combination strategy is used to determine the occlusion relationship of each model in the combined scene.

5. A NeRF multi-model scene construction method based on twisted light according to claim 4, characterized in that: The steps (3-4) specifically include: Used to describe the position of light The termination probability density at ; if Middle The sampling point on the ray , the light parameters are , exist The corresponding sampling points are , the light parameters are ,and The probability density of ray termination at is ,from Select the one with the largest termination probability density; if Maximum, then Middle The light in It is more likely to be terminated at The corresponding sampling points in the NeRF model Therefore, The color and volume density are The color and volume density are: in Respectively Middle The first line The color and volume density of the sampling points, Respectively After The light distorted by the model The first line The color and volume density of the sampling points, ; The termination probability density is: in, represents the probability density of the light being terminated at position t, represents the cumulative transmittance of light at position t, Represents light The volume density at position t.

6. A NeRF multi-model scene construction method based on twisted light according to claim 1, characterized in that: In step S5, the construction of the scene is achieved by adding or reducing models and adjusting the transformation parameters, cropping parameters and viewing angle parameters of each model, and then rendering.

7. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

8. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for drawing neural radiation field in rasterization rendering pipeline in real time

    CN119152100A

  • Real-time neural network radiance caching for path tracing

    US20220284657A1

Cited By

  • Three-dimensional object structure editing method and device based on point cloud representation

    CN121053349A