Three-dimensional forest reconstruction method based on NeRF and GS joint optimization
By improving the combination of NeRF network and Gaussian sputtering, the problems of low efficiency and insufficient precision of forest scene rendering are solved, and efficient and fine forest modeling and visualization are achieved, which is suitable for fields such as ecological protection and virtual reality.
Patent Information
- Application Number
- CN202511000512.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In the prior art, NeRF rendering efficiency is low, GS lacks structural accuracy, making it difficult to achieve efficient and fine modeling of forest scenes.
By improving NeRF network, the introduction of semantic coding and volume density guidance, combined with the Gaussian sputtering method, the Gaussian point set is optimized using a comprehensive weight and structural alignment loss function to achieve efficient rendering of forest scenes.
It improves the rendering efficiency and accuracy of forest scenes, can better capture the details of high-density areas, and provides more accurate semantic visualization effects, suitable for fast forest browsing and three-dimensional interaction on mobile terminals.
Smart Images

Figure CN120510307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a three-dimensional forest reconstruction method based on NeRF and GS joint optimization. Background Art
[0002] NeRF and Gaussian Splatting (GS) are both 3D reconstruction technologies that have demonstrated great potential in natural scene modeling, virtual reality environment construction, and ecological digital twin systems. Forest scene reconstruction has important application needs in areas such as ecological protection, carbon sink assessment, and forest resource management. Accurate 3D reconstruction techniques can provide quantitative analysis support for forest growth patterns, species diversity, and environmental changes. These technologies can help researchers efficiently obtain forest spatial structure, tree distribution, and biodiversity data, providing a reliable data foundation for ecological monitoring, forest management decision-making, and natural disaster early warning systems.
[0003] Neural Radiance Fields (NeRF) is a method for 3D reconstruction using deep learning technology, mainly used to generate high-quality 3D models and scenes. NeRF extracts the geometric shape and texture information of the object from images of multiple perspectives to construct a continuous 3D radiation field, which can present highly realistic 3D models at any angle and distance. The principle of NeRF is to implicitly represent the light distribution and transmission characteristics in the scene by constructing a neural network. The input is a continuous 5-dimensional vector, including any spatial position x (3D space) and viewing direction d (spherical coordinate representation, including azimuth angle θ and pitch angle φ) in the scene. The output is the volume density σ at the position and the color c along the viewing direction at the position. Since the input and output are obtained by function mapping, the function mapping of NeRF is expressed as F θ :(x,d)→(σ,c), where θ is the parameter of the function. Furthermore, spatial position is typically represented by sampling points, which are obtained by sampling along a ray from the camera to a pixel on the imaging plane. Assuming the imaging plane has 1024 pixels, there are 1024 rays.
[0004] Gaussian Splatting (GS) is an advanced 3D modeling and visualization technique. GS uses a large number of three-dimensional Gaussian distributions to represent the scene. Each three-dimensional Gaussian distribution (also known as a 3D Gaussian sphere) contains parameters such as center position, covariance matrix, color, and opacity.
[0005] Currently, traditional NeRF requires uniform or layered sampling of multiple points along each ray. However, in forest scenes, the density distributions of leaf clusters and open areas under the forest vary significantly. Uniform sampling tends to waste sampling points in open areas and lacks precision at the leaves. Furthermore, forests contain a large number of overlapping leaves, intertwined branches, and translucent or reflective features (such as water surfaces and wetlands). Using traditional NeRF solely to estimate the color and density of each pixel can lead to duplicate sampling or difficulty in distinguishing certain areas. NeRF also has slow rendering speeds and high hardware requirements. Gaussian sputtering, on the other hand, is generally based on point cloud data, processing each point as a Gaussian sphere. However, blindly and directly performing uniform point sampling across three-dimensional space can lead to inadequate fitting of key areas, such as high-frequency details, while retaining many points in useless areas. Furthermore, Gaussian sputtering has relatively low training data requirements and can be trained with less image data. This results in a lack of structural accuracy in Gaussian sputtering, making it unsuitable for fine-grained modeling of forest scenes. Therefore, the limitations of existing technologies make it difficult to achieve both high-precision modeling and efficient rendering. Summary of the Invention
[0006] The purpose of the present invention is to provide a three-dimensional forest reconstruction method based on the joint optimization of NeRF and GS to solve the problems of low NeRF rendering efficiency and lack of structural accuracy of GS in forest scene reconstruction, and to achieve high-speed and fine modeling of forest scenes.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows: a three-dimensional forest reconstruction method based on NeRF and GS joint optimization includes the following steps; S1, construct dataset D; Obtain multi-view images of the scene to be reconstructed in the forest, perform pixel-level semantic annotation on the images by category, with a total number of categories K. The annotated images are used as samples, and all samples constitute the dataset D; S2, construct an improved NeRF network; Get a NeRF network whose input is a 5D vector of a spatial point, including 3D coordinates and 2D view angle, and output is the color and volume density of the spatial point; The NeRF network is expanded to add semantic encoding of the category to which the spatial point belongs in the input, add the category probability distribution of the spatial point in the output, and generate sampling points based on the volume density guidance method during volume rendering to obtain an improved NeRF network. The semantic encoding is obtained by semantic annotation encoding. S3, construct the loss function L of the improved NeRF network total and use the dataset D to minimize L total Train until convergence to obtain the NeRF model M of the scene to be reconstructed NeRF , where the spatial points constitute the point set A1; L total =LNeRF +L sem , L NeRF To improve the rendering loss of NeRF network, L sem To reconstruct the semantic loss of pixels in the image based on the improved NeRF network; S4, based on M NeRF , generate Gaussian point set G by comprehensive weight and Gaussian sputtering method; S41, calculate the comprehensive weight of each spatial point in A1, and the comprehensive weight w(p) of spatial point p is calculated according to the following formula; , Where σ(p) is the volume density of the spatial point p, ∇ is the gradient operator, is the L2 norm; S42, randomly sample K spatial points from A1 based on the comprehensive weight to form point set A2; S43, generate a Gaussian sphere for each spatial point in A2 based on the Gaussian sputtering method, and generate the kth spatial point p in A2 k Gaussian sphere G k When G k The center position, color and opacity are p k 3D coordinates, color, and opacity; S44, forming a Gaussian point set G from the K Gaussian balls, and performing rendering based on the Gaussian point set G to obtain a rendered image; S5, construct the total loss of joint optimization ; , , Where, L GS is the rendering loss of Gaussian point set G, L align is the structural alignment loss, λ NeRF ,λ GS ,λ align ,λ sem L NeRF , L GS , L align , L sem The weight of Pointing to p k The 3D coordinates of the sampling point with the largest body density on the ray, μ k G k central location; S6, using dataset D to minimize The parameters of the improved NeRF network and the Gaussian point set G are adjusted to obtain a joint optimization model, and a rendered image is generated based on the joint optimization model.
[0008] Preferably, the semantic annotation is at pixel level, and the categories include tree trunks, leaves, water surface, sky, grass, rocks, bushes, clouds, vehicles and buildings. The semantic coding is generated by Embedding model mapping.
[0009] As a preference, for a spatial point, its category probability distribution l=(l1,l2,⋯,l k ,⋯,l K ), l k is the probability that the spatial point is of category k, 1≤k≤K.
[0010] Preferably, in S2, the volume density guiding method includes steps S21 to S23; S21, generate rays from the camera to each pixel in the imaging plane according to the viewing angle, where the jth ray is r j ; S22, in ray r j Uniformly sample several sampling points and calculate the volume density, where the i-th sampling point pt i The volume density is , S23, generate pt according to the following formula i The subsampling spacing δ i , and in pt i Around by δ i Sampling again to generate several sampling points; , Where C is the adjustment coefficient, and ϵ is a very small number that prevents the denominator from being zero.
[0011] As a preference, the semantic loss L in S3 sem Calculate according to the following formula; , , When training the improved NeRF network, rays from the camera to each pixel in the imaging plane are generated according to the viewing angle, where the jth ray is r j , For the ray r j The probability distribution vector, N is r j The total number of upsampling points, for r j The nth sampling point x j,n , T j,n , α j,n , I j,n x j,n Cumulative transmittance, opacity, and predicted category probability distribution; y j For r jCorresponding to the category of the pixel point in the imaging plane, P(⋅│⋅) is the conditional probability.
[0012] As a preference, in S43, the Gaussian sphere G k =(μ k ,Σ k ,c k ,α k ), μ k ,Σ k 、c k , α k G k The center position, covariance matrix, color and opacity of; , Where, Σ k The size is 3×3, v1, v2, and v3 are the principal axis, the first minor axis, and the second minor axis of the Gaussian sphere, respectively. The three are orthogonal. λ1, λ2, and λ3 are the tensile strengths of v1, v2, and v3, respectively. T is the transpose operation.
[0013] As a preferred embodiment, the joint optimization model obtained in S6 is specifically: S61, preset number of iterations; S62, generated by the improved NeRF network NeRF , then based on M NeRF Generate a Gaussian point set G, then generate a rendered image based on the Gaussian point set G, and calculate the total loss L total , and then minimize L total Adjust and improve the parameters of NeRF network and Gaussian point set G; S63, repeat S62 until the number of iterations is reached, and the last NeRF model M NeRF and Gaussian point set G as a joint optimization model.
[0014] The main improvements of the present invention include: 1. Improve the NeRF network, introduce a structural semantic perception mechanism, add semantic encoding to its 5-dimensional input, and add category probability distribution to the output.
[0015] Based on this improvement, the NeRF network can be transformed from F θ (x,d)=(σ,c) is expanded to F θ (x, d, s) = (σ, c, l), in this expression, θ is the network parameter of NeRF network, F θ (∙) is the output of the NeRF network, F θ (x, d) = (σ, c) means: when the coordinate x and viewing angle d of a spatial point are input, the opacity σ and color c of the spatial point are output. θ(x, d, s) = (σ, c, l) means that when the coordinates x, viewing angle d, and semantic code s of a spatial point are input, the opacity σ, color c, and category probability distribution l of the spatial point are output.
[0016] In the present invention, the semantic code s is based on semantic information and can be obtained through a small amount of manual annotation or an external classification network, and the semantic code is in vector form, which is used to indicate whether the local area belongs to "trunk", "leaf", "water surface", "sky background", etc. The semantic code s can actually be introduced in the input layer or middle layer of the existing NeRF network. The category probability distribution l also contains semantic information, which is used to indicate the category tendency of the current spatial point. The semantic code s is a priori class label embedding, and the category probability distribution l is the semantic activation obtained by automatic learning of the NeRF network, and is converted into a probability distribution through softmax. This method can significantly reduce the repeated estimation of high-similarity areas, allowing the network to express high-density areas (such as leaf areas) and low-density areas (such as water surfaces) in a targeted manner, thereby improving modeling efficiency.
[0017] 2. Improve the NeRF network and introduce a volume density guidance method when sampling rays.
[0018] Traditional NeRF volume rendering requires uniform or layered sampling of multiple points along each ray. However, in forest scenes, the density distributions of leaf clusters and open areas under the forest vary significantly. Uniform sampling can lead to wasted sampling points in open areas and insufficient detail at the leaves. Therefore, the proposed method first uniformly samples the volume density of the sampling points. Then, based on the volume density, a secondary sampling spacing is calculated, generating more sampling points in denser areas and fewer in less dense areas. This method allows for denser sampling of high-density areas (such as leaves) to capture leaf texture and lighting details, while appropriately reducing sampling of low-density areas (such as the sky or simple background areas under the forest) to reduce computational complexity.
[0019] 3. Improve Gaussian sputtering and generate Gaussian point sets based on the NeRF model of forest scenes.
[0020] Because NeRF reconstruction requires sampling dozens or even hundreds of points along each ray, predicting density and color at each point, and finally performing weighted accumulation to produce pixel color, this process is highly accurate but extremely costly. Therefore, this paper uses Gaussian sputtering (GS) rendering to replace the NeRF rendering process. Specifically, it can be seen as the following two steps: Step 1: The concept of comprehensive weight is proposed, and point set A2 is obtained by sampling from point set A1 based on the comprehensive weight. The comprehensive weight is based on the volume density of spatial points. Taking spatial point p as an example, σ(p) is the volume density of spatial point p, which can reflect its "presence in the scene". If it is a transparent area, the volume density is approximately equal to 0; if it is a solid object such as a tree trunk or a leaf, the volume density value will be very large. ∇ is a gradient operator, ∇σ(p) represents the gradient of volume density in space, ∇σ(p) can reflect the local "degree of structural change", such as the edge of the object, the occlusion boundary, and the turning point of the surface, ∇σ(p) will be very large; while inside the object or in the smooth area, ∇σ(p) is close to 0. The comprehensive weight w(p) is obtained based on σ(p) and ∇σ(p). The larger w(p) is, the more it means that the spatial point p is both real and has a structural mutation. The point p is more worthy of subsequent Gaussian fitting. The present invention proposes random sampling based on comprehensive weights, which can be obtained from M NeRF The method tries to pick out as many areas as possible, such as edges, rich details, and prominent objects, while minimizing the background, sky, and other areas. This method can compress the number of Gaussians while retaining effective information, saving rendering resources and improving visual fidelity.
[0021] Step 2: Based on A2 and M NeRF The Gaussian point set G is quickly generated by using the relevant information in the Gaussian sphere. Since the Gaussian sphere includes the center position, covariance matrix, color and opacity, the covariance matrix in the present invention is constructed based on the existing method of Gaussian sputtering, while the center position, color and opacity are directly obtained from M. NeRF Read in, thus avoiding recalculation in Gaussian sputtering.
[0022] 4. Improve the loss function and introduce semantic loss L sem and structure alignment loss L align , semantic loss L sem This enables the model to not only learn the visual information of the scene, but also effectively obtain and utilize the semantic information of the scene, thereby generating more accurate and reliable rendering results. align The purpose is to make the center of the Gaussian sphere as close as possible to the center of the object, with the Gaussian sphere p k For example, μ k For p k central location, It points to p k The 3D coordinates of the sampling point with the largest body density on the ray can be understood as the center of the object on the ray. align , which makes the center of each Gaussian sphere closer to the center of the object rather than deviating from the surface.
[0023] Compared with the prior art, the advantages of the present invention are: (1) By introducing density gradient perception and a sampling strategy based on comprehensive weights, the NeRF model can effectively screen key area point sets in three-dimensional space. This approach not only improves the redundant sampling problem in modeling, but also ensures the accuracy of expressing high-density areas such as leaves and occlusion boundaries, providing stronger data support for forest structure modeling.
[0024] (2) Introducing semantic label supervision and the NeRF semantic extension mechanism as a guiding link in the generation process, each fitted Gaussian point has type distinguishability and regional awareness capabilities. This approach not only focuses on restoring the overall forest distribution, but also accurately models specific semantic areas such as tree trunks, water surfaces, and understory, significantly improving the effects of three-dimensional segmentation and semantic visualization of natural scenes. In particular, this technology significantly improves the practical value of automatic labeling and intelligent visualization in ecological classification and digital forestry systems.
[0025] (3) By constructing a structural alignment and joint optimization mechanism between NeRF and GS, the present invention achieves the integrated collaboration of modeling network and rendering representation, greatly alleviating the bottleneck problem of NeRF in inference speed, while avoiding the problem of GS losing details due to lack of semantic guidance, and providing technical support for fast forest browsing and three-dimensional interaction on mobile terminals.
[0026] In summary, this invention effectively addresses existing issues such as low NeRF rendering efficiency and lack of structural accuracy in GS. It provides a new technical path for high-precision and rapid forest modeling, offering novel solutions for scenarios such as digital forestry, ecological protection, and virtual reality. It significantly enhances the quality and application value of forest modeling and visualization. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Flowchart of the present invention; Figure 2 To improve the NeRF network structure diagram. DETAILED DESCRIPTION
[0028] The present invention will be further described below with reference to the embodiments and accompanying drawings.
[0029] Example 1: See Figure 1 and Figure 2 ,A 3D forest reconstruction method based on the joint optimization of NeRF and GS, comprising the following steps; S1, construct dataset D; Obtain multi-view images of the scene to be reconstructed in the forest, perform pixel-level semantic annotation on the images by category, with a total number of categories K. The annotated images are used as samples, and all samples constitute the dataset D; S2, construct an improved NeRF network; Get a NeRF network whose input is a 5D vector of a spatial point, including 3D coordinates and 2D view angle, and output is the color and volume density of the spatial point; The NeRF network is expanded to add semantic encoding of the category to which the spatial point belongs in the input, add the category probability distribution of the spatial point in the output, and generate sampling points based on the volume density guidance method during volume rendering to obtain an improved NeRF network. The semantic encoding is obtained by semantic annotation encoding. S3, construct the loss function L of the improved NeRF network total and use the dataset D to minimize L total Train until convergence to obtain the NeRF model M of the scene to be reconstructed NeRF , where the spatial points constitute the point set A1; L total =L NeRF +L sem , L NeRF To improve the rendering loss of NeRF network, L sem To reconstruct the semantic loss of pixels in the image based on the improved NeRF network; S4, based on M NeRF , generate Gaussian point set G by comprehensive weight and Gaussian sputtering method; S41, calculate the comprehensive weight of each spatial point in A1, and the comprehensive weight w(p) of spatial point p is calculated according to the following formula; , Where σ(p) is the volume density of the spatial point p, ∇ is the gradient operator, is the L2 norm; S42, randomly sample K spatial points from A1 based on the comprehensive weight to form point set A2; S43, generate a Gaussian sphere for each spatial point in A2 based on the Gaussian sputtering method, and generate the kth spatial point p in A2 k Gaussian sphere G k When G k The center position, color and opacity are p k 3D coordinates, color, and opacity; S44, forming a Gaussian point set G from the K Gaussian balls, and performing rendering based on the Gaussian point set G to obtain a rendered image; S5, construct the total loss of joint optimization ; , , Where, L GS is the rendering loss of Gaussian point set G, L align is the structural alignment loss, λNeRF ,λ GS ,λ align ,λ sem L NeRF , L GS , L align , L sem The weight of Pointing to p k The 3D coordinates of the sampling point with the largest body density on the ray, μ k G k central location; S6, using dataset D to minimize The parameters of the improved NeRF network and the Gaussian point set G are adjusted to obtain a joint optimization model, and a rendered image is generated based on the joint optimization model.
[0030] The semantic annotation is at the pixel level, and the categories include tree trunks, leaves, water surface, sky, grass, rocks, bushes, clouds, vehicles and buildings. The semantic coding is generated by Embedding model mapping.
[0031] For a spatial point, its category probability distribution l=(l1,l2,⋯,l k ,⋯,l K ), l k is the probability that the spatial point is of category k, 1≤k≤K.
[0032] Example 2: See Figure 1 and Figure 2 , based on Example 1, regarding the volume density guidance method of S2, we provide a specific volume density guidance method, including steps S21~S23; S21, generate rays from the camera to each pixel in the imaging plane according to the viewing angle, where the jth ray is r j ; S22, in ray r j Uniformly sample several sampling points and calculate the volume density, where the i-th sampling point pt i The volume density is , S23, generate pt according to the following formula i The subsampling spacing δ i , and in pt i Around by δ i Sampling again to generate several sampling points; , Where C is the adjustment coefficient, and ϵ is a very small number that prevents the denominator from being zero.
[0033] In this embodiment, the semantic loss L in step S3 is sem Calculate according to the following formula; , , When training the improved NeRF network, rays from the camera to each pixel in the imaging plane are generated according to the viewing angle, where the jth ray is r j , For the ray r j The probability distribution vector, N is r j The total number of upsampling points, for r j The nth sampling point x j,n , T j,n , α j,n , I j,n x j,n Cumulative transmittance, opacity, and predicted category probability distribution; y j For r j Corresponding to the category of the pixel point in the imaging plane, P(⋅│⋅) is the conditional probability.
[0034] In S43, Gaussian sphere G k =(μ k ,Σ k ,c k ,α k ), μ k ,Σ k 、c k , α k G k The center position, covariance matrix, color and opacity of; , Where, Σ k The size is 3×3, v1, v2, and v3 are the principal axis, the first minor axis, and the second minor axis of the Gaussian sphere, respectively. The three are orthogonal. λ1, λ2, and λ3 are the tensile strengths of v1, v2, and v3, respectively. T is the transpose operation.
[0035] The specific joint optimization model obtained by S6 is: S61, preset number of iterations; S62, generated by the improved NeRF network NeRF , then based on M NeRF Generate a Gaussian point set G, then generate a rendered image based on the Gaussian point set G, and calculate the total loss L total , and then minimize L total Adjust and improve the parameters of NeRF network and Gaussian point set G; S63, repeat S62 until the number of iterations is reached, and the last NeRF model M NeRF and Gaussian point set G as a joint optimization model.
[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A 3D forest reconstruction method based on NeRF and GS joint optimization, characterized in that: The following steps are included: S1, construct dataset D; Obtain multi-view images of the scene to be reconstructed in the forest, perform pixel-level semantic annotation on the images by category, with a total number of categories K. The annotated images are used as samples, and all samples constitute the dataset D; S2, construct an improved NeRF network; Get a NeRF network whose input is a 5D vector of a spatial point, including 3D coordinates and 2D viewing angle, and output is the color and volume density of the spatial point; The NeRF network is expanded to add semantic encoding of the category to which the spatial point belongs in the input, add the category probability distribution of the spatial point in the output, and generate sampling points based on the volume density guidance method during volume rendering to obtain an improved NeRF network. The semantic encoding is obtained by semantic annotation encoding. S3, construct the loss function L of the improved NeRF network total and use the dataset D to minimize L total Train until convergence to obtain the NeRF model M of the scene to be reconstructed NeRF , where the spatial points constitute the point set A1; L total =L NeRF +L sem , L NeRF To improve the rendering loss of NeRF network, L sem To reconstruct the semantic loss of pixels in the image based on the improved NeRF network; S4, based on M NeRF , generate Gaussian point set G by comprehensive weight and Gaussian sputtering method; S41, calculate the comprehensive weight of each spatial point in A1, and the comprehensive weight w(p) of spatial point p is calculated according to the following formula; , Where σ(p) is the volume density of the spatial point p, ∇ is the gradient operator, is the L2 norm; S42, randomly sample K spatial points from A1 based on the comprehensive weight to form point set A2; S43, generate a Gaussian sphere for each spatial point in A2 based on the Gaussian sputtering method, and generate the kth spatial point p in A2 k Gaussian sphere G k When G k The center position, color and opacity are p k 3D coordinates, color, and opacity; S44, forming a Gaussian point set G from the K Gaussian balls, and performing rendering based on the Gaussian point set G to obtain a rendered image; S5, construct the total loss of joint optimization ; , , Where, L GS is the rendering loss of Gaussian point set G, L align is the structural alignment loss, λ NeRF ,λ GS ,λ align ,λ sem L NeRF , L GS , L align , L sem The weight of Pointing to p k The 3D coordinates of the sampling point with the largest body density on the ray, μ k G k central location; S6, using dataset D to minimize The parameters of the improved NeRF network and the Gaussian point set G are adjusted to obtain a joint optimization model, and a rendered image is generated based on the joint optimization model.
2. A 3D forest reconstruction method based on NeRF and GS joint optimization according to claim 1, characterized in that: The semantic annotation is at the pixel level, and the categories include tree trunks, leaves, water surface, sky, grass, rocks, bushes, clouds, vehicles and buildings. The semantic coding is generated by Embedding model mapping.
3. The three-dimensional forest reconstruction method based on NeRF and GS joint optimization according to claim 1 is characterized in that: For a spatial point, its category probability distribution l=(l1,l2,⋯,l k ,⋯,l K ), l k is the probability that the spatial point is of category k, 1≤k≤K.
4. The 3D forest reconstruction method based on NeRF and GS joint optimization according to claim 1, characterized in that: In S2, the body density guidance method includes steps S21 to S23; S21, generate rays from the camera to each pixel in the imaging plane according to the viewing angle, where the jth ray is r j ; S22, in ray r j Uniformly sample several sampling points and calculate the volume density, where the i-th sampling point pt i The volume density is , S23, generate pt according to the following formula i The subsampling spacing δ i , and in pt i Around by δ i Sampling again to generate several sampling points; , Where C is the adjustment coefficient, and ϵ is a very small number that prevents the denominator from being zero.
5. The three-dimensional forest reconstruction method based on NeRF and GS joint optimization according to claim 1 is characterized in that: Semantic loss L in S3 sem Calculate according to the following formula; , , When training the improved NeRF network, rays from the camera to each pixel in the imaging plane are generated according to the viewing angle, where the jth ray is r j , For the ray r j The probability distribution vector, N is r j The total number of upsampling points, for r j The nth sampling point x j,n , T j,n , α j,n , I j,n x j,n Cumulative transmittance, opacity, and predicted category probability distribution; y j For r j Corresponding to the category of the pixel point in the imaging plane, P(⋅│⋅) is the conditional probability.
6. The three-dimensional forest reconstruction method based on NeRF and GS joint optimization according to claim 1 is characterized in that: In S43, Gaussian sphere G k =(μ k ,Σ k ,c k ,α k ), μ k ,Σ k 、c k , α k G k The center position, covariance matrix, color and opacity of; , Where, Σ k The size is 3×3, v1, v2, and v3 are the principal axis, the first minor axis, and the second minor axis of the Gaussian sphere, respectively. The three are orthogonal. λ1, λ2, and λ3 are the tensile strengths of v1, v2, and v3, respectively. T is the transpose operation.
7. The 3D forest reconstruction method based on NeRF and GS joint optimization according to claim 1, characterized in that: The specific joint optimization model obtained by S6 is: S61, preset number of iterations; S62, generated by the improved NeRF network NeRF , then based on M NeRF Generate a Gaussian point set G, then generate a rendered image based on the Gaussian point set G, and calculate the total loss L total , and then minimize L total Adjust and improve the parameters of NeRF network and Gaussian point set G; S63, repeat S62 until the number of iterations is reached, and the last NeRF model M NeRF and Gaussian point set G as a joint optimization model.
Citation Information
Patent Citations
Three-dimensional forest reconstruction method based on 3DGS technology
CN119942016A
Large-scene lightweight three-dimensional reconstruction method based on NeRF and 3DGS mixed representation
CN120182507A
Indoor real scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering
CN120279159A