Neural radiation field three-dimensional reconstruction method fusing residual feature enhancement and depth supervision
By integrating residual feature enhancement and deep supervision into a neural radiation field 3D reconstruction method, the problems of deep feature degradation and lack of geometric supervision in satellite image 3D reconstruction are solved, achieving high-precision 3D reconstruction in complex scenes and generating more accurate DSM elevation maps.
Patent Information
- Application Number
- CN202510843245.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-11-14
AI Technical Summary
Existing satellite image 3D reconstruction methods are prone to degradation of original spatial features during deep transmission, making it difficult to capture local details, resulting in low reconstruction accuracy and a lack of effective geometric supervision, which limits their application in complex scenes.
A neural radiation field 3D reconstruction method that integrates residual feature enhancement and deep supervision is adopted. Through the RDS-NeRF network model, combined with the basic component generation module, residual component generation module and fusion module, temporal embedding vector and deep supervision loss function are introduced to improve the model's ability to capture local details and geometric constraints.
It significantly improves the accuracy and robustness of 3D reconstruction from satellite imagery, especially in generating more accurate DSM elevation maps in complex scenes such as buildings, roads, water bodies, and vegetation. It enhances the model's ability to perceive high-frequency information and the semantic connection between global and local information, thereby improving the model's generalization ability and robustness.
Smart Images

Figure CN120953477A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of three-dimensional reconstruction, specifically relating to a three-dimensional reconstruction method for neural radiation fields that integrates residual feature enhancement and depth supervision. Background Technology
[0002] 3D reconstruction is an important research direction in the field of remote sensing photogrammetry. Multi-view imagery, which integrates data from different angles, is to some extent more accurate than digital surface models (DSMs) generated from single-view or stereo images, and is widely used in tasks such as digital earth, urban modeling, and disaster assessment. Currently, the main methods for acquiring 3D models include airborne LiDAR point cloud modeling, UAV imagery modeling, and satellite imagery modeling. While the former two can provide high-precision 3D models, they are costly and have limited coverage. Satellite imagery, on the other hand, has advantages such as global coverage, long-term monitoring, and large imaging area, making it an important data source for constructing realistic 3D models.
[0003] Significant progress has been made in 3D reconstruction of satellite imagery, but several shortcomings remain. Traditional multi-view stereo matching techniques generate depth models (DSMs) by obtaining depth or disparity information through dense matching, but errors easily occur in textureless, occluded, and disparity abrupt regions. With the development of deep learning, end-to-end depth estimation methods based on multi-view imagery have gradually emerged, predicting depth through steps such as feature extraction, homography transformation, cost volume construction and regularization, and depth regression. However, these methods do not fully consider the effects of illumination and radiation, resulting in insufficiently refined DSMs. In recent years, Neural Radiation Field (NeRF) technology has shown significant potential in satellite scene reconstruction. Most existing satellite NeRFs use deep MLP backbone structures to model scene geometry. Although they possess good global modeling capabilities, the original spatial features are prone to degradation during deep transfer, making it difficult for the model to capture local details, thus affecting reconstruction accuracy. Furthermore, these methods are mostly end-to-end single-training methods, making it difficult to reuse or optimize scene features, limiting their application in complex scenes. Summary of the Invention
[0004] This invention proposes a neural radiation field 3D reconstruction method that integrates residual feature enhancement and deep supervision. It addresses the technical problems that existing NeRF networks mostly use deep MLP backbone structures to model scene geometry when performing 3D reconstruction. During deep transmission, the original spatial features are prone to degradation, making it difficult for the model to capture local details and thus affecting the reconstruction accuracy.
[0005] This invention can be achieved through the following technical solutions:
[0006] A method for three-dimensional reconstruction of neural radiation fields that integrates residual feature enhancement and depth supervision includes the following steps:
[0007] Step 1: Acquire multi-temporal satellite images of the reconstructed area;
[0008] Step 2: For any time-phase satellite image, preprocessing is performed first, and then ray sampling is performed on any pixel to obtain the feature vector corresponding to each sampling point;
[0009] Step 3: Input the feature vector corresponding to each sampling point into the RDS-NeRF network model to infer the final component corresponding to each sampling point;
[0010] The RDS-NeRF network model uses a Sat-NeRF network structure to extract basic features, processes the basic features, outputs basic components, then uses residual blocks to enhance the basic features, processes the enhanced basic features, outputs residual components, and finally adds and fuses the residual components with the basic components to output the final component.
[0011] Step 4: Using volume rendering technology, process the final components corresponding to the sampling points contained in each pixel to obtain the color and depth information corresponding to each pixel, and convert them into three-dimensional coordinates to generate a DSM elevation map, thus completing the three-dimensional reconstruction of the reconstructed area.
[0012] Furthermore, the basic components, residual components, and final components are all set to five, namely color, volume density, uncertainty coefficient, shadow grayscale value, and ambient light. The feature vectors are set to four, namely the position vector, ray direction vector, sunlight direction vector, and time embedding vector corresponding to each sampling point.
[0013] The RDS-NeRF network model includes a basic component generation module, a residual component generation module, and a fusion module. The basic component generation module adopts a Sat-NeRF network structure, including a backbone network and a first multilayer perceptron (MLP). The residual component generation module includes a residual block and a second multilayer perceptron (MLP).
[0014] The backbone network first encodes the position vector and the ray direction vector, then processes the encoded position vector to extract the basic feature h. base Then, for the basic feature h base The process is performed to generate the basic volume density;
[0015] The first multilayer perceptron (MLP) uses basic features h base The input consists of the sunlight direction vector, the time embedding vector, and the position-encoded ray direction vector. The output components are color, uncertainty coefficient, shadow grayscale value, and ambient light.
[0016] The basic features are concatenated with the position vector after position encoding, and this is used as the input to the residual block to extract the residual features h. res And the residual characteristics h res The process is performed to generate the residual component volume density;
[0017] The second multilayer perceptron (MLP) uses residual features h res The input consists of the sunlight direction vector, the time embedding vector, and the position-encoded light direction vector. The output consists of five residual components: color, uncertainty coefficient, shadow grayscale value, and ambient light.
[0018] The fusion module is used to add and fuse the five residual components with the five basic components, and output the five final components: color, uncertainty coefficient, shadow grayscale value, and ambient light.
[0019] Furthermore, let c be the fundamental components: color, uncertainty coefficient, shadow grayscale value, and ambient light. base ,σ base ,β base ,s base sky base The residual component color, uncertainty coefficient, shadow grayscale value, and ambient light are respectively c res ,σ res ,β res ,s res sky res The position vector, ray direction vector, sunlight direction vector, and time embedding vector of each sampling point are x, d, ω, and t, respectively.
[0020] The backbone network is configured as eight fully connected networks, each with a scale of h. The first multilayer perceptron (MLP) and the second multilayer perceptron (MLP) have the same structure.
[0021] The residual characteristics are calculated using the following formula.
[0022] h res =F res2 (ReLU(F res1 (z)))
[0023] z = concat(γ(x), h base )
[0024] Among them, F res1 Indicates a two-layer fully connected network, F res2 Let represent a fully connected layer, and γ(x) represent the position vector after position encoding.
[0025] Calculate the five fundamental components using the following formula.
[0026]
[0027]
[0028] in, These represent the fully connected networks in the first multilayer perceptron (MLP). represents a three-layer fully connected network, and the rest represent a single-layer fully connected network. γ(d) represents the position-encoded ray direction vector.
[0029] Calculate the five residual components using the following formula.
[0030]
[0031] in, These represent the fully connected networks in the second multilayer perceptron (MLP). represents a three-layer fully connected network, and the rest represent a single-layer fully connected network. γ(d) represents the position-encoded ray direction vector.
[0032] Calculate the five final components I using the following formula. final ,
[0033] I final =I base +I res
[0034] Among them, I base Take c respectively base ,σ base ,β base ,s base sky base I res Take c respectively res ,σ res ,β res ,s res sky res .
[0035] Furthermore, the total loss function of the RDS-NeRF network model is set as follows:
[0036] L = L RGB +λ SC L SC +λ MD L MD (R)
[0037]
[0038] A = [D] ref 1]
[0039] Among them, LRGB L represents the uncertainty-weighted loss function. SC L represents the solar correction term function. MD Let λ represent the deep supervision loss function. SC , λ MD All represent hyperparameters, R represents the set of pixels contained in the satellite image, r represents any pixel, p represents the scale factor, q represents the translation offset, and D represents the hyperparameter. ref This represents the dense depth map output by the Depth Anything v2 model, D. pred This represents the predicted depth map obtained by processing the final components of the RDS-NeRF network model output using volume rendering technology, where A represents the alignment matrix.
[0040] Furthermore, in step two, when preprocessing any temporal satellite image, the maximum elevation value h corresponding to the satellite image is first obtained from the image metadata. max and minimum elevation value h min Simultaneously, the solar direction vector is calculated based on the solar altitude angle and azimuth angle, and its PRC parameters are adjusted and optimized. Then, the maximum elevation value h is used as the basis for the calculation. max Minimum elevation value h min Using the optimized PRC parameters as input, ray sampling is performed on any pixel in the satellite image to obtain the position vector and ray direction vector of each sampling point. Finally, the position vector, ray direction vector, sunlight direction vector, and time embedding vector are used as the feature vector of each sampling point.
[0041] Furthermore, in step four, the depth information of each pixel is converted into normalized three-dimensional coordinates using the following formula.
[0042]
[0043] Where o represents the origin of ray g, d represents the direction of ray g, and z represents the depth information of the pixel corresponding to ray g;
[0044] Subsequently, it was converted into actual spatial coordinates in the Earth-centered Earth-fixed coordinate system ECEF through inverse normalization. Then, use the following formula to convert the actual spatial coordinates into geographic coordinates, including latitude. Longitude λ g and elevation h g ,
[0045]
[0046] Where ecef_to_latlon is the geographic transformation function;
[0047] Finally, the geographic coordinates are projected into the east and north coordinates in the UTM plane coordinate system. The spatial extent of the DSM is further determined based on the ROI boundary and resolution. Then, the plyflatten tool is used to spatially interpolate the point cloud to generate the DSM elevation map.
[0048] The beneficial technical effects of this invention are as follows:
[0049] 1. This invention proposes a highly applicable deep learning network model, RDS-NeRF, for 3D reconstruction of satellite imagery. This model enables high-precision 3D reconstruction based on multi-view satellite imagery in various complex scenarios (buildings, roads, water bodies, vegetation), generating more accurate and robust DSM elevation maps. The model fully considers uncertainties such as illumination differences, imaging time, and transient disturbances in satellite imagery, and introduces temporal embedding vectors to model multi-temporal images, significantly improving the model's generalization ability and robustness under different observation conditions.
[0050] 2. The RDS-NeRF network model of this invention introduces a residual feature enhancement structure to solve the problem of spatial feature degradation in deep networks, improve feature reuse capability, effectively capture scene spatial details, and improve the depth estimation accuracy of building edges, low-texture areas, etc., ensuring more accurate and consistent DSM elevation map generation. This structure, guided by basic features, focuses on supplementing local details that are difficult for the backbone network to model through lightweight residual branches, forming a main-auxiliary structure collaborative modeling mechanism. This not only improves the network's ability to perceive high-frequency information but also enhances the semantic connections and structural consistency between the global and local layers.
[0051] 3. This invention proposes a simple and effective depth estimation loss function. Combined with monocular depth estimation results, it effectively supervises the depth estimated by the model, enhancing geometric constraints in the DSM elevation map generation process and making the prediction results closer to the real terrain. Compared with traditional methods that rely solely on color consistency loss or sparse depth point supervision, this invention introduces a dense depth map as a geometric prior and unifies the depth scale through a scale alignment strategy. This significantly improves the model's ability to perceive and complete the overall terrain contour, especially demonstrating superior DSM accuracy and spatial integrity in weakly textured or dynamic scenes such as water bodies. Attached Figure Description
[0052] Figure 1 This is an overall flowchart of the three-dimensional reconstruction method of the neural radiation field proposed in this invention;
[0053] Figure 2 This is a diagram of the RDS-NeRF network model structure of the three-dimensional reconstruction method of neural radiation field proposed in this invention;
[0054] Figure 3This is a schematic diagram of the DSM results for four areas of a building scene in a specific embodiment of the present invention;
[0055] Figure 4 This is a schematic diagram of the DSM results for four areas of a road scene in a specific embodiment of the present invention;
[0056] Figure 5 This is a schematic diagram of the DSM results for four areas in a water scene in a specific embodiment of the present invention;
[0057] Figure 6 This is a schematic diagram of the DSM results for four areas of a vegetation scene in a specific embodiment of the present invention;
[0058] Figure 7 This is a schematic diagram of vegetation from different viewing angles in a specific embodiment of the present invention. Detailed Implementation
[0059] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings and preferred embodiments.
[0060] like Figure 1 As shown, this invention provides a three-dimensional reconstruction method for neural radiation fields that integrates residual feature enhancement and depth supervision. First, the reconstruction region is determined, multi-temporal satellite images of the region are acquired, PRC parameters are adjusted and optimized, and the maximum elevation value h of the satellite images is obtained. max and minimum elevation value h min Simultaneously, the direction vector of sunlight in satellite imagery is acquired, and then based on h max h min Ray sampling is performed on any pixel in the image using the optimized RPC parameters to obtain the position vector and ray direction vector of each sampling point, and the position is encoded respectively. Then, the three encoded vectors are input into the RDS-NeRF network model to infer parameters such as color and volume density of each sampling point. Next, volume rendering technology is used to synthesize the color information of camera rays, that is, the color and depth information of each pixel. A color loss function is constructed using the rendered color and the real color of the pixel, and a depth supervision loss function is constructed using the predicted depth and the introduced monocular estimated depth. The loss function is backpropagated to the RDS-NeRF network model to optimize the network parameters. Finally, the trained and optimized RDS-NeRF network is used to infer the reconstructed region, generate the DSM, and synthesize a new view to complete the 3D reconstruction.
[0061] Specifically as follows:
[0062] Step 1: Acquire multi-temporal satellite images of the reconstructed area
[0063] This invention utilizes multi-view, multi-temporal RGB imagery acquired by the WorldView-3 satellite. The 2019 IEEE GRSS Data Fusion Contest (DFC) provided a WorldView-3 (WV-3) satellite imagery dataset, covering the area of Jacksonville, Florida, USA, from 2014 to 2016. This dataset provides panchromatic sharpened RGB imagery with a spatial resolution of 0.3m and a size of 2048×2048 pixels, with each image having its corresponding RPC parameters. Additionally, the dataset provides an airborne LiDAR DSM with a sampling distance of 0.5m and a size of 512×512 pixels. Considering existing computer resources, we cannot directly train the dataset using the 2048×2048 pixel multi-view imagery; therefore, we cropped the original images to approximately 800×800 pixels based on the range of the LiDAR DSM.
[0064] Meanwhile, to test the applicability and robustness of our network model, we selected four different types of complex scenes in the WV-3 satellite image dataset, including buildings, roads, water, and vegetation, and selected four different sets of images for each scene.
[0065] Each selected scene presents different terrain features and reconstruction difficulties. Building areas have regular structures with clear roof and wall features, but tall buildings severely obstruct the view, making shadowed areas difficult to recover. Road areas are flat and continuous, with significant elevation differences between elevated and ground-level roads, and interference from transient objects such as vehicles. Water areas are highly reflective, making depth difficult to obtain, and are easily affected by dynamic objects such as boats. Vegetation areas have relatively uniform textures with many repetitive structures such as tree canopies; tree heights vary greatly and are significantly affected by seasonal changes. We specifically selected multi-temporal images with significant changes in viewing angle and imaging time to test the robustness of the algorithm in these complex scenes. Detailed information on the image data used in the experiment is shown in Table 1.
[0066] Table 1. Detailed information on data for different types of scenarios.
[0067]
[0068] Step 2: For any time-phase satellite image, preprocessing is performed first, and then ray sampling is performed on any pixel to obtain the feature vector corresponding to each sampling point.
[0069] When preprocessing satellite imagery at any given time, the maximum elevation value h corresponding to the satellite imagery is first obtained from the image metadata. max and minimum elevation value h min Simultaneously, the solar direction vector is calculated based on the solar altitude angle and azimuth angle, and its PRC parameters are adjusted and optimized. Then, the maximum elevation value h is used as the basis for the calculation. max Minimum elevation value hmin Using the optimized PRC parameters as input, ray sampling is performed on any pixel in the satellite image to obtain the position vector and ray direction vector of each sampling point. In addition, in order to model the instantaneous differences in satellite images caused by temporary objects (such as vehicles, cloud shadows) or changes in illumination at different observation times, a temporal embedding vector is introduced for each satellite image. This temporal embedding vector is related to the observation time of the satellite image and does not come directly from the metadata. Instead, it is used as a learnable function of the image index j and is obtained through backpropagation optimization along with other network parameters during network training. Its vector dimension is manually set to 4 to represent a compact temporal feature expression capability. Finally, the position vector, ray direction vector, sunlight direction vector, and temporal embedding vector are used as the feature vectors of each sampling point.
[0070] Step 3: Input the feature vector corresponding to each sampling point into the RDS-NeRF network model to infer the final component corresponding to each sampling point.
[0071] 3.1 Constructing the RDS-NeRF network model
[0072] Our proposed RDS-NeRF network model includes a basic component generation module, a residual component generation module, and a fusion module. First, the basic features are extracted using the Sat-NeRF network structure and processed to output basic components. Then, residual blocks are used to enhance the basic features, and the enhanced basic features are processed to output residual components. Finally, the residual components and basic components are added together and fused to output the final component. These basic components, residual components, and final components are all set to five, namely color, volume density, uncertainty coefficient, shadow gray value, and ambient light. Among them, color represents the inherent color of static objects in the image, and ambient light is only related to the incident sunlight and represents the illumination color from the sky.
[0073] In other words, the Sat-NeRF structure is adopted, which is a backbone network composed of multiple fully connected layers. The network is initialized with SIREN to enhance its representation capability. The input position information is encoded and then enters the backbone network to extract basic features h. base It is also passed as a basic feature vector to multiple downstream MLP branches and used to predict multiple output components, including volume density, color, shadow, ambient light and uncertainty coefficient.
[0074] Based on this, we designed a lightweight residual branch that utilizes the basic features h extracted from the backbone network. base And the encoded location information, and extract residual features h through a lightweight MLP network. res Similar to the backbone network, h resIt is also used to predict five components, including volume density. This residual branch is essentially based on the existing basic feature h. base With the assistance of the backbone network, the scene is remodeled. Unlike the backbone network, which builds the overall representation from scratch, the residual branch builds upon existing basic features h. base Guided by the core principles, targeted supplementary modeling is performed, focusing on supplementing and optimizing detailed information. By combining encoded location information and basic features, it can learn more refined local features based on existing structure perception. This is achieved through a feature backflow mechanism to alleviate feature degradation, retaining key information extracted by the backbone network while further learning scene details and enhancing the model's ability to depict local details. This module not only performs targeted refinement modeling based on the output of the backbone network but also establishes an effective connection between global and local features by sharing input information such as location encoding with the backbone. This effectively improves the accuracy of prediction results and the ability to restore details, thereby enhancing the accuracy and robustness of the overall 3D reconstruction.
[0075] Finally, the basic predictions and residual compensations are added together to form the final output, enabling the network to better establish connections between the global structure and local details.
[0076] 1) Basic component generation module
[0077] The basic component generation module adopts the Sat-NeRF network structure, such as the Sat-NeRF network model disclosed in the 2022 method for 3D reconstruction of satellite images based on neural radiation fields proposed by Marí et al., which includes a backbone network and a first multilayer perceptron (MLP).
[0078] The backbone network can be composed of multiple fully connected layers. It is initialized with SIREN to enhance representation capabilities. First, the position vector and ray direction vector are position-encoded. Then, the position-encoded position vectors are processed to extract basic features h. base Then, for the basic feature h base The process generates a basic component volume density; this first multilayer perceptron (MLP) uses the basic features h base The input consists of the sunlight direction vector, the time embedding vector, and the position-encoded ray direction vector. The output consists of the basic components: color, uncertainty coefficient, shadow grayscale value, and ambient light.
[0079] We can follow the input and output of the existing Sat-NeRF network model, with the mapping function as follows:
[0080]
[0081] Where x represents the location of the sampling point, d represents the direction of the light rays at the sampling point, ω represents the direction of sunlight, t represents the embedding vector of the time encoding, c represents the color, σ represents the volume density, β represents the uncertainty coefficient, s represents the shadow gray value, and sky represents the ambient light.
[0082] In the basic component generation module, the position encoding mapping first maps the position of the sampling point and the direction of the ray to a higher dimension:
[0083] γ(x)=concat(sin(2 0 πx),cos(2 0 πx),...,sin(2 9 πx),cos(2 9 πx))
[0084] γ(d)=concat(sin(2 0 πd),cos(2 0 πd),...,sin(2 3 πd),cos(2 3 πd))
[0085] Then, after processing through 8 fully connected layers, the basic feature h is obtained. base :
[0086]
[0087] Then, calculate the five basic components based on the basic characteristics:
[0088]
[0089] in, These represent the fully connected networks in the first multilayer perceptron (MLP). This indicates a three-layer fully connected network, while the others indicate a one-layer fully connected network.
[0090] 2) Residual parameter generation module
[0091] This residual parameter generation module directly passes shallow features to deep feature maps via residual connections, enabling the network to learn more feature representations, fully utilizing feature information while maintaining the accuracy of geometric information. It mainly includes residual blocks and a second multilayer perceptron (MLP); it concatenates basic features with position-encoded position vectors, using this as input to the residual blocks to extract residual features h. res And for the residual characteristics h res The process generates the residual component volume density; this second multilayer perceptron (MLP) uses the residual feature h resThe input consists of the sunlight direction vector, the time embedding vector, and the position-encoded ray direction vector. The output consists of five residual components: color, uncertainty coefficient, shadow grayscale value, and ambient light.
[0092] First, the basic features are concatenated with the encoded location coordinates:
[0093] z = concat(γ(x), h base )
[0094] Then feature enhancement is performed using residual blocks:
[0095] h res =F res2 (ReLU(F res1 (z)))
[0096] Among them, F res1 It is a two-layer fully connected network, F res2 It is a fully connected layer, h res It is a residual characteristic.
[0097] Then, the output is mapped through a second multilayer perceptron (MLP):
[0098]
[0099] This fusion module is used to add the five residual components to the five basic components and fuse them accordingly, outputting the five final components: color, uncertainty coefficient, shadow grayscale value, and ambient light. The specific calculation formula is as follows:
[0100] I final =I base +I res
[0101] Among them, I base Take c respectively base ,σ base ,β base ,s base sky base I res Take c respectively res ,σ res ,β res ,s res sky res .
[0102] 4) Construction of the total loss function
[0103] 4.1 Monocular Depth Supervision Loss Function
[0104] While state-of-the-art neural radiation field methods can achieve high-quality reconstruction of simple scenes under multi-view conditions, their performance significantly degrades when dealing with larger-scale, more complex, or sparser scenes. This problem mainly stems from relying solely on color consistency to optimize the model (i.e., supervision through the difference between predicted and actual image colors), lacking explicit geometric constraints in areas with missing textures or fewer views, leading to inaccurate reconstruction results. Existing Sat-NeRF networks use sparse 3D points generated by incorporating image features through bundled adjustments as depth supervision signals. However, the number of these sparse points is limited and their distribution is uneven, making it difficult to support accurate estimation of dense depth. In contrast, dense geometric supervision, as effective prior information, has greater potential to improve the accuracy of geometric structure. To address this, this invention employs an advanced monocular depth estimation model, Depth Anything V2, to estimate depth from satellite imagery and combines it with a scale alignment method to provide coarse but globally consistent dense depth guidance, thereby enhancing geometric constraints in the DSM elevation map generation process and making the prediction results closer to the real terrain.
[0105] For a training image I, the Depth Anything v2 model F Θ Output dense depth map D ref Note that in large-scale scenes, the absolute scale is difficult to estimate, so it is necessary to consider D as a constant. ref This serves as a relatively deep reference clue.
[0106] Meanwhile, the predicted depth map D is obtained by calculating the output components of the RDS neural network using volume rendering technology. pred With D ref There is a scale inconsistency between the two, and directly calculating their differences will lead to errors and lack physical meaning. Therefore, we adopt a scale-offset alignment method to scale the two depths so that they are in the same scale space.
[0107] The following formula can be used to analyze the dense depth map D. ref Perform an alignment operation to align it with the predicted depth map D. pred Alignment:
[0108] D aligned =pD ref +q
[0109] Where p is the scale factor and q is the translation offset.
[0110] Solve for the scale p and offset q using the least squares method:
[0111]
[0112] Where p is the scale factor, representing D pred Compared to D ref The scaling factor, q is the translation offset, representing D pred The overall offset, D ref It is the depth output by the Depth Anything v2 model, A = [D ref [1] is the scale alignment matrix, based on D ref Constructed.
[0113] After completing the scale alignment, this invention calculates the aligned D. ref With D pred The mean squared error (MSE) between the two sides is defined as follows: The deep supervision loss function is defined as:
[0114]
[0115] Among them, L MD (R) represents the depth supervision loss, where R represents the set of pixels contained in the satellite image, r represents any pixel, p is the scale factor, and q is the translation offset.
[0116] 4.2 Solar Correction Loss Function and Uncertainty Weighted Loss Function
[0117] The RDS-NeRF network model uses a solar correction loss function similar to S-NeRF and Sat-NeRF. Since the shadow scalar may produce unreasonable responses to sunlight directions not present in the training, a regularization term is introduced to constrain the shadow representation under solar orientation. This term uses transmittance Ti and opacity αi to supervise the learning of the shadow scalar, improving the model's ability to model shadows.
[0118]
[0119] Among them, L SC (R SC R represents the solar correction loss. SC It is a collection of sunlight rays, where r represents each sunlight ray, and N... SC T represents the number of sampling points on each ray. i s is the cumulative transmittance at the i-th sampling point. i It is the grayscale value of the shadow at the i-th sampling point, α i It is the opacity of the i-th sampling point.
[0120] The RDS-NeRF network model also uses an uncertainty-weighted loss similar to NeRF-W, which uses the uncertainty coefficient β to weight the color reconstruction error. This reduces the interference of transient objects or changing regions on training and improves the model's robustness to dynamic scenes.
[0121]
[0122] β′(r)=β(r)+β min ,
[0123] Among them, L RGB (R) represents the uncertainty-weighted loss, R represents the set of rays, r represents each ray, c(r) represents the predicted color of each ray, and c gt (r) represents the true color of each pixel, T i It is the cumulative transmittance at the i-th sampling point, α i β is the opacity of the i-th sampling point. min Take 0.05, β(x) i ,t) represents the i-th sampling point x on ray r. i At that point, combined with the transient embedding vector t, the uncertainty coefficient is predicted, where N represents the total number of discrete points sampled along ray r.
[0124] The final total loss function consists of three parts: uncertainty-weighted loss L RGB Solar correction item L SC and depth of supervision loss L DS Its overall expression is:
[0125] L = L RGB +λ SC L SC +λ MD L MD
[0126] Among them, L RGB L represents the uncertainty-weighted loss function. SC L represents the solar correction term function. MD Let λ represent the deep supervision loss function. SC , λ MD Both represent hyperparameters.
[0127] Step 4: Using volume rendering technology, process the final components corresponding to the sampling points contained in each pixel to obtain the color and depth information corresponding to each pixel, and convert them into three-dimensional coordinates to generate a DSM elevation map, thus completing the three-dimensional reconstruction of the reconstructed area.
[0128] 1. Use volumetric rendering technology to calculate the color and depth information corresponding to each pixel.
[0129] Based on the final components output by the RDS-NeRF network model, the color c of each sampling point on the light ray is calculated using the shadow-sensing infrared radiation model proposed in S-NeRF:
[0130] c = c final ·(s final +(1-s final )·sky final )
[0131] Among them, c final It is the final color component predicted by the RDS-NeRF network model for each sampling point, s final It is the final shadow grayscale value component predicted by the RDS-NeRF network model for each sampling point, sky final It is the final ambient light component predicted by the RDS-NeRF network model for each sampling point.
[0132] Then, the colors and densities of all sampled points on the light source are combined to calculate the color value C corresponding to the pixel:
[0133]
[0134] α i =1-exp(-σ i δ i )
[0135] Where, δ i =t i+1 -t i t represents the distance between adjacent sampling points i and i+1. i T represents the distance from the camera origin to sampling point i. i α represents the cumulative transmittance at sampling point i. i c represents the opacity at sampling point i. i The color at sampling point i is represented by , and N represents the total number of discrete points sampled along ray r.
[0136] Meanwhile, the depth value z observed for each ray can be expressed in a similar way to the color value C.
[0137]
[0138] Where N represents the total number of sampling points, T i α represents the cumulative transmittance at sampling point i. i The opacity at sampling point i is represented by t. i This represents the distance from sampling point i to the origin of the ray.
[0139] Using the depth value of each ray, along with its origin and direction, the 3D position of each pixel is calculated. The normalized 3D coordinates of the predicted point are then:
[0140]
[0141] Where o represents the origin of ray g, d represents the direction of ray g, and z represents the depth of ray g. It is then converted to actual spatial coordinates in the Earth-centered Earth-fixed coordinate system ECEF through inverse normalization. Next, the actual spatial coordinates are converted into geographic coordinates, namely latitude φ, longitude λ, and elevation h;
[0142]
[0143] Where ecef_to_latlon is the geographic transformation function;
[0144] Finally, the geographic coordinates of the points are projected into the east and north coordinates in the UTM plane coordinate system. The spatial extent of the DSM is further determined based on the ROI boundary and resolution. Then, the plyflatten tool is used to spatially interpolate the point cloud to generate the DSM elevation map.
[0145] To comprehensively evaluate the performance of the RDS-NeRF network model under different complex scenarios, this invention conducted a series of comparative experiments, comparing it with mainstream methods such as S-NeRF, Sat-NeRF, Bungee-NeRF, and Season-NeRF. Furthermore, it analyzed the experimental results of each method in building, road, water, and vegetation scenarios based on multiple evaluation metrics. Specifically, it used five metrics—MAE, RMSE, PAG<2.5m, PAG<7.5m, and integrity—to evaluate the quality of the final DSM, and two metrics—PSNR and SSIM—to evaluate the quality of the synthesized new view.
[0146] 1. Scenes featuring buildings, roads, waterways, and vegetation
[0147] Table 2. Evaluation of DSM accuracy and new view synthesis quality of different methods in architectural scenes.
[0148]
[0149] (1) Architectural Scene
[0150] Region JAX_068 contains large, iconic buildings; region JAX_156 features buildings with repetitive textures; regions JAX_214 and JAX_251 have dense buildings with significant elevation differences. We compared the performance of the four methods in DSM accuracy and new perspective image synthesis quality across four regions with varying building density and structural complexity, as shown in Table 2. The DSM results for the four regions of the architectural scene are as follows: Figure 3 As shown. From Figure 3The results show that S-NeRF reconstructs elevations in repetitive texture areas with significant discrepancies from the actual elevations. Sat-NeRF outperforms S-NeRF on some metrics, showing improved reconstruction quality. Bungee-NeRF performs relatively poorly, struggling to accurately reconstruct building geometry and details. Season-NeRF can detect major buildings, but building boundaries are blurred. While it achieves the highest completeness, covering the entire area, its elevation accuracy is insufficient, and building geometric details are oversimplified. RDS-NeRF achieves the best results in all regions, preserving clear structure and uniform height in building edges and flat areas. In terms of DSM error metrics, RDS-NeRF has an average MAE of 1.46 and an average RMSE of 2.83. It also leads in the proportion of elevation errors in PAG < 2.5m and in completeness, with an average <2.5m error rate of 85.76% and an average Comp. percentage of 95.32%. Furthermore, its synthesized images also outperform other models in PSNR and SSIM metrics. In the four regions, RDS-NeRF outperforms other regions in all metrics in the JAX_156 region.
[0151] While Bungee-NeRF introduces residual structures to progressively capture higher-frequency details, it relies solely on RGB loss without any direct or indirect geometric supervision signals. In areas with weak or repetitive textures, the RGB loss is insufficient to constrain depth accuracy, limiting its ability to recover geometric structures. Secondly, in satellite imagery, object shadows not only affect image appearance but also directly impact geometric reconstruction accuracy. Since it doesn't incorporate lighting factors like solar elevation angle for modeling, it struggles to handle geometric ambiguities caused by shadows, making it difficult to effectively adapt to high-precision, high-consistency DSM reconstruction tasks. Season-NeRF employs a two-stage training approach, initially relying on height maps obtained through Space Carving to guide density learning. However, these height maps themselves contain significant noise. Although this supervision signal is removed later, the initial noise interferes with the network's convergence path, resulting in less refined overall geometric structures. Therefore, while height map-guided training can improve coverage completeness, it doesn't necessarily contribute to refined modeling. Secondly, while Season-NeRF is designed to handle seasonal variations, data from different times exhibit visual shadow variations due to the lack of explicit modeling of lighting or semantic changes, which affects geometric consistency. Consequently, it performs worse than other methods in terms of geometric accuracy, edge sharpness, and elevation consistency.
[0152] Table 3. Evaluation of DSM accuracy and new view synthesis quality of different methods in road scenes.
[0153]
[0154]
[0155] (2) Road Scene
[0156] Region JAX_079 is characterized by regular roads and minimal occlusion; region JAX_113 contains multiple interchange structures, exhibiting occlusion and elevation variations; region JAX_159 features dense roads and a complex scene; while region JAX_204 is a large-scale highway area with a wide viewing angle. We compared the performance of four methods in DSM accuracy and new perspective image synthesis quality across these four regions with different road characteristics and surrounding environmental complexities, as shown in Table 3. The DSM results for the four road scene regions are as follows: Figure 4 As shown. From Figure 4 The results show that S-NeRF can reproduce terrain features well in flat areas, but exhibits significant discontinuities and blurred edges in complex structures such as overpasses and abrupt elevation changes. Sat-NeRF shows good integrity in some areas, but falls short in areas with severe occlusion or complex details. Bungee-NeRF and Season-NeRF are generally blurry and inaccurate in representing elevation changes. In contrast, RDS-NeRF displays clearer terrain structures and geometric details in all four regions, achieving the best overall visual quality. Regarding DSM error metrics, RDS-NeRF has an average MAE of 1.29 and an average RMSE of 2.06, significantly lower than the other methods. It leads in the proportion of elevation errors (PAG < 2.5m) with an average of 85.33% and in terms of completeness with an average of 97.72%. Furthermore, its synthesized images also have relatively high PSNR and SSIM metrics, indicating advantages in new view synthesis. Of the four regions, RDS-NeRF achieved the best accuracy results in the JAX_204 region, reconstructing terrain that is closer to the true value.
[0157] Table 4. DSM accuracy and new view synthesis quality assessment of different methods in aquatic scenes.
[0158]
[0159] (3) Water Scene
[0160] Regions JAX_020 and JAX_149 are inland water areas with some vegetation and buildings. Regions JAX_260 and JAX_269 are outer water areas; JAX_260 has a clear boundary between the water and the shore, while JAX_269 has boats and other dynamic objects on the shore. We compared the performance of four methods in DSM accuracy and new perspective image synthesis quality in these four regions with different water features and surrounding environmental complexities, as shown in Table 4. The DSM results for the four water areas are as follows: Figure 5 As shown. From Figure 5The results from S-NeRF reconstruction in the water area are rather blurry, failing to distinguish between water and land features; Sat-NeRF has the same problem; Bungee-NeRF performs even worse, completely failing to identify feature categories. Season-NeRF almost completely fails to identify water body boundaries, losing most water features and exhibiting significant errors in water height. RDS-NeRF, on the other hand, reconstructs both clear water areas and boundaries, and fully reconstructs land features, showing high consistency with the ground truth. Regarding DSM error metrics, RDS-NeRF has an average MAE of 2.28 and an average RMSE of 3.46, showing a significant advantage over methods like Bungee-NeRF. It also has an average of 70.74% of PAG < 2.5m elevation errors, and an average of 94.43% in terms of completeness, placing it at the forefront. Among the four regions, RDS-NeRF performs best on average in the JAX_020 region.
[0161] Table 5. Evaluation of DSM accuracy and new view synthesis quality of different methods in vegetation scenes.
[0162]
[0163]
[0164] (4) Vegetation Scene
[0165] Regions JAX_004 and JAX_028 have high vegetation coverage, and the vegetation distribution exhibits a patchy clustering characteristic; regions JAX_118 and JAX_122 have low-lying, sparsely distributed vegetation with obvious intervals between vegetation patches, failing to form large, continuous vegetation clusters. In these four regions with different vegetation characteristics and surrounding environmental complexities, we compared the performance of four methods in DSM accuracy and new perspective image synthesis quality, as shown in Table 5. The DSM results for the four vegetation scene regions are as follows: Figure 7 As shown. (Combined with Table 5 and...) Figure 7 Analysis shows that in the JAX_004 and JAX_028 regions, S-NeRF and Sat-NeRF exhibit better DSM accuracy, with vegetation heights closer to the true values; while RDS-NeRF shows only moderate DSM accuracy. Figure 6While RDS-NeRF reconstruction of vegetation height also exhibits some errors, it provides clearer and more complete reconstructions of surrounding features such as buildings, roads, and water bodies compared to S-NeRF and Sat-NeRF. In the JAX_118 and JAX_122 regions, the DSM reconstruction results of S-NeRF and Sat-NeRF are comparable to RDS-NeRF, but their DSM accuracy is slightly lower. Bungee-NeRF generally performs poorly across all regions, failing to reconstruct the shape of vegetation at all. Season-NeRF can basically identify vegetation, but its ability to reconstruct details is limited, internal texture is insufficient, and vegetation height variations are not clearly represented. RDS-NeRF demonstrates the best DSM accuracy and also performs excellently in PSNR and SSIM metrics. Among these four regions, RDS-NeRF performs best overall in the JAX_118 region.
[0166] 2. Ablation test
[0167] To verify the contributions of the proposed Residual Vector Generation (RFE) module and Monocular Depth Supervision (MDS) mechanism to model performance, we designed ablation experiments. New components were added to the base model to evaluate their impact on DSM accuracy and rendering quality. We selected one region in each of four scenarios for the experiments and used the same evaluation metrics as the main experiment. The quantitative evaluation results are shown in Table 6.
[0168] Table 6. Quality assessment of DSM and new view synthesis for different module combinations in different scenarios.
[0169]
[0170]
[0171] In Table 6, JAX_068, JAX_079, JAX_020, and JAX_004 correspond to building, road, water, and vegetation scenes, respectively. In these four scenes, different module combinations (+RFE, +MDS) all improved the accuracy and completeness of the DSM compared to the base model, indicating that both designs effectively enhance geometric reconstruction capabilities. Among them, +RFE showed a more significant improvement in DSM accuracy. This is because +RFE, guided by basic features, provides targeted residual information supplementation, focusing on details that are easily overlooked or difficult to accurately represent in the backbone network, optimizing and refining them, thereby improving the model's geometric accuracy and surface reconstruction quality. This makes it unrestricted by scene type; whether it's a regular building, road, or water scene, or a complex and irregular vegetation scene, it can bring significant accuracy improvements.
[0172] In terms of completeness, +MDS performs better, indicating that introducing monocular depth as a global prior enhances the model's perception and completion of the overall terrain contour. Especially in water scenes, +MDS outperforms +RFE in both DSM accuracy and completeness. This is because water scenes have weak textures and indistinct features, making residual enhancement difficult to effectively extract high-frequency information and prone to introducing errors. Monocular depth, as a global guide, better constrains the geometry of low-texture areas, thereby improving the reconstruction quality of water scenes.
[0173] Furthermore, we found that the complete model integrating both modules had the highest completeness in all four scenarios. In the building, road, and water scenarios, the accuracy of DSM was also better than that of the individual module combinations. This indicates that +RFE and +MDS are complementary in terms of detailed modeling and global constraints, and their combined use can improve the model's adaptability and robustness to diverse scenarios.
[0174] However, in vegetation scenes, the DSM accuracy of the complete model is slightly lower than that of the base model. This is because monocular depth estimation tends to group regions with similar textures to the same height. Vertical or flat structures like buildings, roads, and water bodies can effectively utilize monocular depth estimation to improve terrain accuracy and completeness. However, for vegetation, such as… Figure 7 As shown, since satellite imagery is taken from above, only the tree canopy is visible. The large and dense canopy obscures the tree trunk, making it difficult to obtain accurate height and spatial structure information of vegetation from monocular depth estimation in satellite imagery. Furthermore, the obtained depth information contains significant errors. When these errors are fed back into the training process as supervisory signals, they may interfere with model learning, leading to a decrease in DSM accuracy. Therefore, for densely canopied vegetation, introducing a monocular depth supervision module does not necessarily bring positive benefits and may even negatively impact accuracy. This can be verified by the results in regions JAX_004 and JAX_028 in Table 5. In regions JAX_118 and JAX_122, the vegetation is low and sparsely distributed, similar to... Figure 7 In the case of the small tree on the left side, depth supervision still effectively improves the accuracy of DSM.
[0175] Figure 7 Schematic diagrams of vegetation from different viewing angles. From a satellite perspective, satellite imagery can only present the planar outline of tree canopies when acquiring vegetation information, making it difficult to show the three-dimensional structure of vegetation. From a direct viewpoint, the true height of trees can be clearly seen, but because the dense canopy obscures the trunk, only a portion of the height can be estimated from satellite images; that is, the depth estimate is often lower than the true height.
[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for three-dimensional reconstruction of neural radiation fields that integrates residual feature enhancement and depth supervision, characterized in that... Includes the following steps: Step 1: Acquire multi-temporal satellite images of the reconstructed area; Step 2: For any time-phase satellite image, preprocessing is performed first, and then ray sampling is performed on any pixel to obtain the feature vector corresponding to each sampling point; Step 3: Input the feature vector corresponding to each sampling point into the RDS-NeRF network model to infer the final component corresponding to each sampling point; The RDS-NeRF network model uses a Sat-NeRF network structure to extract basic features, processes the basic features, outputs basic components, then uses residual blocks to enhance the basic features, processes the enhanced basic features, outputs residual components, and finally adds and fuses the residual components with the basic components to output the final component. Step 4: Using volume rendering technology, process the final components corresponding to the sampling points contained in each pixel to obtain the color and depth information corresponding to each pixel, and convert them into three-dimensional coordinates to generate a DSM elevation map, thus completing the three-dimensional reconstruction of the reconstructed area.
2. The method for three-dimensional reconstruction of neural radiation fields by fusing residual feature enhancement and depth supervision according to claim 1, characterized in that: The basic components, residual components, and final components are all set to five, namely color, volume density, uncertainty coefficient, shadow gray value, and ambient light. The feature vectors are set to four, namely the position vector, ray direction vector, sunlight direction vector, and time embedding vector corresponding to each sampling point. The RDS-NeRF network model includes a basic component generation module, a residual component generation module, and a fusion module. The basic component generation module adopts a Sat-NeRF network structure, including a backbone network and a first multilayer perceptron (MLP). The residual component generation module includes a residual block and a second multilayer perceptron (MLP). The backbone network first encodes the position vector and the ray direction vector, then processes the encoded position vector to extract the basic feature h. base Then, for the basic feature h base The process is performed to generate the basic volume density; The first multilayer perceptron (MLP) uses basic features h base The input consists of the sunlight direction vector, the time embedding vector, and the position-encoded ray direction vector. The output components are color, uncertainty coefficient, shadow grayscale value, and ambient light. The basic features are concatenated with the position vector after position encoding, and this is used as the input to the residual block to extract the residual features h. res And the residual characteristics h res The process is performed to generate the residual component volume density; The second multilayer perceptron (MLP) uses residual features h res The input consists of the sunlight direction vector, the time embedding vector, and the position-encoded light direction vector. The output consists of five residual components: color, uncertainty coefficient, shadow grayscale value, and ambient light. The fusion module is used to add and fuse the five residual components with the five basic components, and output the five final components: color, uncertainty coefficient, shadow grayscale value, and ambient light.
3. The method for three-dimensional reconstruction of neural radiation fields by fusing residual feature enhancement and depth supervision according to claim 2, characterized in that: Let c be the fundamental components: color, uncertainty coefficient, shadow grayscale value, and ambient light. base ,σ base ,β base ,s base sky base The residual component color, uncertainty coefficient, shadow grayscale value, and ambient light are respectively c res ,σ res ,β res ,s res sky res The position vector, ray direction vector, sunlight direction vector, and time embedding vector of each sampling point are x, d, ω, and t, respectively. The backbone network is configured as eight fully connected networks, each with a scale of h. The first multilayer perceptron (MLP) and the second multilayer perceptron (MLP) have the same structure. The residual characteristics are calculated using the following formula. h res =F res2 (ReLU(F res1 (z))) z=concat(γ(x),h base ) Among them, F res1 Indicates a two-layer fully connected network, F res2 Let represent a fully connected layer, and γ(x) represent the position vector after position encoding. Calculate the five fundamental components using the following formula. in, These represent the fully connected networks in the first multilayer perceptron (MLP). represents a three-layer fully connected network, and the rest represent a single-layer fully connected network. γ(d) represents the position-encoded ray direction vector. Calculate the five residual components using the following formula. in, These represent the fully connected networks in the second multilayer perceptron (MLP). represents a three-layer fully connected network, and the rest represent a single-layer fully connected network. γ(d) represents the position-encoded ray direction vector. Calculate the five final components I using the following formula. final , I final =I base +I res Among them, I base Take c respectively base ,σ base ,β base ,s base sky base I res Take c respectively res ,σ res ,β res ,s res sky res .
4. The method for three-dimensional reconstruction of neural radiation fields by fusing residual feature enhancement and depth supervision according to claim 2, characterized in that: The total loss function of the RDS-NeRF network model is set as follows: L=L RGB +λ SC L SC +λ MD L MD (R) Among them, L RGB L represents the uncertainty-weighted loss function. SC L represents the solar correction term function. MD Let λ represent the deep supervision loss function. SC , λ MD All represent hyperparameters, R represents the set of pixels contained in the satellite image, r represents any pixel, p represents the scale factor, q represents the translation offset, and D represents the hyperparameter. ref This represents the dense depth map output by the Depth Anything v2 model, D. pred This indicates that the final components output by the RDS-NeRF network model are processed using volume rendering technology to obtain the predicted depth map, where A represents the alignment matrix.
5. The method for three-dimensional reconstruction of neural radiation fields by fusing residual feature enhancement and depth supervision according to claim 1, characterized in that: In step two, when preprocessing any temporal satellite image, the maximum elevation value h corresponding to the satellite image is first obtained from the image metadata. max and minimum elevation value h min Simultaneously, the solar direction vector is calculated based on the solar altitude angle and azimuth angle, and its PRC parameters are adjusted and optimized. Then, the maximum elevation value h is used as the basis for the calculation. max Minimum elevation value h min Using the optimized PRC parameters as input, ray sampling is performed on any pixel in the satellite image to obtain the position vector and ray direction vector of each sampling point. Finally, the position vector, ray direction vector, sunlight direction vector, and time embedding vector are used as the feature vector of each sampling point.
6. The method for three-dimensional reconstruction of neural radiation fields by fusing residual feature enhancement and depth supervision according to claim 1, characterized in that: In step four, the depth information of each pixel is converted into normalized three-dimensional coordinates using the following formula. Where o represents the origin of ray g, d represents the direction of ray g, and z represents the depth information of the pixel corresponding to ray g; Subsequently, it was converted into actual spatial coordinates in the Earth-centered Earth-fixed coordinate system ECEF through inverse normalization. Then, use the following formula to convert the actual spatial coordinates into geographic coordinates, including latitude. Longitude λ g and elevation h g , Where ecef_to_latlon is the geographic transformation function; Finally, the geographic coordinates are projected into the east and north coordinates in the UTM plane coordinate system. The spatial extent of the DSM is further determined based on the ROI boundary and resolution. Then, the plyflatten tool is used to spatially interpolate the point cloud to generate the DSM elevation map.
Citation Information
Patent Citations
Large-scale building three-dimensional reconstruction method fusing multi-scale image features
CN116402942A
Neural radiation field-based few-view multi-temporal remote sensing image three-dimensional reconstruction method
CN117765185A
Mine image three-dimensional reconstruction method based on neural radiation field
CN118279490A
Light coding method and device for projection three-dimensional display, medium and product
CN118474333A
Deep learning three-dimensional climate simulation method for adding climate to real scene
CN120088383A