Automated 3D Reconstruction Method for Urban Components from Sub-meter High-Resolution Remote Sensing Imagery

By combining multi-source image fine correction and dense point cloud generation with semantic NeRF network and vector data fusion, the problems of large image registration error and sparse point cloud with a lot of noise in traditional methods are solved, and high-precision automated 3D reconstruction of urban components is realized.

CN120747387BActive Publication Date: 2025-11-14SHAANXI TIRAIN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511262880.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-14
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Traditional automated 3D reconstruction methods for urban components from sub-meter resolution remote sensing images are prone to local distortion under complex terrain or multi-temporal image conditions. Stereo matching methods are not robust to occluded areas and areas with poor texture. Point clouds contain holes and outliers, making it difficult to meet the requirements for high-precision and automated reconstruction.

Method used

Through multi-source remote sensing image fine correction, dense point cloud generation, semantic NeRF enhancement, and vector data fusion, an initial ground-image projection mapping relationship is established using a rational polynomial coefficient model. Layered factor map optimization is performed by combining ground control points and external multi-source data to perform image registration. Dense point clouds are generated by fusing LiDAR point clouds and image features, and the semantic information of the point clouds is enhanced by a semantic neural radiation field network.

Benefits of technology

It achieves high-precision, automated 3D reconstruction of urban components, improves the accuracy and stability of image registration, reduces the proportion of point cloud noise and outliers, and ensures the geometric fidelity and continuity of point clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747387B_ABST
    Figure CN120747387B_ABST
Patent Text Reader

Abstract

This invention discloses an automated 3D reconstruction method for urban components using sub-meter resolution high-resolution remote sensing imagery, relating to the field of image processing technology. The method includes the following steps: acquiring sub-meter resolution high-resolution remote sensing imagery of the target area; establishing an initial ground-image projection mapping relationship based on a rational polynomial coefficient model; acquiring ground control points; correcting the mapping relationship; outputting a registered image; performing multi-view geometric matching on the registered imagery to generate a dense surface point cloud, and inputting it into a semantic neural radiation field network to generate another dense point cloud; performing cross-modal feature alignment and fusion of the dense point cloud and vector data; and outputting a fused multi-source dense point cloud. This application solves the problems of large registration errors, sparse point clouds with high noise, and semantic missing information through multi-source remote sensing image fine correction, achieving high-precision, automated 3D reconstruction of urban components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to a method for automated 3D reconstruction of urban components from sub-meter resolution remote sensing images. Background Technology

[0002] In the process of urban digital transformation and smart city construction, the construction of high-precision 3D models has become an important support for applications such as urban planning, infrastructure management, disaster monitoring, traffic simulation, and digital twin cities. However, traditional urban 3D reconstruction methods mostly rely on manual surveying, low-resolution aerial imagery, or single-source LiDAR point clouds. These methods suffer from high costs, limited coverage, long update cycles, and a lack of semantic information, making it difficult to meet the current urgent needs for refined urban modeling and intelligent understanding. With the rapid development of remote sensing technology, sub-meter high-resolution remote sensing imagery, due to its advantages of high resolution, wide acquisition range, and strong timeliness, has gradually become an important data source for urban 3D reconstruction. However, several technical challenges remain in the process of using sub-meter high-resolution remote sensing imagery for automated 3D reconstruction of urban components.

[0003] For example, the invention patent with publication number CN113343346B discloses a rapid modeling method for 3D traffic scenes based on high-precision maps, including the following steps: S1: geometrically correcting remote sensing images and using a two-dimensional polynomial correction model for correction; S2: extracting building boundaries and roof feature lines from the geometrically corrected remote sensing images to generate feature point chains; setting thresholds to remove adjacent point chains, generating building feature lines, and outputting building vector data; S3: fusing the building vector data into the high-precision map vector data using a one-dimensional polynomial correction model; S4: constructing a data encoding mapping dictionary table and integrating classification information into the high-precision map vector data; S5: using a rule-based modeling method to batch construct 3D models of various traffic elements and buildings along the route; S6: using a refined modeling method to build a refined traffic element model; S7: constructing a 3D traffic scene with individual characteristics based on the modeling results of steps S5 and S6.

[0004] For example, the invention patent with publication number CN120471962A discloses a method and apparatus for three-dimensional reconstruction of urban buildings. The method includes: acquiring multi-view satellite orthophotos from the same scene, satellite-borne laser altimetry data of the target area, a digital terrain model, and two-dimensional building vector data. Precise geometric registration is performed on the multi-view satellite orthophotos, and dense matching is executed to generate an initial disparity point cloud. An accurate planar positioning disparity map is constructed by combining the azimuth information of the multi-view orthophotos. The elevation values ​​above the ground surface are calculated using the satellite-borne laser altimetry data and the digital terrain model. The disparity elevation ratio coefficient is estimated by jointly using the elevation values ​​and disparity values ​​to convert the disparity map into a height map. Three-dimensional building vector data of the target area is generated by combining the two-dimensional building vector data. This solves the problem in related technologies where the cost of mapping-grade remote sensing data is high and data acquisition is limited, making it unable to meet the global demand for three-dimensional urban reconstruction.

[0005] The above-disclosed technical solutions have at least the following technical problems:

[0006] Traditional automated 3D reconstruction methods for urban components based on sub-meter resolution remote sensing images typically only perform one-time global optimization for image correction using RPC. This can easily lead to local distortions under complex terrain or multi-temporal image conditions. Furthermore, traditional stereo matching methods are not robust to occluded areas and areas with poor texture, and point clouds contain many holes and outliers.

[0007] To address the above problems, this invention proposes a solution. Summary of the Invention

[0008] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an automated 3D reconstruction method for urban components from sub-meter resolution remote sensing images. By performing multi-source remote sensing image fine correction, dense point cloud generation, semantic NeRF enhancement, and vector data fusion, the method solves the problems of large registration errors, sparse point clouds with high noise, and semantic missingness in traditional methods, thereby achieving high-precision and automated 3D reconstruction of urban components.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] An automated 3D reconstruction method for urban components from sub-meter resolution high-resolution remote sensing imagery includes: acquiring sub-meter resolution high-resolution remote sensing imagery of the target area; establishing an initial ground-image projection mapping relationship based on a rational polynomial coefficient model; acquiring ground control points; correcting the mapping relationship; outputting a registered image; performing multi-view geometric matching on the registered imagery to generate a dense surface point cloud, and inputting it into a semantic neural radiation field network to generate a dense point cloud; and performing cross-modal feature alignment and fusion of the dense point cloud and vector data to output a fused multi-source dense point cloud.

[0011] In a preferred embodiment, the acquisition of sub-meter high-resolution remote sensing images of the target area and the establishment of an initial ground-image projection mapping relationship based on a rational polynomial coefficient model are as follows: Multiple sub-meter high-resolution remote sensing images of the target area and their accompanying rational polynomial coefficient files are acquired, and the projection coordinate system of the images is unified with the geographic coordinate system and elevation datum of the elevation reference to form sub-meter high-resolution remote sensing images and a corresponding initial set of RPC functions; data features are extracted from the sub-meter high-resolution remote sensing images, filtered, and the optimal stereo combination is selected based on the information gain criterion; the RPC function set is analyzed, a projection function is established, and the Jacobian is output. The system employs a matrix and a differentiable projection operator. It divides the image region into sub-blocks based on the terrain slope and projection linearization error of the target area, generating multi-scale pyramid images within each sub-block. Ground point coordinates are sampled within each sub-block, and projected onto the pixel planes of the optimal stereo combination using a differentiable projection operator to obtain the corresponding pixel coordinates. The error covariance of each pixel is calculated using the prior elevation variance and the Jacobian matrix. A covariance matrix is ​​constructed based on the pixel domain covariance, and the corresponding confidence ellipse parameters are derived. The ground point coordinates, pixel coordinates, and confidence ellipse parameters are stored as table entries and organized into a perceptual lookup table based on spatial location indexing.

[0012] In a preferred embodiment, the extraction of data features from sub-meter high-resolution remote sensing images, the screening of these features, and the selection of the optimal stereo combination based on the information gain criterion are as follows: Sub-meter high-resolution remote sensing images are eliminated according to a preset threshold to obtain a candidate image set; features such as imaging time, off-axis angle, overlap, and cloud cover are used as input variables to construct a comprehensive evaluation function, which is then mapped to the quality score of the candidate images; the contribution of each feature to the final quality score is calculated based on the quality score using the information gain criterion; the features are sorted according to their information gain, and the sorting results are normalized to obtain the weight value of each feature within a fixed interval; the comprehensive evaluation function of all candidate images is weighted and solved according to the determined feature weights to obtain the corresponding weighted comprehensive score; the candidate images are sorted from highest to lowest weighted comprehensive score, and the image pair with the highest comprehensive score is selected as the optimal stereo combination.

[0013] In a preferred embodiment, the steps of acquiring ground control points, correcting the mapping relationship, and outputting the registered image are as follows: Acquire the known 3D coordinates of the ground control points in the geographic coordinate system, and calculate the predicted pixel coordinates of the 3D coordinates in the optimal stereo combination using a differentiable projection operator; extract the difference between the observed pixel coordinates and the predicted pixel coordinates of the ground control points in the optimal stereo combination, and output the ground control point projection residuals; acquire external multi-source data, project the corresponding 3D geographic coordinates onto the optimal stereo combination using a differentiable projection operator, and calculate the difference with the observed pixel coordinates to form observation constraint residuals; combine the ground control point projection residuals with the observation constraint residuals to generate a projection residual set; parameterize the bias field of the RPC model as nodes to be optimized, and construct a perception map optimization energy function using the projection residuals as factor constraints; use the perception map optimization energy function as input, and perform hierarchical optimization based on the projection residual set and the RPC bias field parameterized nodes; reproject the optimal stereo combination onto a unified geographic coordinate system according to the hierarchically optimized RPC, and output the error-compensated registered image.

[0014] In a preferred embodiment, the step of using the perceptual map optimization energy function as input and performing hierarchical optimization based on the projection residual set and RPC bias field parameterized nodes is as follows: The perceptual map optimization energy function is used as the optimization objective; the variables to be optimized are divided according to their scope of action, forming global parameter nodes, local parameter nodes, and temporal parameter nodes; the projection residual set is used as a factor constraint and introduced into the parameter nodes to form a hierarchical factor graph; in the hierarchical factor graph, only the global parameter nodes are activated, and the global parameters of the RPC are coarsely corrected using a sparse nonlinear least squares optimization method; after the global layer optimization is completed, the local parameter nodes are activated, and the local bias field within the sub-block is finely corrected using the factor graph optimization method, and local discontinuities are eliminated using a sub-block boundary consistency factor; under multi-temporal image conditions, the temporal parameter nodes are activated, and a time smoothing constraint term is introduced; during the layer graph optimization process, the residual factor weights are dynamically adjusted using confidence information in the perceptual lookup table, and the RPC bias field parameters are updated using an iterative reweighted least squares method until the energy function converges.

[0015] In a preferred embodiment, the step of performing multi-view geometric matching on the registered images to generate a dense surface point cloud is as follows: acquiring LiDAR point cloud data and projecting it onto the registered image coordinate system to generate the pixel position and depth value of each projected point; applying a convolutional neural network to each registered image to extract depth convolution features; generating depth feature vectors for corresponding pixels in the projected LiDAR point cloud; aligning the image features and LiDAR depth features according to pixel position, and forming a cross-modal feature descriptor through feature concatenation and weighted fusion; constructing a multi-resolution image pyramid to match images at different scales; calculating the matching error using a differentiable stereo matching method at each scale layer, and using the error as a differentiable loss to iteratively optimize the depth value; projecting the optimized depth value onto the geographic coordinate system to generate an initial dense surface point cloud; applying confidence weights and neighborhood consistency constraints to the initial dense surface point cloud for noise filtering, removing outliers, and generating a high-precision dense surface point cloud.

[0016] In a preferred embodiment, the input to the semantic neural radiation field network to generate a dense point cloud is as follows: The dense surface point cloud is locally voxelized, dividing it into several spatial voxel units, each containing several points, and each point is feature-encoded; the features of each point are normalized to generate a point cloud enhancement feature vector; the multi-view image pixels and the point cloud enhancement feature vector are input into the differentiable rendering module of the semantic NeRF network, and the volume density and color contribution of each pixel ray are calculated through ray sampling, and the network is trained through volume rendering error and self-supervised consistency constraints; the trained network predicts the three-dimensional coordinates, color, and semantic probability distribution of each voxel unit point to form a preliminary dense point cloud; low-confidence points and isolated points are preprocessed, and the processed points are global coordinate mapped and formatted to generate the final high-precision dense point cloud.

[0017] In a preferred embodiment, the process of cross-modal feature alignment and fusion of dense point clouds and vector data to output a fused multi-source dense point cloud is as follows: Vector data is acquired; coordinate system transformation is performed on the vector data to place it in a unified geographic coordinate system with the dense point cloud; geometric feature encoding is performed on the vector data to form vector features; local neighborhood feature encoding is performed on the dense point cloud to form point cloud features; the point cloud features and vector features are mapped to the same high-dimensional feature space, and the local features of the dense point cloud are matched with the features of the vector data based on nearest neighbor search or graph matching methods; cross-modal feature data is extracted based on the matching results, and a joint optimization energy function is constructed; the joint optimization energy function is used as the objective to iteratively optimize the 3D coordinates and local feature vectors of the dense point cloud; in each iteration, the point cloud position is adjusted according to the gradient descent method, iterating until the energy function converges; after the iterative optimization converges, the updated 3D coordinates of each point are acquired, and the updated results are stored as a preliminarily optimized dense point cloud; the optimized dense point cloud is fused with the vector data to generate a multi-source dense point cloud.

[0018] In a preferred embodiment, the process of fusing the optimized dense point cloud with vector data to generate a multi-source dense point cloud is as follows: The optimized dense point cloud and corresponding vector data are acquired; a set of constraint points is generated for key urban structures in the vector data; preliminary spatial matching is performed on each point in the dense point cloud to determine its nearest neighbor relationship with the vector constraint points; for each point in the dense point cloud, the spatial distance between it and its nearest neighbor in the constraint point set is calculated as a geometric error index, and the point position is adjusted according to the error to make the dense point cloud geometrically aligned with the vector data; the semantic probability of each point is weighted and updated to make its semantic category distribution consistent with the vector data, achieving semantic fusion; neighborhood consistency filtering is performed on the fused dense point cloud to eliminate local noise and isolated points, and outliers are removed to form the final dense point cloud.

[0019] The technical effects and advantages of the automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery of this invention are as follows:

[0020] 1. This invention establishes an initial ground-image projection relationship based on a rational polynomial coefficient (RPC) model and, combined with ground control points and external multi-source data, employs a hierarchical factor graph optimization method to progressively correct the projection residuals. This method not only enables dynamic adjustment of global and local parameters but also introduces temporal smoothing constraints under multi-temporal image conditions, thereby ensuring the spatial and temporal continuity and consistency of the registered images. This hierarchical optimization mechanism effectively suppresses the local distortion and projection discontinuity problems that occur in traditional single-parameter correction methods, significantly improving the accuracy and stability of image registration.

[0021] 2. This invention achieves high-precision alignment of image features and depth features by fusing registered images with LiDAR point clouds across modalities and stitching together depth convolutional and depth projection features. Supported by a multi-scale image pyramid, a differentiable stereo matching method is used to iteratively optimize depth values, resulting in a high-confidence, dense surface point cloud. Compared to traditional stereo matching or single-modal point cloud reconstruction methods, this invention effectively reduces the proportion of noise and outliers while maintaining point cloud density, ensuring higher geometric fidelity and continuity in the point cloud results. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the process for an automated 3D reconstruction method of urban components from sub-meter resolution remote sensing images according to the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1, Figure 1 This invention presents an automated 3D reconstruction method for urban components from sub-meter resolution high-resolution remote sensing imagery, comprising the following steps:

[0025] S1, acquire sub-meter high-resolution remote sensing images of the target area, and establish the initial land-image projection mapping relationship based on the rational polynomial coefficient model;

[0026] In this embodiment, sub-meter resolution high-resolution remote sensing images of the target area are acquired, and an initial land-image projection mapping relationship is established based on a rational polynomial coefficient model, as follows:

[0027] Acquire multiple sub-meter high-resolution remote sensing images of the target area and their accompanying rational polynomial coefficient (RPC) files, and unify the image's projection coordinate system with the geographic coordinate system (such as WGS84 / UTM) and elevation datum to form sub-meter high-resolution remote sensing images and the corresponding initial RPC function set;

[0028] Data features are extracted from sub-meter resolution high-resolution remote sensing images, filtered, and the optimal stereo combination is selected based on the information gain criterion. The data features include imaging time, off-axis angle, overlap, and cloud cover. Specifically: imaging time is obtained by extracting the shooting timestamp of each image and calculating the imaging time difference between candidate images; off-axis angle is obtained by calculating the angle between the image center point and the normal of the target area based on the satellite attitude and imaging geometric parameters; overlap is obtained by statistically analyzing the overlap ratio between the candidate image coverage area and the target area after geometric registration; and cloud cover is obtained by extracting the image cloud mask using cloud detection algorithms (such as threshold-based spectral discrimination or deep learning segmentation models) and calculating the cloud coverage rate.

[0029] The RPC function set is analyzed, and forward (ground → image pixel) and backward (image pixel → ground) projection functions are established. The Jacobian matrix and differentiable projection operator are output.

[0030] The image region is divided into sub-blocks based on the terrain slope and projection linearization error of the target area, and a multi-scale pyramid image is generated in each sub-block.

[0031] Within each sub-block, the ground point coordinates are sampled, and the ground point coordinates are projected onto the pixel planes of the optimal stereo combination using a differentiable projection operator to obtain the corresponding pixel coordinates. The error covariance of each pixel is then calculated using the prior elevation variance and the Jacobian matrix.

[0032] Construct a covariance matrix based on the pixel domain covariance, and derive the corresponding confidence ellipse parameters (major axis, minor axis, and orientation angle).

[0033] The ground point coordinates, pixel coordinates, and confidence ellipse parameters are stored as table entries and organized into a perceptual lookup table based on spatial location index.

[0034] The data features extracted from sub-meter resolution high-resolution remote sensing images are filtered, and the optimal stereo combination is selected based on the information gain criterion, as follows:

[0035] Image pairs with an imaging time difference greater than a preset threshold are removed to ensure temporal consistency of the image combination.

[0036] Images with excessively large off-axis angles (exceeding the preset angle range) are removed to avoid severe geometric distortion;

[0037] Image pairs with an overlap of less than a threshold are removed to ensure the effectiveness of subsequent stereo matching.

[0038] Images with cloud cover exceeding a set ratio will be removed to ensure a proportion of effective information.

[0039] Based on the above elimination rules, sub-meter resolution high-resolution remote sensing images are eliminated to obtain a candidate image set;

[0040] Using features such as imaging time, off-axis angle, overlap, and cloud cover as input variables, a comprehensive evaluation function is constructed, which is mapped to the quality score of candidate images. The contribution of each feature to the final quality score is calculated using the information gain criterion.

[0041] The features are sorted according to their information gain, and the sorting results are normalized to obtain the weight values ​​of each feature in the [0,1] interval, so as to ensure the comparability of the weight distribution. The information gain refers to the contribution.

[0042] Based on the determined feature weights, the comprehensive evaluation function of all candidate images is solved by weighting to obtain the corresponding weighted comprehensive score;

[0043] Candidate images are sorted from highest to lowest according to their weighted composite scores to ensure that high-quality image pairs are given priority, and the image pair with the highest composite score is selected as the optimal stereo combination.

[0044] S2: Acquire ground control points, correct the mapping relationship, and output registered images;

[0045] In this embodiment, ground control points are acquired, the mapping relationship is corrected, and the registered image is output, as detailed below:

[0046] Obtain the known 3D coordinates of ground control points in the geographic coordinate system, and calculate the predicted pixel coordinates of the 3D coordinates in the optimal stereo combination using a differentiable projection operator;

[0047] Extract the difference between the observed pixel coordinates and the predicted pixel coordinates of the ground control points in the optimal stereo combination, and output the ground control point projection residual.

[0048] External multi-source data is acquired, and the corresponding three-dimensional geographic coordinates are projected onto the optimal stereo combination through a differentiable projection operator. The difference is calculated by combining the observed pixel coordinates to form the observation constraint residual. The external multi-source data includes LiDAR elevation points, SAR strong scattering points, and vector road intersections.

[0049] The projection residuals of ground control points are combined with the observation constraint residuals to generate a set of projection residuals;

[0050] The bias field of the RPC model is parameterized as the variable nodes to be optimized, and the projection residual is used as a factor constraint to construct the perceptual map optimization energy function. The perceptual map optimization energy function is to parameterize the systematic bias field existing in RPC, set each component of the bias field as the variable nodes to be optimized, so that the model can dynamically correct the bias of different regions and different features during the optimization process. Then, the projection residual is used as a factor constraint. Specifically, for the projection difference of the same geographical location in the image domain and the LiDAR domain, it is defined as a residual vector, and a corresponding covariance matrix is ​​assigned to each residual through a perceptual lookup table to characterize the confidence level and uncertainty level of the residual, thereby constructing the perceptual map optimization energy function.

[0051] Using the energy function optimized by the perception map as input, hierarchical optimization is performed based on the projection residual set and RPC bias field parameterized nodes;

[0052] Based on the RPC obtained from the hierarchical optimization solution, the optimal stereo combination is reprojected onto a unified geographic coordinate system. During the reprojection process, the confidence ellipse information updated by the perception lookup table is fused, and masking is performed on occluded areas and low-confidence areas to ensure the geometric consistency and reliability of the generated image, and the registered image after error compensation is output.

[0053] In this embodiment, the energy function of the perception map is used as input, and a hierarchical optimization solution is performed based on the projection residual set and the parameterized nodes of the RPC bias field, as detailed below:

[0054] The energy function of the perception map is used as the optimization objective. The variables to be optimized are divided according to their scope of application to form global parameter nodes (used to describe the overall bias and scale distortion of the RPC model), local parameter nodes (used to characterize the local geometric deviations in each sub-block region), and temporal parameter nodes (used to model the continuity between multi-temporal images).

[0055] The projection residual set is used as a factor constraint and introduced into the parameter node to form a hierarchical factor map. The residual factor of the ground control point and LiDAR point constrains the global parameter node, the residual factor of the road intersection and building corner point constrains the local parameter node, and the time series parameter node is constrained by the time smoothing factor.

[0056] In the hierarchical factor graph, only the global parameter nodes are activated, and a sparse nonlinear least squares optimization method is used to coarsely correct the global parameters of the RPC model in order to constrain long baseline distortion.

[0057] After the global layer optimization is completed, the local parameter nodes are activated, and the local deviation field within the sub-block is finely corrected using the factor graph optimization method. Local discontinuities are eliminated by the sub-block boundary consistency factor.

[0058] Under multi-temporal image conditions, the temporal parameter node is activated, and a temporal smoothing constraint term is introduced;

[0059] During the layer graph optimization process, the residual factor weights are dynamically adjusted by sensing the confidence information in the lookup table, and the bias field parameters of the RPC are updated by the iterative reweighted least squares method until the energy function converges.

[0060] The hierarchical optimization solution includes a global layer, a local layer, and a temporal layer, as detailed below:

[0061] Global layer: A sparse nonlinear least squares optimization method is used to coarsely correct the global parameters of the RPC model in order to constrain long baseline distortion;

[0062] Local layer: Based on the sub-block partitioning, the local deviation field is finely corrected using the factor graph optimization method to constrain local geometric errors;

[0063] Time series layer: Under multi-temporal data, a time smoothing constraint term is introduced to ensure the continuity of the RPC bias field over time.

[0064] S3 performs multi-view geometric matching on the registered images to generate dense surface point clouds, which are then input into the semantic neural radiation field network to generate dense point clouds.

[0065] In this embodiment, multi-view geometric matching is performed on the registered image to generate a dense surface point cloud, as detailed below:

[0066] Acquire LiDAR point cloud data and project it onto the registered image coordinate system to generate the pixel position and depth value of each projection point;

[0067] For each registered image, a convolutional neural network (CNN) is applied to extract deep convolutional features, which include texture, edge, and corner information;

[0068] A depth feature vector is generated for the corresponding pixel of the projected LiDAR point cloud. The depth feature vector includes three-dimensional coordinates, normal vector and local curvature information.

[0069] Image features and LiDAR depth features are aligned by pixel position, and cross-modal feature descriptors are formed by feature stitching and weighted fusion. Each descriptor simultaneously encodes image texture information and spatial depth information.

[0070] Construct a multi-resolution image pyramid to match images at different scales, thereby enhancing the adaptability to geometric structures at different scales;

[0071] In each scale layer, the matching error is calculated using a differentiable stereo matching method, and the error is used as a differentiable loss to iteratively optimize the depth value;

[0072] The optimized depth values ​​are projected onto the geographic coordinate system to generate an initial dense surface point cloud, with each point containing three-dimensional coordinates, image texture, and depth information.

[0073] Noise filtering is performed on the initial dense surface point cloud by applying confidence weights and neighborhood consistency constraints to remove outliers and generate a high-precision dense surface point cloud.

[0074] In this embodiment, the input is fed into a semantic neural radiation field network to generate a dense point cloud, as detailed below:

[0075] The dense surface point cloud is locally voxelized, dividing it into several spatial voxel units, each containing several points. Feature encoding is then performed on each point, including:

[0076] a. Color features: The color of the corresponding pixel is obtained through multi-view image projection sampling;

[0077] b. Geometric features, including three-dimensional coordinates, normal vectors, and local curvature;

[0078] c. Semantic features: Extracting the probabilities of categories such as buildings, roads, and vegetation through a shallow semantic segmentation network;

[0079] The features of each point are normalized to generate a unified representation of the point cloud enhanced feature vector;

[0080] The multi-view image pixels and point cloud enhanced feature vectors are input into the differentiable rendering module of the semantic NeRF network. The volume density and color contribution of each pixel ray are calculated through ray sampling. The network is trained through volume rendering error and self-supervised consistency constraints, including view consistency, depth consistency and semantic consistency constraints.

[0081] During training, point cloud confidence weighting is incorporated to increase the loss contribution to high-confidence regions and reduce the loss weight to low-confidence regions, thereby enhancing the accuracy of key structural regions.

[0082] Local sampling optimization is performed on each voxel unit in the training network to improve rendering efficiency and reduce redundant calculations;

[0083] The trained network predicts the three-dimensional coordinates, color, and semantic probability distribution of each voxel unit, forming a preliminary dense point cloud.

[0084] Neighborhood smoothing and removal are performed on low-confidence points or isolated points to reduce noise and outliers, thereby improving the overall reliability and accuracy of dense point clouds.

[0085] The processed points are then subjected to global coordinate mapping and formatting to generate a final high-precision dense point cloud. The formatting includes the standardization of point attributes (coordinates, color, semantic probability) and the generation of voxelized indexes.

[0086] S4 performs cross-modal feature alignment and fusion of dense point cloud and vector data, and outputs the fused multi-source dense point cloud;

[0087] In this embodiment, dense point cloud and vector data are cross-modal feature alignment and fusion to output a fused multi-source dense point cloud, as detailed below:

[0088] The vector data is acquired and its coordinate system is transformed to place it in the same geographic coordinate system as the dense point cloud. The vector data includes the vector geometric information of road centerlines, building outlines, bridges and other urban components.

[0089] Vector data is geometrically encoded to form vector features, wherein the geometric feature encoding includes line segment direction, curvature, length, topological relationship and semantic category;

[0090] Local neighborhood feature encoding is performed on dense point clouds to form point cloud features. The local neighborhood feature encoding includes the three-dimensional coordinates, normal vector, color, semantic probability, and point cloud confidence of the points.

[0091] Point cloud features and vector features are mapped to the same high-dimensional feature space, and local features of dense point cloud are matched with vector data features based on nearest neighbor search or graph matching methods.

[0092] Based on the matching results, cross-modal feature data is extracted, and a joint optimization energy function is constructed. The cross-modal feature data includes the geometric error, semantic difference, and local smoothing constraints of the vector data.

[0093] The joint optimization energy function is used as the objective to iteratively optimize the 3D coordinates and local feature vectors of dense point cloud points.

[0094] In each iteration, the point cloud position is adjusted according to gradient descent or higher-order optimization methods to minimize geometric error, semantic error and smoothness constraints, and the iteration continues until the energy function converges.

[0095] After the iterative optimization converges, the updated 3D coordinates of each point are obtained, and the update results are stored as a preliminary optimized dense point cloud, where each point is a 3D point in the dense point cloud.

[0096] The optimized dense point cloud is fused with vector data to generate a multi-source dense point cloud, where each point contains three-dimensional coordinates, color, semantic probability, and vector constraint information.

[0097] In this embodiment, the optimized dense point cloud is fused with vector data to generate a multi-source dense point cloud, as detailed below:

[0098] The optimized dense point cloud and corresponding vector data are obtained. For key urban structures in the vector data (including but not limited to building outlines, road centerlines and intersections), corresponding constraint point sets are generated. For each point in the dense point cloud, preliminary spatial matching is performed to determine its nearest neighbor relationship with the vector constraint points.

[0099] For each point in each dense point cloud, calculate its spatial distance to the nearest neighbor in the set of constraint points, use it as a geometric error index, and fine-tune the point position according to the error to make the dense point cloud geometrically aligned with the vector data.

[0100] The semantic probability of each point is updated with weights to make its semantic category distribution consistent with the vector data, thus achieving semantic-level fusion;

[0101] The fused dense point cloud is subjected to neighborhood consistency filtering to eliminate local noise and isolated points, and outliers are removed to form a smooth, continuous and high-precision final dense point cloud.

[0102] The generation of the corresponding constraint point set is as follows:

[0103] Building outline: Discretize the boundary of the vector polygon into equally spaced sampling points as constraint points;

[0104] Road centerline: Sampling is performed at fixed intervals along the vector centerline to obtain road constraint points;

[0105] Road intersections and building corners: Extract vector topology nodes (intersections and corners) and use them directly as high-weight constraint points;

[0106] Other linear or regional features: uniformly sample or mesh them according to their geometric properties to form corresponding constraint points.

[0107] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0108] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0109] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0110] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0111] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0112] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery, characterized in that: include: Acquire sub-meter resolution high-resolution remote sensing images of the target area and establish an initial land-image projection mapping relationship based on a rational polynomial coefficient model; Obtain ground control points, correct the mapping relationship, and output registered images; Multi-view geometric matching is performed on the registered images to generate dense surface point clouds, which are then input into a semantic neural radiation field network to generate dense point clouds. Cross-modal feature alignment and fusion are performed between dense point clouds and vector data to output a fused multi-source dense point cloud; The process of acquiring ground control points, correcting the mapping relationship, and outputting registered images is as follows: Obtain the known 3D coordinates of ground control points in the geographic coordinate system, and calculate the predicted pixel coordinates of the 3D coordinates in the optimal stereo combination using a differentiable projection operator; Extract the difference between the observed pixel coordinates and the predicted pixel coordinates of the ground control points in the optimal stereo combination, and output the ground control point projection residual. Acquire external multi-source data, project the corresponding 3D geographic coordinates onto the optimal stereo combination using a differentiable projection operator, and calculate the difference by combining the observed pixel coordinates to form the observation constraint residual; The projection residuals of ground control points are combined with the observation constraint residuals to generate a set of projection residuals; The bias field of the RPC model is parameterized as the node of the variable to be optimized, and the projection residual is used as a factor constraint to construct the energy function of the perception map. Using the energy function optimized by the perception map as input, hierarchical optimization is performed based on the projection residual set and RPC bias field parameterized nodes; Based on the RPC obtained from the hierarchical optimization solution, the optimal stereo combination is reprojected onto a unified geographic coordinate system, and the registered image after error compensation is output.

2. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 1, characterized in that, The acquisition of sub-meter resolution high-resolution remote sensing images of the target area, and the establishment of an initial land-image projection mapping relationship based on a rational polynomial coefficient model, are as follows: Acquire multiple sub-meter high-resolution remote sensing images of the target area and their accompanying rational polynomial coefficient files, and unify the projected coordinate system of the images with the geographic coordinate system and elevation datum of the elevation reference to form sub-meter high-resolution remote sensing images and the corresponding initial RPC function set. Data features are extracted from sub-meter resolution high-resolution remote sensing images, filtered, and the optimal stereo combination is selected based on the information gain criterion. Analyze the RPC function set, establish the projection function, and output the Jacobian matrix and differentiable projection operator; The image region is divided into sub-blocks based on the terrain slope and projection linearization error of the target area, and a multi-scale pyramid image is generated in each sub-block. Within each sub-block, the ground point coordinates are sampled, and the ground point coordinates are projected onto the pixel planes of the optimal stereo combination using a differentiable projection operator to obtain the corresponding pixel coordinates. The error covariance of each pixel is then calculated using the prior elevation variance and the Jacobian matrix. Construct a covariance matrix based on the pixel-domain covariance and derive the corresponding confidence ellipse parameters; The ground point coordinates, pixel coordinates, and confidence ellipse parameters are stored as table entries and organized into a perceptual lookup table based on spatial location index.

3. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 2, characterized in that, The data features extracted from sub-meter resolution high-resolution remote sensing images are filtered, and the optimal stereo combination is selected based on the information gain criterion, as follows: Sub-meter high-resolution remote sensing images are eliminated based on a preset threshold to obtain a candidate image set; Using features such as imaging time, off-axis angle, overlap, and cloud cover as input variables, a comprehensive evaluation function is constructed and mapped to the quality score of candidate images; The contribution of each feature to the final quality score is calculated based on the quality score using the information gain criterion. The features are sorted according to their information gain, and the sorting results are normalized to obtain the weight values ​​of each feature within a fixed interval. Based on the determined feature weights, the comprehensive evaluation function of all candidate images is solved by weighting to obtain the corresponding weighted comprehensive score; Candidate images are sorted from highest to lowest according to their weighted composite scores, and the image pair with the highest composite score is selected as the optimal stereo combination.

4. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 3, characterized in that, The process involves using the energy function optimized from the perception map as input, and performing hierarchical optimization based on the projection residual set and RPC bias field parameterized nodes, as detailed below: The energy function of the perception graph is used as the optimization objective, and the variables to be optimized are divided according to their scope of action to form global parameter nodes, local parameter nodes and time-series parameter nodes; The projection residual set is used as a factor constraint and introduced into the parameter node to form a hierarchical factor graph. In the hierarchical factor graph, only the global parameter nodes are activated, and the global parameters of RPC are coarsely corrected by a sparse nonlinear least squares optimization method. After the global layer optimization is completed, the local parameter nodes are activated, and the local deviation field within the sub-block is finely corrected using the factor graph optimization method. Local discontinuities are eliminated by the sub-block boundary consistency factor. Under multi-temporal image conditions, the temporal parameter node is activated, and a temporal smoothing constraint term is introduced; During the layer graph optimization process, the residual factor weights are dynamically adjusted by sensing the confidence information in the lookup table, and the bias field parameters of the RPC are updated by the iterative reweighted least squares method until the energy function converges.

5. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 4, characterized in that, The process of performing multi-view geometric matching on the registered images to generate a dense surface point cloud is as follows: Acquire LiDAR point cloud data and project it onto the registered image coordinate system to generate the pixel position and depth value of each projection point; For each registered image, a convolutional neural network is applied to extract deep convolutional features; Generate depth feature vectors for the corresponding pixels of the projected LiDAR point cloud; Image features and LiDAR depth features are aligned by pixel position, and cross-modal feature descriptors are formed by feature stitching and weighted fusion. Construct a multi-resolution image pyramid to match images at different scales; In each scale layer, the matching error is calculated using a differentiable stereo matching method, and the error is used as a differentiable loss to iteratively optimize the depth value; The optimized depth values ​​are projected onto the geographic coordinate system to generate an initial dense surface point cloud; Noise filtering is performed on the initial dense surface point cloud by applying confidence weights and neighborhood consistency constraints to remove outliers and generate a high-precision dense surface point cloud.

6. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 5, characterized in that, The input is fed into a semantic neural radiation field network to generate a dense point cloud, as detailed below: The dense surface point cloud is locally voxelized, dividing it into several spatial voxel units, each voxel containing several points, and each point is feature-encoded. The features of each point are normalized to generate an enhanced feature vector for the point cloud. The multi-view image pixels and point cloud enhanced feature vectors are input into the differentiable rendering module of the semantic NeRF network. The volume density and color contribution of each pixel ray are calculated through ray sampling, and the network is trained through volume rendering error and self-supervised consistency constraints. The trained network is used to predict the three-dimensional coordinates, color, and semantic probability distribution of each voxel unit, forming a preliminary dense point cloud. Low-confidence points and isolated points are preprocessed, and the processed points are then subjected to global coordinate mapping and formatting to generate a final high-precision dense point cloud.

7. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 6, characterized in that, The process of performing cross-modal feature alignment and fusion between dense point clouds and vector data to output a fused multi-source dense point cloud is as follows: Acquire vector data and perform coordinate system transformation on the vector data to place it in the same geographic coordinate system as the dense point cloud; Geometric feature encoding is performed on vector data to form vector features; Local neighborhood feature encoding is performed on dense point clouds to form point cloud features; Point cloud features and vector features are mapped to the same high-dimensional feature space, and local features of dense point cloud are matched with vector data features based on nearest neighbor search or graph matching methods. Based on the matching results, cross-modal feature data is extracted, and a joint optimization energy function is constructed; The joint optimization energy function is used as the objective to iteratively optimize the 3D coordinates and local feature vectors of dense point clouds; In each iteration, the point cloud position is adjusted according to the gradient descent method, and the iteration continues until the energy function converges; After the iterative optimization converges, the updated 3D coordinates of each point are obtained, and the update results are stored as a dense point cloud after preliminary optimization. The optimized dense point cloud is fused with vector data to generate a multi-source dense point cloud.

8. The automated 3D reconstruction method for urban components from sub-meter resolution remote sensing imagery according to claim 7, characterized in that, The process of fusing the optimized dense point cloud with vector data to generate a multi-source dense point cloud is as follows: The optimized dense point cloud and corresponding vector data are obtained. A set of corresponding constraint points is generated for the key urban structures in the vector data. Preliminary spatial matching is performed on each point in the dense point cloud to determine its nearest neighbor relationship with the vector constraint points. For each point in each dense point cloud, calculate its spatial distance to the nearest neighbor in the set of constraint points, use it as a geometric error index, and adjust the point position according to the error to make the dense point cloud geometrically aligned with the vector data. The semantic probability of each point is weighted and updated to make its semantic category distribution consistent with the vector data, thereby achieving semantic fusion. The fused dense point cloud is subjected to neighborhood consistency filtering to eliminate local noise and isolated points, and outliers are removed to form the final dense point cloud.

Citation Information

Patent Citations

  • A rapid modeling method for 3D traffic scenes based on high-precision maps

    CN113343346B

  • Urban building three-dimensional reconstruction method and device

    CN120471962A

  • Method of realizing satellite remote sensing image high precision geometric correction through slightly modifying RPC parameters

    CN105761228A

  • Three-dimensional reconstruction method integrating laser radar and oblique photography

    WO2025036361A1