A building three-dimensional reconstruction method based on multi-coplanar geometry and graph neural network

By explicitly modeling multiple sets of local planar perspective relationships and multi-scale pyramid processing, combined with graph neural networks, the problems of matching accuracy and completeness in cross-platform architectural 3D reconstruction were solved, achieving high-precision restoration of architectural facade details and 3D reconstruction.

CN121685875BActive Publication Date: 2026-05-19NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH
Filing Date
2026-02-10
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies for 3D architectural reconstruction in cross-platform, high-viewpoint-difference scenarios suffer from problems such as decreased matching accuracy, local geometric mismatch, and significant impacts from reconstruction holes, occlusion, and lighting differences, making it difficult to achieve high-precision and high-completeness 3D reconstruction.

Method used

By explicitly modeling multiple sets of local planar perspective relationships in image pairs, and combining multi-scale pyramid processing and graph neural networks, depth features of point and line structures are extracted. The homography matrix is ​​optimized using a multi-local homography matrix and a random sampling consensus algorithm. A multi-scale pyramid representation is constructed and point-line joint matching is performed to generate a robust cross-platform control point set.

Benefits of technology

It significantly improves the success rate of cross-platform matching and the spatial distribution uniformity of matching points, improves point cloud density and mesh integrity, and enhances the ability to restore building facade details, making it suitable for 3D modeling tasks in urban areas and large scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685875B_ABST
    Figure CN121685875B_ABST
Patent Text Reader

Abstract

The application relates to a building three-dimensional reconstruction method based on multi-coplanar geometry and a graph neural network, which comprises the following steps: pre-processing aerial images and ground images and obtaining an initial point matching set; iteratively extracting multiple groups of local homography matrices between the image pairs; scaling the aerial images and the ground images and constructing a multi-scale pyramid representation; performing matching by using a deep matching network; after the matching results are converged, the inverse transformation of the homography matrix is used to project back to the original image coordinate system to form a candidate cross-platform corresponding point line set; and verified cross-platform control point set is used to replace the matching input of the reconstruction process to output a final building three-dimensional model. The application can obtain sufficient, uniform and geometrically consistent cross-platform control points in cross-platform image pairs with large angle differences, scale changes and local occlusions, thereby replacing or reinforcing the matching input of traditional three-dimensional reconstruction, and realizing high-precision, high-completeness three-dimensional reconstruction and detail recovery of a city building scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image registration and 3D reconstruction technology, specifically to a method for 3D reconstruction of buildings based on multi-coplanar geometry and graph neural networks. Background Technology

[0002] In classic matching workflows based on local descriptors, a common approach is to first extract local invariant features from the image, such as SIFT, SURF, and ORB. Then, pairwise matching is performed based on the descriptors. Next, robust fitting methods such as RANSAC are used to estimate homography, the fundamental matrix, or projective geometry. Finally, the corresponding points obtained through matching are fed into the SfM and BA workflows to recover the camera pose and perform 3D point cloud reconstruction. This type of method has a clear technical workflow, mature engineering, and can be directly integrated into mainstream reconstruction software such as COLMAP and MetaShape. However, it has the following shortcomings when facing cross-platform matching between air and ground: First, when there are large differences in viewpoint and strong scale transformations, the repeatability and discriminability of local descriptors decrease significantly, resulting in sparse or incorrect matching points. Second, when there are multi-planar structures and a large number of repetitive textures (window panes, railings) on building facades, a single global transformation is often insufficient to describe local perspective differences, resulting in local geometric mismatches and reconstruction holes. Third, occlusion, strong shadows, and differences in lighting further reduce the stability of feature detection and description, making it difficult for traditional workflows to obtain uniform and high-density cross-platform control points in complex urban environments. In summary, the applicability of the classic matching process is limited in cross-platform scenarios with high perspective differences, which directly affects the quality and integrity of subsequent 3D reconstruction.

[0003] In recent years, deep learning methods, especially matching methods based on attention mechanisms and graph neural networks, have provided new ideas for overcoming the aforementioned limitations. Representative methods, such as point feature-based deep matchers and attention-based matching networks, have significantly improved the robustness of matching in complex backgrounds by learning end-to-end feature representations and matching decisions. Meanwhile, deep matching combined with multi-scale or pyramid strategies also shows advantages in handling scale differences. However, existing deep matching methods still have several shortcomings in cross-platform scenarios between aviation and ground: First, a single global matching network struggles to simultaneously handle large-scale perspective distortion and local planar differences, leading to decreased matching accuracy under large viewpoint variations; second, most end-to-end models are primarily point-level features, lacking explicit modeling of linear structures such as building outlines, eaves, and window frames, resulting in deficiencies in restoring building facade details; third, single-scale or simple pyramid fusion strategies are still insufficient when dealing with complex multi-planar geometry, easily causing uneven matching coverage and concentrated mismatches.

[0004] Patent CN116206063B discloses an adaptive UAV image matching pair selection method and system. This invention provides a method that uses global feature descriptor vectors to replace word frequency calculation based on local features and offers a graph-based indexing strategy to achieve efficient overlapping image search. However, this method cannot achieve high-precision matching and reconstruction results, and the reconstruction effect is difficult to guarantee in large scenes. Patent CN120526066B discloses a 3D reconstruction method for panoramic images based on 3DGS. This invention provides a method for rapid 3D reconstruction using neural radiation fields. However, this method requires a high degree of overlap between views and is not suitable for 3D reconstruction between large scene images. Patent CN120747366A discloses a dynamic 3D reconstruction method for monocular video scenes based on optical flow. This invention achieves spatiotemporal consistency constraints on optical flow and can accurately realize end-to-end reconstruction from monocular video to 3D scenes, but it still does not solve the problem of 3D reconstruction in large scenes. Patent CN120747387A discloses an automated 3D reconstruction method for urban components from sub-meter resolution remote sensing images. This method solves the problems of large registration errors, sparse point clouds with high noise, and semantic loss in traditional methods, significantly improving the accuracy and stability of image registration. It can not only achieve dynamic adjustment of global and local parameters, but also effectively reduce the proportion of noise and outliers. However, using only remote sensing images still cannot meet the accuracy requirements. Patent CN116797806B discloses a multi-source heterogeneous large-scene remote sensing image matching method with a global optimal solution. It achieves high-precision matching results by constructing a sub-pixel point group feature detection algorithm and a multi-mode matching algorithm based on wavelet analysis. However, due to its insufficient operating efficiency, it cannot be widely used in 3D reconstruction of large scenes. Summary of the Invention

[0005] The purpose of this invention is to provide a method for 3D reconstruction of buildings based on multi-coplanar geometry and graph neural networks. By explicitly modeling multiple sets of local planar perspective relationships in image pairs, and performing depth feature extraction and matching of point and line structures in each local perspective subdomain, a sufficient number of uniformly distributed and geometrically consistent cross-platform control points can be obtained in cross-platform image pairs with large perspective differences, scale variations and local occlusions. This replaces or enhances the matching input of traditional 3D reconstruction, and achieves high-precision, high-completeness 3D reconstruction and detail restoration of urban architectural scenes.

[0006] The technical solution adopted in this invention is: a method for 3D reconstruction of buildings based on multi-coplanar geometry and graph neural networks, comprising the following steps:

[0007] S1: Acquire aerial images and ground images, and preprocess the aerial images and ground images to obtain an initial line segment matching set and thereby obtain an initial point matching set;

[0008] S2: Based on the initial point matching set in step S1, multiple sets of local homography matrices are iteratively extracted between image pairs. The iterative extraction is achieved through the Random Sample Consensus Algorithm (RANSAC) and an iterative filtering method. Specifically, each time, RANSAC is used to identify a set of points (interior points) within the image from the initial point matching set, and the corresponding homography matrix is ​​calculated. Then, the interior points in this set are removed, and a new set of interior points is identified in the initial point matching set to continue iterating until no new interior points are obtained. The homography matrices of all interior points are used as local coplanar perspective constraints to form multiple local homography matrices. A line segment matching error is constructed based on line segment features for the homography matrix, and the homography matrix is ​​optimized by minimizing this error.

[0009] S201: Using the initial point matching set as input, call the Random Sampling Consensus Algorithm (RANSAC) to calculate the homography matrix and extract the corresponding interior points; if the interior points are not empty, record the interior point set as a set of locally coplanar interior points and delete the interior point set from the initial point matching set. Then repeat this step in the remaining initial point matching set until no more interior points satisfying the homography constraint are found.

[0010] S202: The multi-local homography matrix is ​​stored in the form of a three-dimensional array, which serves as the basis for subsequent local sub-image construction and inverse transformation back-projection mapping;

[0011] S203: Based on the initial line segment matching set, calculate the matching error between the endpoints of corresponding line segments in the aerial image and the ground image respectively. The matching error is the cumulative value of the sum of squares of the Euclidean distances between the endpoints of the line segments in the aerial image and the endpoints of the corresponding line segments obtained by transforming the corresponding endpoints in the ground image through the homography matrix. By minimizing the matching error, an optimized homography matrix is ​​obtained.

[0012] S3: Based on the image pyramid structure, the aerial and ground images in step S1 are scaled into several local sub-images, and a multi-scale pyramid representation is constructed for the local sub-images; a point-line joint depth matching network is used for matching at each layer of the pyramid. The depth matching network includes a key point detection and description module, a line segment detection module, a position and line segment encoder, a point and line message passing module, and a dual softmax allocation module for matching decisions.

[0013] S4: After pooling the matching results of the multi-scale layers, the inverse transformation of the homography matrix is ​​used to back-project the original image coordinate system to form a set of candidate cross-platform corresponding points and lines; the verified cross-platform control point set is used to replace the matching input of the reconstruction process to output the final 3D building model.

[0014] Furthermore, the preprocessing includes converting the image to an appropriate color space and applying a limited contrast adaptive histogram equalization method to the luminance channel to enhance image contrast, as well as extracting edge lines. Specifically, the method of enhancing image contrast involves applying a limited contrast adaptive histogram equalization method to the Y channel in the YCbCr color gamut, and then recombining the Y channel with the original Cb and Cr channels after enhancement to obtain an enhanced image. Simultaneously, the Canny edge detection operator is used to extract edges and line segment candidates are obtained through probabilistic Hough transform. The line segment candidates are then subjected to confidence-based weighted voting and non-maximum suppression to remove redundancy and noise.

[0015] Furthermore, the multi-scale pyramid is a Gaussian pyramid, and the number of layers of the multi-scale pyramid is denoted as s. Point-line joint matching is performed on each pyramid layer using a point-line joint deep matching network. The deep matching network includes at least: a key point detection and descriptor module based on the SuperPoint architecture, a line segment detector module based on the line segment detection algorithm LSD, a multilayer perceptron (MLP) encoder for position encoding and line segment description, a self-attention connection module for points, a line segment message passing module (LMP), and dual softmax matching decision modules for points and lines. The matching results of each layer are aggregated according to scale information and back-projected to the original image through inverse transformation.

[0016] The SuperPoint architecture's keypoint detection and descriptor module outputs keypoint locations and corresponding keypoint description vectors from the input sub-image of each pyramid layer. The line segment detector module based on the Line Segment Detection Algorithm (LSD) outputs a set of line segments and the endpoint geometric information of each line segment. The Multilayer Perceptron (MLP) encoder for position encoding and line segment description maps keypoint coordinates and line segment endpoints to point and line features of a unified dimension. The self-attention connection module for points is used for context aggregation to update point features. The line segment message passing module (LMP) is used to jointly update point and line features. The dual softmax matching decision modules for points and lines respectively output point matching probability matrices and line matching probability matrices.

[0017] The point-line joint matching uses a weighted sum of point-level negative log-likelihood loss and line-segment-level negative log-likelihood loss for training in the matching loss calculation to ensure the collaborative optimization of point and line matching. In the inference stage, matching pairs are selected using confidence threshold and geometric consistency threshold.

[0018] The beneficial effects of this invention are as follows:

[0019] (1) This invention extracts multiple sets of local coplanar homography transformations within an image pair by explicitly extracting them, decomposing the complex global perspective transformation into several local approximate perspective subdomains. Combined with multi-scale pyramid processing on the local subdomains, it significantly reduces the negative impact of large viewpoint differences and scale changes on matching. In each local subdomain, a point-line joint deep matching network is introduced. Based on the graph neural network structure and combined with transfer learning, it better represents the structured geometric information of the building facade, such as eaves, window frames and outlines, through joint encoding, message passing and joint decision-making of key points and line segments. Thus, it can still obtain robust and geometrically reliable corresponding points under conditions of scarce texture, repetitive texture or local occlusion. It not only improves the success rate of cross-platform matching, but also improves the spatial distribution uniformity of matching points, providing a higher quality control point set for subsequent 3D reconstruction.

[0020] (2) After the verified cross-platform control points are replaced or reinforced in the traditional 3D reconstruction process, the present invention shows significant improvement in point cloud density, mesh integrity and detail recovery. It reduces holes and local distortions caused by sparse or uneven distribution of matching. It has good modularity and engineering compatibility and can be connected with existing commercial or open source reconstruction software. It is convenient to promote its application in air-ground joint 3D modeling tasks in city-level or larger scenarios. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the method of an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of multi-plane homography transformation in an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the homography matrix optimization workflow in an embodiment of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of the deep matching graph neural network (GNN) in an embodiment of the present invention;

[0026] Figure 5 This is a diagram showing the matching results of an embodiment of the present invention;

[0027] Figure 6 This is a diagram showing the connection points superimposed on a dense point cloud in an embodiment of the present invention;

[0028] Figure 7 This is a comparative schematic diagram of building modeling based on white film according to an embodiment of the present invention;

[0029] Figure 8 This is a comparative schematic diagram of textured building models according to an embodiment of the present invention. Detailed Implementation

[0030] To better understand the above-described objects, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the invention; however, the invention may be practiced in other ways different from those described herein, and therefore, the invention is not limited to the specific embodiments disclosed below.

[0031] like Figure 1 As shown, this embodiment of the invention provides a method for 3D reconstruction of buildings based on multi-coplanar geometry and graph neural networks, including the following steps:

[0032] S1: Acquire aerial and ground images, and preprocess the aerial and ground images. In this embodiment of the invention, the preprocessing includes converting the image to an appropriate color space and applying limited contrast adaptive histogram equalization to the luminance channel to enhance image contrast, as well as extracting edge lines; the edge line extraction uses an edge detection operator to extract edges from the enhanced image and obtains line segment candidates through probabilistic Hough transform, and applies confidence-based weighted voting and non-maximum suppression to the line segment candidates to remove redundancy and noise.

[0033] The specific method to enhance image contrast is to implement a contrast-limited adaptive histogram equalization method on the Y channel in the YCbCr color gamut, and then recombine the Y channel with the original Cb and Cr channels after enhancement to obtain an enhanced image.

[0034] The edge detection operator is the Canny edge detection operator, which extracts edges from the enhanced image and outputs a corresponding edge confidence map. The edge confidence map is used to characterize the edge strength of each edge point and serves as the basis for subsequent line feature detection and calculation.

[0035] The edge line detection employs a line segment detection method based on probabilistic Hough transform. This method involves probabilistically sampling edge points and estimating the parameters of candidate line segments based on the geometric relationships of the sampled pixels. Specifically, it maps edge points in the image space to a parameter space and accumulates a vote on possible line parameters, as described in the following manner:

[0036] ;

[0037] in, Indicates the first in the image An edge point, These are the pixel coordinates of the edge point in the image coordinate system; The discretized angle between the line and the horizontal axis of the image is the first... The horizontal axis is defined as a coordinate axis that extends to the right along the horizontal direction of the image, with the lower left corner of the image as the origin. For the first Each edge point at the angle The corresponding distance value, in pixels, is used to describe the position of the line in the Hough transform space; for any edge point in the image All satisfy the above formula, with the angle The above formula gives the edge points. A curve in parameter space, such that multiple edge points located on the same straight line are identical or close to each other in parameter space. The parameter unit generates a consistent voting response.

[0038] Meanwhile, to enhance line features, this embodiment of the invention incorporates a probabilistic weighting mechanism that assigns weights to votes based on the confidence level of edge points. Specifically, different weights are assigned to the votes of each edge point in the edge confidence map. By introducing these weights, pixels with greater edge strength contribute more to the accumulator, which is a voting matrix stored in the parameter space to record the weighted voting results corresponding to different line parameters. This enhances the ability to detect lines corresponding to strong edges, improves the response to high-quality edges, and suppresses unnecessary invalid votes caused by noise.

[0039] In the accumulator parameter space, only the most significant peak where the voting values ​​reach a local maximum is retained. This most significant peak corresponds to the line parameter combination with the highest support, thereby reducing redundant line segment detection results and improving the accuracy and stability of line feature extraction. A probability weighting mechanism introduces weights during the accumulation phase of the transformation. For each pair in the accumulator array Voting. The specific formula is:

[0040] ;

[0041] in, Distance parameter Discretized Each distance cell takes a value. Angle parameters Discretized Each angle unit takes a value. and These correspond to the indices of the accumulator's two-dimensional array on the distance and angle axes, respectively. For each edge point... At a fixed angle Calculation And quantize to the nearest distance unit. , Indicates the parameter unit in the accumulator The number of votes, Represents the q-th edge point The weight, This represents the Dirac function, which indicates the parameter matching relationship and is used to... The points are added together. Through the above voting accumulation process, when multiple edge points located on the same straight line are in the same or similar group in the parameter space, the points are added together. When the parameter produces a consistent weighted response, the parameter unit will form a significant peak, thus corresponding to a candidate line feature.

[0042] In edge line detection, when the probabilistic Hough transform generates too many candidate lines, many of these candidate lines may almost overlap or be very close to each other. To address this issue, this invention uses non-maximum suppression (NMS) to enhance the probabilistic Hough transform, aiming to identify the most accurate candidate sequence corresponding to the actual sequence from the candidate sequences, thereby reducing duplicate detections and redundant information. Specifically, the method involves first sorting the line segments in descending order of their lengths to create a sorted set, and then initializing an empty set to store the line segments to be suppressed. Each line segment is then evaluated; if the length difference between the line segment and any line segment already in the suppression set is below a set threshold, the line segment is ignored; otherwise, it is added to the suppression set. Non-maximum suppression effectively utilizes the results of line detection to eliminate redundancy in the accumulator array. The threshold is determined using the standardized deviation based on the line segment length. Considering that the length of a line segment is often a key factor affecting its similarity, a relative length difference is established as the threshold standard. Consider two line segments L1 and L2 with lengths |L1| and |L2| respectively, their relative length difference... It can be calculated as:

[0043] ;

[0044] In this invention, max(L1, L2) is used to normalize the length difference, enabling comparison of line segments of different lengths on a relative scale. Through multiple experiments, this embodiment of the invention selected an appropriate percentile as the threshold to ensure the identification of most similar line segments. If 90% of the normalized length differences between line segments in the image dataset fall within the range of 0 to 0.05, it indicates that the length differences between most line segment pairs are small, typically representing different views of the same object or different parts of the same line. Therefore, the threshold can be set to 0.05 or 0.1 to filter out these nearly identical line segments. The remaining 10% of line segment pairs have a normalized difference greater than 0.05, indicating significant length differences, possibly representing different objects or significantly different features. It is recommended to retain these line segments. In this embodiment, 0.05 is selected as the threshold, retaining 10% of the line segments to represent different edge line features. Then, line segment matching is performed based on the extracted edge line features to obtain an initial line segment matching set and an initial point matching set.

[0045] S2: Based on the initial point matching set, multiple sets of local homography matrices are iteratively extracted between image pairs. The iterative extraction is achieved through the Random Sample Consensus Algorithm (RANSAC) and an iterative selection method. Specifically, each time, RANSAC is used to identify a set of points (interior points) within the image from the initial point matching set, and the corresponding homography matrix is ​​calculated. Then, the interior points in this set are removed, and a new set of interior points is identified in the initial point matching set to continue iterating until no new interior points are obtained. The homography matrices of all interior points are used as local coplanar perspective constraints to form multiple local homography matrices. Line segment matching errors are constructed based on line segment features for the homography matrices, and the homography matrices are optimized by minimizing these errors.

[0046] S201: Using the initial point matching set as input, determine the coplanar relationship between the aerial image and the ground image, as well as the homography matrix corresponding to each local region between the two images, and identify the corresponding points in the aerial-ground stereo image pair.

[0047] The specific method is as follows: First, establish the corresponding coplanar set S, with the following expression:

[0048] ;

[0049] in, Representing ground imagery and aerial imagery A set of images that are coplanar. This indicates the existence of ground images with coplanar S. , Aerial images indicating the existence of coplanar S .

[0050] Unlike traditional photogrammetric methods that rely on fundamental matrix constraints for matching, camera intrinsic parameters and extrinsic orientation elements are not necessary conditions for establishing a spatial correspondence between aerial oblique images and ground images. This invention presents a multi-viewpoint coplanar extraction method based on a multi-dimensional homography model M, which constructs different geometric spatial relationships by retrieving matching points from different facades. Since multi-view aerial and ground images contain multiple planes of buildings, multiple local homography matrices are used instead of a global homography matrix to characterize the geometric relationships between stereo image pairs. For multi-view aerial images of the same building, the mathematical model can be expressed as:

[0051] ;

[0052] in, This represents the m-th ground image. Let P represent the nth aerial image, and let P represent the set of multi-view aerial and ground images that are coplanar. This represents the perspective geometric transformation matrix between images within the set of multi-view aerial and ground images P.

[0053] Subsequently, the Random Sample Consensus Algorithm (RANSAC) is used to identify a set of interior points from the initial point matching set and calculate the corresponding homography matrix to determine the local perspective transformation matrix, i.e., the local homography matrix, between the coplanar set S and the multi-view images. Then, a spatial mapping relationship between the ground images is established using a 3D array R.

[0054] ;

[0055] in, Multi-view images representing coplanar relationships and The l-th perspective geometric transformation between them, This represents a multi-view image that has a coplanar relationship. The formula indicates that if two multi-view images... and If two images have a set of corresponding coplanar points, then there is a perspective geometric transformation between them; if not, then there is no such transformation.

[0056] Subsequently, an iterative filtering method was used to extract multidimensional perspective geometric transformations from the two-dimensional image. Point matching set It is obtained through the initial line segment matching set, corresponding to the coplanar set. Under the constraint of perspective transformation, if the Random Sample Consensus Algorithm (RANSAC) can identify interior points, then the perspective geometric transformation λ is calculated. ][ ][l], and from the corresponding coplanar set Remove interior points from the middle, then update to... The process is iteratively repeated using the Random Sample Consensus (RANSAC) algorithm to find interior points until no further interior points are detected, at which point the loop iteration is terminated.

[0057] S202: The multi-local homography matrix is ​​stored in the form of a three-dimensional array, serving as the basis for subsequent local sub-image construction and inverse projection transformation mapping.

[0058] S203: Based on the initial line segment matching set, calculate the matching error between the endpoints of corresponding line segments in the aerial image and the ground image respectively. The matching error is defined as the cumulative value of the sum of squares of the Euclidean distances between the endpoints of the line segments in the aerial image and the endpoints of the corresponding line segments obtained by transforming the corresponding endpoints in the ground image through the homography matrix. By minimizing the matching error, an optimized homography matrix is ​​obtained.

[0059] In this embodiment of the invention, the perspective transformation of a single local space is calculated based on the line segment candidates obtained in step S1. For two sets of line segments... and In the image and In the process, line segment L2 is transformed into homogeneous coordinate form, that is, point L2 is transformed into homogeneous coordinate form. and points Transformed into corresponding homogeneous coordinates and :

[0060] ;

[0061] Then, calculate the homogeneous coordinates corresponding to line segment L2. The line segments are specifically obtained by using the initial homography matrix 𝑴 After transformation, we get:

[0062] ;

[0063] ;

[0064] in, and They represent and Transformed coordinates Represents homogeneous coordinates Convert to homogeneous coordinates Scale factor, Represents homogeneous coordinates Convert to homogeneous coordinates The scaling factor.

[0065] The homogeneous coordinates of line segment L2 ( , The standardization is performed using the following formula:

[0066] ;

[0067] ;

[0068] in, Homogeneous coordinates Standardized coordinates Homogeneous coordinates Standardized coordinates.

[0069] To further optimize the initial matrix using the extracted line features, an error equation is constructed by calculating the line segment matching error E. The line segment matching error E is defined as the sum of the squares of the distances between the endpoints of the original line segment and the transformed line segment, specifically expressed as:

[0070] ;

[0071] Where ‖∙‖ represents the Euclidean distance, z represents the Euclidean coordinates in the image plane, and h represents the Euclidean coordinates obtained by normalizing homogeneous coordinates. and It is an image The coordinates of the two endpoints of the p-th line segment are and These are the endpoint coordinates of the line segments obtained through L2 perspective transformation. Then, by minimizing the line segment matching error... Thus, the optimized matrix is ​​obtained. ,Right now:

[0072] .

[0073] S3: Based on the image pyramid structure, the aerial and ground images from step S1 are scaled into several local sub-images, and a multi-scale pyramid representation is constructed for the local sub-images, thereby making GNN suitable for large-scale image matching; a point-line joint deep matching network is used for matching at each layer of the pyramid. The deep matching network includes a key point detection and description module, a line segment detection module, a position and line segment encoder, a point and line message passing module, and a dual softmax allocation module for matching decisions.

[0074] The multi-scale pyramid is a Gaussian pyramid, used to construct a multi-resolution level deep feature extraction model, enabling the depth matching grid in this embodiment to adapt to cross-scale variations in ground-based images. To address the problem of requiring large amounts of data to train the point-line joint deep matching network, transfer learning techniques are employed to share some mature network modules, such as the parameters of GlueStick, thereby reducing the need for sample generation and enhancing the generalization ability of the network model. Mathematically, images... scale space Gaussian convolution is defined as variable-scale Gaussian convolution, and the Gaussian function is defined as...

[0075] ;

[0076] ;

[0077] ;

[0078] ;

[0079] in, Represents the Gaussian function. The standard deviation is used to determine the scale of the Gaussian function; x represents the pixel's horizontal coordinate, and y represents the pixel's vertical coordinate. Here, s represents the initial scaling parameter, s represents the number of levels within each octave (i.e., the number of pyramid layers), m represents the number of octaves in the multi-scale pyramid, and n represents the number of Gaussian smoothing operations performed within each octave. The scale growth factor between adjacent scale layers .

[0080] Initial scaling parameters This is typically determined based on features and image resolution. When processing lower-resolution images, a smaller [size / size] is usually used. Value; conversely, for higher resolution images, The number should be increased accordingly. In practical applications, this can be achieved through cross-validation or comparing the effects at different scales. The optimization is as follows. For typical images, m is usually set to 3 to 5 layers. Higher n values ​​are suitable for more complex images that require stronger smoothing effects, while lower n values ​​are beneficial for images that require more detail preservation. In this embodiment of the invention, m is set to 3 and n is set to 3.

[0081] At each pyramid level, a point-line joint depth matching network is used to perform point-line joint matching. This point-line joint depth matching network is an improvement on GlueStick, such as... Figure 4As shown, the system includes a keypoint detection and descriptor based on the SuperPoint architecture, a line segment detector based on the Line Segment Detection (LSD) algorithm, a multilayer perceptron (MLP) encoder for position encoding and line segment description, self-attention connections for points, a line segment message passing (LMP) module, and dual softmax matching decision modules for points and lines. The matching results of each layer are aggregated according to scale information and back-projected to the original image through inverse transformation. The point-line joint matching is trained by simultaneously using a weighted sum of point-level negative log-likelihood loss and line segment-level negative log-likelihood loss to ensure the collaborative optimization of point-line matching. In the inference stage, matching pairs are selected using confidence thresholds and geometric consistency thresholds.

[0082] S4: After pooling the matching results of the multi-scale layers, the inverse transformation of the homography matrix is ​​used to back-project the original image coordinate system to form a set of candidate cross-platform corresponding points and lines; the verified cross-platform control point set is used to replace the matching input of the reconstruction process to output the final 3D building model.

[0083] This invention uses the deep matching graph neural network model trained in S3 to test the matching results. Simultaneously, it utilizes scale information and the inverse projection mapping matrix to propagate the matching results of each layer back to the original image. The test results are then used to evaluate accuracy. The formula for the evaluation metric is as follows:

[0084] ;

[0085] in, Indicates the number of correct matches. MP represents the total number of matches, and MP represents the matching accuracy, i.e., the matching precision.

[0086] In this invention, ten image datasets from different datasets were selected for comparison (see Table 1) to analyze the generalization performance and robustness of the method. Representative manual methods such as SIFT, RIFT, ORB, ASIFT, MeshMatch, and PSC, as well as deep learning-based methods including GlueStick, LightGlue, and SuperGlue, were selected for comparative analysis with the method described in this invention. The matching results and accuracy are detailed in Table 2.

[0087] Table 1. Detailed information on four representative datasets

[0088] Dataset sensor Image size Number of images Reference coordinate system Center Sony Nex-7 4000×6000 Aerial oblique imagery: 1212; Ground imagery: 446 WGS1984 Zeche Sony Nex-7 4000×6000 Aerial oblique imagery: 147; Ground imagery: 172 WGS1984 SWJTU-BLD Aerial oblique imagery: Sony ICLE-5100; Ground imagery: Canon EOS M6 4000×6000 Aerial oblique imagery: 207; Ground imagery: 88 WGS1984 / UTM SWJTU-RES Aerial oblique imagery: Sony ICLE-5100; Ground imagery: Canon EOS M6 Aerial oblique imagery: 4000×6000; Ground imagery: 3040×4056 Aerial oblique imagery: 92; Ground imagery: 192 WGS1984 / UTM

[0089] Table 2. Image matching accuracy evaluation results of ten methods

[0090] method 1 2 3 4 5 6 7 8 SIFT 29.03 0 0 23.33 0 0 0 0 RIFT 25.00 0 0 0 0 0 0 0 ASIFT 21.64 0 0 37.14 0 0 53.85 0 ORB 10.00 0 0 37.5 0 0 0 0 MeshMatch 90.27 82.44 57.72 90.70 57.69 63.64 57.47 47.76 SuperGlue 44.01 44.86 19.20 53.69 25.40 22.37 19.58 0 LightGlue 33.33 0 0 77.78 0 0 0 0 PSC 59.63 47.06 24.14 54.05 0 0 6.67 12.00 GlueStick 85.64 68.42 74.42 90.06 58.06 80.95 61.36 76.05 Embodiments of the present invention 92.66 90.32 80.69 99.27 82.81 88.89 92.56 84.66

[0091] The embodiments of this invention outperform other methods in terms of matching accuracy for all image pairs, achieving a maximum matching accuracy (MP) of 99.27% ​​and an average matching accuracy (MP) of 88.68%, surpassing the second-ranked MeshMatch algorithm, which has an average matching accuracy (MP) of 68.46%. This demonstrates that the embodiments of this invention can accurately and efficiently achieve image matching on empty ground images, exhibiting superior performance. The matching results are shown below. Figure 5 As shown.

[0092] Experimental results show that the proposed method can effectively improve the density of local 3D structural point clouds, thereby enhancing the representation ability of complex local details. This is especially true in multi-facade matching scenarios, where the method performs even better. Figure 6 As shown, the matching connection points superimposed on the dense point cloud are marked in green, mainly concentrated in the building facade area. These matching points have a wide coverage and good continuity of the facade, and can relatively completely depict the facade geometry. This further verifies the matching accuracy and engineering applicability of the method under planar constraints. In this embodiment of the invention, a verified cross-platform matching point set is used to replace the matching input in the reconstruction process to output the final 3D building model, and the accuracy of the reconstruction effect is evaluated. Several popular photogrammetry software were selected for comparative analysis, including the original ContextCapture, MetaShape, Pix4Dmapper, and Colmap. The evaluation indicators include: the number of images that contribute to effective modeling, the number of 3D points generated in the model, the number of faces in the generated model, and the root mean square error of the reprojected pixels. The specific results are shown in Table 3.

[0093] Table 3. Quantitative Comparison Results of Ground Building Modeling

[0094]

[0095] As can be seen from Table 3, the embodiments of the present invention successfully achieved 3D modeling of all datasets. Compared with other methods, the embodiments of the present invention have significant advantages in terms of the number of images effectively modeled, the number of 3D points generated in the model, the number of faces in the generated model, and the root mean square error of reprojection. In particular, the embodiments of the present invention generated more 3D points and triangular faces, thereby improving the accuracy of the reconstructed 3D model.

[0096] To quantitatively evaluate the effectiveness of the embodiments of the present invention in building modeling, the density of the irregular triangular network (TIN) and the level of detail of the white model constructed from the 3D point cloud of the building were also compared. Specific results are as follows: Figure 7 and Figure 8As shown. The construction of the irregular triangle network TIN across four datasets demonstrates that the embodiments of this invention integrate ground-space matching using perspective-invariant matching, achieving more matching points compared to existing irregular triangle network TINs, resulting in a denser irregular triangle network TIN. This advantage is particularly evident in the SWJTU-RES dataset. Furthermore, as... Figure 7 As shown, the white models derived from the four datasets demonstrate that the embodiments of the present invention produce richer local details, with particularly complex local features in the areas marked by the red frames. Therefore, the embodiments of the present invention can capture more refined details and features of buildings, especially in densely populated or structurally complex areas. The embodiments of the present invention more accurately reflect the shape, structure, and texture of buildings, alleviate occlusion problems, and thus improve the accuracy, realism, and completeness of the 3D building models.

[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for 3D architectural reconstruction based on multi-coplanar geometry and graph neural networks, characterized in that, Includes the following steps: S1: Acquire aerial images and ground images, and preprocess the aerial images and ground images to obtain an initial line segment matching set and thereby obtain an initial point matching set; S2: Based on the initial point matching set in step S1, multiple sets of local homography matrices are iteratively extracted between image pairs. The iterative extraction is achieved by the Random Sampling Consensus Algorithm (RANSAC) and the iterative screening method. Specifically, each time, the Random Sampling Consensus Algorithm (RANSAC) is used to identify a set of points in the image from the initial point matching set, i.e., inliers, and the corresponding homography matrix is ​​calculated. Then, the inliers in this group are removed, and a new group of inliers is identified in the initial point matching set to continue iterating until no new inliers are obtained; the homography matrix of all inliers is used as a local coplanar perspective constraint to form a multilocal homography matrix; A line segment matching error is constructed based on line segment features for the homography matrix, and the homography matrix is ​​optimized by minimizing this error; S201: Using the initial point matching set as input, call the Random Sampling Consensus Algorithm (RANSAC) to calculate the homography matrix and extract the corresponding interior points; if the interior points are not empty, record the interior point set as a set of locally coplanar interior points and delete the interior point set from the initial point matching set. Then repeat this step in the remaining initial point matching set until no more interior points satisfying the homography constraint are found. S202: The multi-local homography matrix is ​​stored in the form of a three-dimensional array, which serves as the basis for subsequent local sub-image construction and inverse transformation back-projection mapping; S203: Based on the initial line segment matching set, calculate the matching error between the endpoints of corresponding line segments in the aerial image and the ground image respectively. The matching error is the cumulative value of the sum of squares of the Euclidean distances between the endpoints of the line segments in the aerial image and the endpoints of the corresponding line segments obtained by transforming the corresponding endpoints in the ground image through the homography matrix. By minimizing the matching error, an optimized homography matrix is ​​obtained. S3: Based on the image pyramid structure, the aerial and ground images in step S1 are scaled into several local sub-images, and a multi-scale pyramid representation is constructed for the local sub-images; a point-line joint depth matching network is used for matching at each layer of the pyramid. The depth matching network includes a key point detection and description module, a line segment detection module, a position and line segment encoder, a point and line message passing module, and a dual softmax allocation module for matching decisions. S4: After pooling the matching results of the multi-scale layers, the inverse transformation of the homography matrix is ​​used to back-project the original image coordinate system to form a set of candidate cross-platform corresponding points and lines; the verified cross-platform control point set is used to replace the matching input of the reconstruction process to output the final 3D building model.

2. The method for 3D architectural reconstruction based on multi-coplanar geometry and graph neural networks according to claim 1, characterized in that, The preprocessing includes converting the image to an appropriate color space and applying a limited contrast adaptive histogram equalization method to the luminance channel to enhance image contrast, as well as extracting edge lines. Specifically, the method of enhancing image contrast involves applying a limited contrast adaptive histogram equalization method to the Y channel in the YCbCr color gamut, and then recombining the Y channel with the original Cb and Cr channels after enhancement to obtain an enhanced image. Simultaneously, the Canny edge detection operator is used to extract edges, and line segment candidates are obtained through probabilistic Hough transform. Furthermore, the line segment candidates are subjected to confidence-based weighted voting and non-maximum suppression to remove redundancy and noise.

3. The method for 3D architectural reconstruction based on multi-coplanar geometry and graph neural networks according to claim 2, characterized in that, The multi-scale pyramid is a Gaussian pyramid, and the number of layers of the multi-scale pyramid is denoted as s; At each pyramid layer, a point-line joint deep matching network is used to perform point-line joint matching. The deep matching network includes at least: a key point detection and descriptor module based on the SuperPoint architecture, a line segment detector module based on the line segment detection algorithm LSD, a multilayer perceptron (MLP) encoder for position encoding and line segment description, a point self-attention connection module, a line segment message passing module (LMP), and dual softmax matching decision modules for points and lines. The matching results of each layer are aggregated according to scale information and back-projected to the original image through inverse transformation. The SuperPoint architecture's keypoint detection and descriptor module outputs keypoint locations and corresponding keypoint description vectors from the input sub-image of each pyramid layer. The line segment detector module based on the Line Segment Detection Algorithm (LSD) outputs a set of line segments and the endpoint geometric information of each line segment. The Multilayer Perceptron (MLP) encoder for position encoding and line segment description maps keypoint coordinates and line segment endpoints to point and line features of a unified dimension. The self-attention connection module for points is used for context aggregation to update point features. The line segment message passing module (LMP) is used to jointly update point and line features. The dual softmax matching decision modules for points and lines respectively output point matching probability matrices and line matching probability matrices. The point-line joint matching uses a weighted sum of point-level negative log-likelihood loss and line-segment-level negative log-likelihood loss for training in the matching loss calculation to ensure the collaborative optimization of point and line matching. In the inference stage, matching pairs are selected using confidence threshold and geometric consistency threshold.