Non-uniform-category building facade point cloud segmentation method based on topology perception graph network

By designing the topology-aware graph network FTG-Net, combining the elevation topology extraction module, enhanced sampling geometry extraction module and dual feature attention fusion module, the problem of insufficient accuracy of the building elevation point cloud segmentation method in the existing technology when dealing with occlusion, point density changes and category imbalance is achieved, and high-precision and robust building elevation point cloud segmentation is achieved.

CN120147325AActive Publication Date: 2025-06-13NANJING UNIV OF INFORMATION SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510149804.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-13
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing deep learning-based point cloud segmentation method of building facades is difficult to effectively extract and integrate facade topology and geometric features when dealing with severe occlusion, point density changes and category imbalance, resulting in low segmentation accuracy.

Method used

A topology-aware graph network FTG-Net is designed, which includes a facade topology extraction module, an enhanced sampling geometry extraction module and a dual feature attention fusion module. Through these modules, different granularity and types of features are extracted and fused, the small sample category segmentation performance is improved and the overall performance and robustness are improved.

Benefits of technology

Accurate and robust segmentation of point clouds for facades of uneven categories of buildings is achieved, and segmentation accuracy and robustness are improved, especially when dealing with small sample categories and category imbalances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147325A_ABST
    Figure CN120147325A_ABST
Patent Text Reader

Abstract

The invention provides a non-uniform-category building facade point cloud segmentation method based on a topology perception graph network, and the method specifically comprises the steps: constructing a facade topology extraction module, and capturing abstract object-level topology features between facade elements; constructing an enhanced sampling geometric extraction module, and extracting point-level hierarchical geometric features among facade elements; constructing a double-feature attention fusion module, fusing the two obtained features with different granularities and different types, and realizing advantage complementation of the two features; the three modules are combined to form a topology perception graph network; an input data set is divided into a training set, a verification set and a test set, and after the network is trained, testing is performed and a result is evaluated. According to the method provided by the invention, the problem that topological features among building facade elements cannot be effectively extracted and effectively fused with geometric features in the prior art is solved, and effective segmentation of building facade point clouds with non-uniform categories can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent processing of three-dimensional point clouds and application research of geospatial information data, and particularly relates to a method for segmenting point clouds of building facades with uneven categories based on a topology-aware graph network. Background Art

[0002] Building facade point cloud segmentation is a technology that extracts strongly distinguishable spatial features to segment and analyze semantic elements of building facades (such as windows, walls, doors, balconies, etc.). The building facade point cloud segmentation technology is an important core of urban scene understanding and has been widely applied in various fields, such as solar potential calculation, building structure health monitoring, digital twin city construction, etc.

[0003] Nowadays, due to its high-precision and high-density characteristics, laser scanning technology has become an important means for collecting spatial point cloud data. The point cloud data contains rich geometric features, shape attributes, topological information, etc., providing key information for building facade segmentation. At the same time, due to its high computing power and rich feature description ability, deep learning has been widely popularized in the field of point cloud processing. However, the existing deep learning-based building facade point cloud segmentation methods still have the following difficulties: the building facade point clouds often have serious occlusion and point density changes; in the building facade scene, the quantity ratios of different facade elements are usually seriously unbalanced, which leads to insufficient feature learning for small-sample targets (such as doors, balconies), thereby affecting their segmentation accuracy; the spatial topological features between facade elements are crucial for segmentation, but how to effectively extract and integrate facade topologies and geometric features at different granularities is still challenging.

[0004] Existing methods for extracting point cloud features based on deep learning can be roughly divided into: point-based methods, graph-based methods, and attention-based methods. Point-based methods focus on extracting and aggregating the features of each point in the point cloud. Among them, the pioneering research on point cloud segmentation and classification based on deep learning is PointNet and PointNet++. Identifying local features can provide an important basis for segmentation. Some studies have proposed neighborhood feature aggregation methods. For example, Point SIFT divides the space into quadrants to search for neighboring points in different directions and aggregates feature information in multiple directions through convolutional operations. For large-scale point clouds, some scholars have proposed an efficient and lightweight neural network RandLA. This method adopts a random sampling strategy and uses a local feature aggregation module to extract detailed neighborhood features. Graph-based methods mainly use the permutation invariance of the graph structure to capture geometric features. By aggregating the graph structure and node features, convolution is applied to extract significant geometric features. For example, PointGCN uses global pooling and multi-resolution pooling to capture global and local information. In addition, some scholars have proposed RGCNN, which uses graph convolution approximated by Chebyshev polynomials to extract features and adds a regularization term to the loss function to constrain the model parameters. In addition, there are methods based on spatial graph convolution. For example, DGCNN constructs a local dynamic neighborhood graph and effectively captures hierarchical features in the point cloud by dynamically updating the neighborhood. Some scholars have proposed RadDGCNN, which replaces the KNN method with a spherical neighborhood to achieve more effective facade segmentation. Attention-based methods introduce an attention mechanism to adaptively learn the importance of different regions. For example, some scholars have proposed AFGL-Net, which enhances local features through position and orientation encoding. The Transformer module is integrated into AFGL-Net, and attention-based feature fusion is used to flexibly balance local features and global features. Some scholars have also proposed DLA-Net, which is embedded in the encoder-decoder network architecture and uses self-attention and attention pooling modules to learn detailed local features, achieving significant progress in facade segmentation.

[0005] In addition to the above methods, scholars have also focused on using deep learning networks to capture abstract topological features. In the field of 3D reconstruction, FoldingNet and AtlasNet capture abstract topological information in the 3D target point cloud by fusing folding and Riemann two-dimensional grids. The role of topological features in semantic segmentation has also gradually emerged. For example, some scholars have proposed a plug-and-play module SRN to extract the intrinsic topological information in the local structure of 3D objects. Some scholars have also introduced the topological theory of persistent homology into the deep learning framework and proposed the TopoSeg module and a loss function compatible with point cloud segmentation. However, existing deep learning methods still have limitations in capturing the abstract topological information between facade elements, unable to effectively extract the abstract topological features between building facade elements and effectively fuse them with geometric features. Summary of the Invention

[0006] To address the problems existing in the above-mentioned prior art, the method proposed by the present invention extracts two different granularity and type of features, namely topological features and geometric features, by designing a Topology-Aware Graph Network FTG-Net including an elevation topology extraction module, an enhanced sampling geometric feature extraction module, and a dual-feature attention fusion module. While improving the performance of small-sample class segmentation, the complementary advantages of the two features are combined to improve the overall performance and robustness of the class-imbalanced building facade point cloud segmentation.

[0007] To achieve the above technical objectives, the present invention provides the following technical solutions:

[0008] A method for class-imbalanced building facade point cloud segmentation based on a topology-aware graph network, which specifically includes the following steps:

[0009] S1. Construct an elevation topology extraction module to capture the abstract object-level topological features between elevation elements;

[0010] S2. Construct an enhanced sampling geometric extraction module to extract the point-level hierarchical geometric features between elevation elements;

[0011] S3. Construct a dual-feature attention fusion module to fuse the two different granularity and different type of features obtained in steps S1 and S2 to achieve the complementary advantages of the two features;

[0012] S4. The elevation topology extraction module, the enhanced sampling geometric extraction module, and the dual-feature attention fusion feature module are combined to form a topology-aware graph network FTG-Net; the building facade sample data is divided into a training set, a validation set, and a test set and then input into FTG-Net. After training the overall network, it is tested and the results are evaluated to achieve class-imbalanced building facade point cloud segmentation.

[0013] Further, step S1 includes:

[0014] S11. Use K-Means uniform sampling for each building facade sample data to obtain a complete facade point cloud. Take the sampled facade point cloud as an input unit and use a topology information encoder to encode the input unit to capture high-dimensional encoded features with strong discrimination;

[0015] S12. Use a topology information decoder enhanced by topology attention to extract the abstract object-level topological features between elevation elements from the high-dimensional encoded information.

[0016] More specifically, step S11 is specifically:

[0017] The sampled elevation point cloud first expands the receptive field through two consecutive edge convolution operations to learn the global coarse-grained features containing topological information among elevation elements; then, a three-layer convolution operation with dimension D 1 is used for dimension expansion to make up for the limitation that edge convolution can only extract coarse-grained features, and high-dimensional encoded features with strong discrimination are obtained The formula is expressed as:

[0018]

[0019] The sampled elevation point cloud first expands the receptive field through two consecutive edge convolution operations to learn the global coarse-grained features containing topological information among elevation elements; then, a three-layer convolution operation with dimension D 1 is used for dimension expansion to make up for the limitation that edge convolution can only extract coarse-grained features, and high-dimensional encoded features with strong discrimination are obtained The formula is expressed as:

[0020]

[0021] Among them, is the point in the sampled elevation point cloud, with a total of N; is its neighboring point, with a total of K; and respectively represent the elevation input point and its neighboring point after the first edge convolution operation; EConv(·) and Conv (1) (·) are the edge convolution operation and the three-layer convolution operation with dimension D 1 respectively; || is the feature concatenation operation.

[0022] More specifically, step S12 specifically includes:

[0023] S121. Introduce and set a two-dimensional Riemannian square grid for the topological feature decoder; set the two-dimensional Riemannian square grid as a×a, then the number N of elevation input points is a 2 ; learn the initial topological information f from the high-dimensional encoded feature Grid and the two-dimensional Riemannian square grid F c The formula is expressed as:

[0024]

[0025] Among them, Conv (2) (·) is the three-layer convolution operation with dimension D 2 ; || is the feature concatenation operation;

[0026] S122. Add a topological attention enhancement module during the decoding process; based on f cCalculate the attention weight matrix W and perform a linear transformation; then, based on f c , V, and the attention weight matrix W, generate the offset feature f sa , to reduce weakly related topologies and retain strongly related information; f sa Then, successively pass through a linear transformation, batch normalization, and the ReLU activation function to obtain the object-level topology feature f′ sa ; The formula is expressed as:

[0027] W = Softmax(Q·K T );

[0028] f sa = f c - V·W;

[0029] f′ sa = ReLU{BN[Linear(f sa )]};

[0030] Among them, Q, K, and V are the query, key, and value in the attention mechanism; Softmax(.) is the Softmax activation function; Linear(.) is the linear transformation; BN(.) is batch normalization.

[0031] Furthermore, step S2 specifically includes:

[0032] S21. Adopt the adaptive weighted sampling technique to adaptively adjust the sampling weights of various facade elements to alleviate the class imbalance problem, and thus sample the complete facade point cloud and divide it into multiple overlapping facade point cloud blocks;

[0033] S22. For the facade point cloud blocks obtained in step S21, learn multi-scale geometric features by stacking three hierarchical geometric extraction modules to effectively capture hierarchical geometric features and enhance the description of facade elements.

[0034] More specifically, the calculation process of the sampling weights in step S21 is as follows:

[0035] For the input building facade sample data, use i to represent a certain type of facade element, and p i represents the proportion of the sample size of this type of facade element in the overall building facade sample data; construct an adaptive weighted sampling function w(p i ) to balance the sampling weights between small-sample classes and large-sample classes; The formula of the adaptive weighted sampling function is expressed as:

[0036]

[0037] Among them, q 1 , q 2 , q 3 ; k1 , k 2 , k 3 ; v 1 , v 2 , v 3 All are piecewise curve coefficients; s and l represent the boundary values of the small sample class and the large sample class, and the boundary values are preliminarily determined based on the approximate proportions of the small sample class and the large sample class.

[0038] Then, calculate the sampling number n of each type of elevation element according to the sampling weight i :

[0039]

[0040] Among them, N represents the number of points in each elevation point cloud block, and C is the number of categories.

[0041] More specifically, step S22 specifically includes:

[0042] S221. In each hierarchical geometric extraction module, for each point x after adaptive weighted sampling i , search for its neighboring points through the K-nearest neighbor search method Then, obtain the edge feature ε through the edge convolution operation i , and the formula is expressed as:

[0043]

[0044] Among them, EConv(·) is the edge convolution operation; || is the feature concatenation operation;

[0045] S222. Adopt the strip pooling operation to enhance the long-distance geometric features; perform the strip pooling operation, convolution operation, and batch normalization in the horizontal and vertical directions in sequence, and obtain the horizontal feature f i h and the vertical feature f i v ; the formula is expressed as:

[0046] f i h = Repeat(BN{conv[SP h (ε i )]});

[0047] f i v = Repeat(BN{conv[SP v (ε i )]});

[0048] Among them, SP h (·), SP v(·) are horizontal strip pooling operation and vertical strip pooling operation respectively; Repeat(.) is a copy operation;

[0049] S223. After the horizontal feature and the vertical feature are superimposed, they pass through the ReLU activation function, convolution operation, and Sigmoid normalization in sequence to obtain the fused feature f i hv , and the formula is expressed as:

[0050] f i hv = Sigmoid{conv[ReLU(f i h + f i v )]};

[0051] Then calculate the inner product between the edge feature ε i and the fused feature f i hv , and add them point by point to obtain the extraction feature of the hierarchical geometric extraction block; the extraction feature of the previous hierarchical geometric extraction block is used as the input of the next hierarchical geometric extraction block;

[0052] S224. Aggregate the extraction features of the three hierarchical geometric extraction blocks, perform convolution operation and max pooling, and then copy them to the same dimension as the original feature for further aggregation and convolution, and finally output the hierarchical geometric feature at the point level

[0053] Furthermore, step S3 specifically includes:

[0054] S31. Construct a mapping index mechanism to accurately associate and aggregate the object-level topological feature and the hierarchical geometric feature;

[0055] S32. Use the channel attention mechanism to fuse the topological and geometric features, retain the important topological and geometric features when segmenting the facade elements, and achieve robust and accurate segmentation of the facade elements.

[0056] More specifically, step S31 specifically includes:

[0057] S311. Create a mapping index mechanism to associate each facade point cloud with the divided facade point cloud blocks; for a facade point cloud with an index of {I m |0 ≤ m ≤ d}, the corresponding facade point cloud block index is Use the facade point cloud block index as the value and the index of the facade point cloud to which it belongs as the key to construct a mapping mechanism; where d represents the number of facade point clouds, and j is the number of facade point cloud blocks divided by each facade point cloud

[0058] For each elevation point cloud block, its corresponding elevation point cloud is searched by the mapping mechanism index; for each point in the elevation point cloud block, its nearest point is searched in the elevation to form a pair; a pair of points after pairing is represented as: Among them, represents the r-th point in the point cloud block, represents the t-th point in its corresponding elevation;

[0059] S312. Connect the hierarchical geometric features carried by all points in each point cloud block with the object-level topological features carried by the corresponding elevation points to obtain the dual-feature fusion feature F(G, T):

[0060]

[0061] Among them, is an aggregation operation; K and L are the sets of indices of each point in the formed point pairs in the elevation and the point cloud block.

[0062] More specifically, step S32 specifically includes:

[0063] S321. Based on the dual-feature fusion feature F(G, T), calculate the context attention coefficient Φ, and the formula is expressed as:

[0064] Φ = Sigmoid{conv{Gpool[F(G, T)]}};

[0065] Among them, Gpool[·] represents the global average pooling operation;

[0066] S322. Use the context attention coefficient Φ as the feature weight, copy it to the same dimension as the original feature and multiply to make the finally extracted features focus on the most relevant information; finally, obtain the predicted values of each category through the MLP operation and the Dropout operation.

[0067] Based on the above technical solutions, the present invention has at least the following beneficial effects:

[0068] 1. Through the elevation topology extraction module, the topological features between building elevation elements are extracted, that is, the spatial structure information of the elevation elements and the spatial relationship between each elevation element are captured, which is beneficial to distinguish building elevation elements;

[0069] 2. Design an enhanced sampling geometric extraction module, combine adaptive weighted sampling and hierarchical geometric feature extraction, extract local and global geometric features, and improve the segmentation performance of small sample categories;

[0070] 3. Design a dual - feature attention fusion module to adaptively fuse topological and geometric features at different granularities. By combining the complementary advantages of the two types of features, the overall performance and robustness of facade point cloud segmentation are improved;

[0071] 4. Through the above three modules, the method proposed by the present invention can accurately and robustly identify building facade elements with severely imbalanced samples. Brief Description of the Drawings

[0072] Figure 1 It is the framework diagram of the topological - aware graph network FTG - Net constructed by the method proposed by the present invention;

[0073] Figure 2 It is the structural diagram of the facade topology extraction module in FTG - Net;

[0074] Figure 3 It is the curve graph of the adaptive weighted sampling function in the method proposed by the present invention;

[0075] Figure 4 It is the structural diagram of the enhanced sampling geometry extraction module in FTG - Net;

[0076] Figure 5 It is the structural diagram of the dual - feature attention fusion module in FTG - Net;

[0077] Figure 6 It is the schematic diagram of the building facade point cloud data set collected in the application example of the present invention;

[0078] Figure 7 It is the segmentation result graph of the test set in the application example of the present invention. Detailed Embodiments

[0079] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0080] Although the steps in the present invention are numbered, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0081] A method for segmenting building facade point clouds with class imbalance based on a topological - aware graph network proposed by the present invention provides a topological - aware graph network FTG - Net for segmenting building facade point clouds with class imbalance, as Figure 1As shown, it realizes the segmentation function through three modules. In the facade topology extraction module, it captures the abstract topological information between facade elements to enhance the description of facade elements; in the geometric extraction module with sampling enhancement, through adaptive sampling, it extracts hierarchical geometric features at the point level to further improve the description of small-sample facade elements; in the dual-feature attention fusion module, through adaptive fusion of topological and geometric features with different granularities and different types, it realizes the complementary advantages of the two features and improves the overall performance of facade segmentation.

[0082] The method proposed by the present invention can effectively segment the building facade point cloud with serious class imbalance, which specifically includes the following steps:

[0083] S1. Construct a facade topology extraction module to capture the abstract object-level topological features between facade elements;

[0084] As a preferred implementation manner, step S1 specifically includes:

[0085] S11. Use K-Means uniform sampling for each building facade sample data to obtain the facade point cloud, take the sampled facade point cloud as the input unit, and use the topological information encoder to encode the input unit to capture high-dimensional encoded features with strong discrimination;

[0086] As Figure 2 shown, in the topological information encoder:

[0087] The sampled facade point cloud first passes through two consecutive edge convolution operations to expand the receptive field to learn the global coarse-grained features containing topological information between facade elements; then through a three-layer convolution operation with dimension D 1 to perform dimension expansion, making up for the limitation that the edge convolution can only extract coarse-grained features, and obtaining high-dimensional encoded features with strong discrimination The formula is expressed as:

[0088]

[0089] Among them, is the point in the sampled facade point cloud, with a total of N; is its neighboring point, with a total of K; and respectively represent the facade input point and its neighboring point after the first edge convolution operation; EConv(·) and Conv (1) (·) are the edge convolution operation and the three-layer convolution operation with dimension D 1 respectively; || is the feature splicing operation; in this embodiment, take D 1 ={128, 256, 512};

[0090] S12. Use a topology information decoder enhanced by topological attention to extract the abstract object-level topological features between facade elements from the high-dimensional encoded information;

[0091] Similarly, as Figure 2 shown, the specific process of topology information decoding is as follows:

[0092] S121. Introduce and set a two-dimensional Riemannian square grid for the topology feature decoder; set the two-dimensional Riemannian square grid as a×a, then the number N of facade input points is a 2 From the high-dimensional encoded features and the two-dimensional Riemannian square grid F Grid learn the initial topology information f c , which is expressed by the formula:

[0093]

[0094] where Conv (2) (·) is a three-layer convolution operation with dimension D 2 ; || is the feature concatenation operation; in this embodiment, take a = 45, take D 2 ={512, 512, 3}

[0095] S122. Add a topology attention enhancement module during the decoding process; calculate the attention weight matrix W based on f c and perform a linear transformation; then generate the offset feature f c based on f sa , V and the attention weight matrix W to reduce weakly related topologies and retain strongly related information; f sa then successively pass through a linear transformation, batch normalization, and ReLU activation function to obtain the object-level topological feature f′ sa ; the formula is expressed as:

[0096] W = Softmax(Q·K T )

[0097] f sa = f c - V·W

[0098] f′ sa = ReLU{BN[Linear(f sa )]}

[0099] where Q, K, V are the query, key, and value in the attention mechanism; Softmax(.) is the Softmax activation function; Linear(.) is the linear transformation; BN(.) is the batch normalization;

[0100] In this embodiment, the facade topology extraction module is trained using the Adam optimizer, with an initial learning rate of 0.001, a training batch of 600, and a batch size of 2. Due to network limitations, 2025 three-dimensional points are input each time. In addition, this module ensures the integrity and accuracy of the topological information between the extracted facade objects by reconstructing the point cloud of the facade. At the same time, to constrain the facade topology extraction module, the chamfer distance is introduced to evaluate the similarity between the original point cloud P and the reconstructed point cloud , which is expressed by the formula:

[0101]

[0102] where dist(·) is the minimum distance between pairs of points in two point clouds. The first term represents the sum of the minimum distances from any point X in P to , and the second term represents the sum of the minimum distances from any point in to P

[0103] S2. Construct an enhanced sampling geometry extraction module to extract the hierarchical geometric features at the point level between facade elements

[0104] As a preferred implementation, step S2 specifically includes:

[0105] S21. Adopt the adaptive weighted sampling technique to adaptively adjust the sampling weights of various facade elements to alleviate the problem of class imbalance. After sampling, multiple overlapping facade point cloud blocks are obtained

[0106] The specific process of the sampling weight is as follows:

[0107] For the input building facade sample data, let i represent a certain type of facade element, and p i represents the proportion of the sample size of this type of facade element in the overall building facade sample data. Construct an adaptive weighted sampling function w(p i ) to balance the sampling weights between the small-sample class and the large-sample class. The formula of the adaptive weighted sampling function is expressed as:

[0108]

[0109] where q 1 , q 2 , q 3 ; k 1 , k 2 , k 3 ; v 1 , v 2 , v 3They are all piecewise curve coefficients; s and l represent the boundary values of the small-sample class and the large-sample class, and the boundary values are initially determined based on the approximate proportions of the small-sample class and the large-sample class; in this embodiment, as Figure 3 shown, a curve graph of the adaptive weighted sampling function is given

[0110] Then, the sampling number n of each type of facade element is calculated according to the sampling weights i :

[0111]

[0112] where, N represents the number of points in each facade point cloud block, and C is the number of categories;

[0113] S22. For the facade point cloud blocks obtained in step S21, three hierarchical geometric extraction modules are stacked to learn multi-scale geometric features, so as to effectively capture hierarchical geometric features and enhance the description of facade elements;

[0114] More specifically, as Figure 4 shown, the sampled facade point cloud is divided into multiple facade point cloud blocks; and for each point in the facade point cloud block, the process of capturing hierarchical geometric features is as follows:

[0115] S221. In each hierarchical geometric extraction module, for each point x i after adaptive weighted sampling, its neighboring points are searched through the K-nearest neighbor search method Then, edge features ε i are obtained through edge convolution operations, and the formula is expressed as:

[0116]

[0117] where, EConv(·) is the edge convolution operation; || is the feature concatenation operation;

[0118] S222. Strip pooling operations are adopted to enhance long-distance geometric features; strip pooling operations, convolution operations, and batch normalization are sequentially performed in the horizontal and vertical directions respectively to obtain horizontal feature f i h and vertical feature f i v ; the formula is expressed as:

[0119] f i h = Repeat(BN{conv[SP h (ε i )]});

[0120] f i v = Repeat(BN{conv[SP v(ε i )]});

[0121] Among them, SP h (·) and SP v (·) are horizontal strip pooling operation and vertical strip pooling operation respectively; Repeat(.) is a replication operation;

[0122] After the horizontal feature and the vertical feature are superimposed, they pass through the ReLU activation function, the convolution operation, and the Sigmoid normalization in sequence to obtain the fused feature f i hv , and the formula is expressed as:

[0123] f i hv =Sigmoid{conv[ReLU(f i h +f i v )]};

[0124] Then calculate the inner product between the edge feature ε i and the fused feature f i hv , and add them point by point to obtain the extraction feature of the hierarchical geometric extraction block; the extraction feature of the previous hierarchical geometric extraction block is used as the input of the next hierarchical geometric extraction block;

[0125] S224. Aggregate the extraction features of the three hierarchical geometric extraction blocks, perform convolution operation and max pooling, and then copy them to the same dimension as the original feature for further aggregation and convolution, and finally output the hierarchical geometric feature at the point level

[0126] In this embodiment, the enhanced sampling geometric extraction module is trained using the SGD optimizer, the initial learning rate is set to 0.001, the training batch is 100, and the batch size is 8; considering the high point density and the large size of the facade scene, the facade is divided into point cloud blocks of 6m×6m, and the overlap between adjacent blocks is 3m; at the same time, the cross-entropy loss function is used to constrain the parameters of the network enhanced sampling geometric extraction module. This function can effectively measure the difference between the predicted probability distribution and the true label distribution, especially performs well in dealing with classification and segmentation problems, and is denoted as L CE , and the specific formula is expressed as:

[0127]

[0128] Among them, N and C respectively represent the number of samples and the number of categories, X uw represents the true label of the u-th sample in the w-th category, and p uw represents the probability that the model predicts that the u-th sample belongs to the w-th category;

[0129] In addition, since the Lovász-Softmax loss can significantly optimize the mIoU index, thereby reducing the impact of category imbalance, this embodiment also constructs a Lovász-Softmax loss function:

[0130]

[0131] Among them, e (i) represents the ranking error for each category i, which is obtained based on the difference between the predicted probability and the true label; y represents the one-hot encoded true label of the point; Δ i (y) represents the IoU change of category i;

[0132] The above two loss functions and the chamfer distance are combined as the total loss function of the facade topology extraction module and the enhanced sampling geometry extraction module to train and obtain the optimal model; the total loss function is expressed as:

[0133]

[0134] Among them, α and β are balance coefficients.

[0135] S3, construct a dual-feature attention fusion module to fuse the two features of different granularity and type obtained in steps S1 and S2 to achieve complementary advantages of the two features;

[0136] As a preferred implementation, step S3 specifically includes:

[0137] S31, build a mapping index mechanism to accurately associate and aggregate object-level topological features and hierarchical geometric features;

[0138] More specifically, Figure 5 As shown, step S31 specifically includes:

[0139] S311, create a mapping index mechanism to associate each facade point cloud with its divided facade point cloud blocks; for an index {I m |0≤m≤d}, the corresponding facade point cloud block index is The facade point cloud block index is used as the value and the facade point cloud index to which it belongs is used as the key to construct a mapping mechanism; where d represents the number of facade point clouds, and j represents the number of facade point cloud blocks into which each facade point cloud is divided.

[0140] For each facade point cloud block, the mapping mechanism indexes and searches for its corresponding facade point cloud; for each point in the facade point cloud block, search for its nearest point in the facade to form a pair; a pair of paired points is expressed as: in, represents the r-th point in the point cloud block, and represents the t-th point on its corresponding facade;

[0141] S312. Connect the hierarchical geometric features carried by all points in each point cloud block with the object-level topological features carried by the corresponding facade points to obtain the dual-feature fusion feature F(G, T):

[0142]

[0143] where, is an aggregation operation; K and L are the sets of indices of each point in the formed point pairs in the facade and the point cloud block.

[0144] S32. Adopt a channel attention mechanism to fuse topological and geometric features, retain the important topological and geometric features during the segmentation of facade elements, and achieve robust and accurate segmentation of facade elements;

[0145] Similarly, as Figure 5 shown, step S32 specifically includes:

[0146] S321. Based on the dual-feature fusion feature F(G, T), calculate the context attention coefficient Φ, and the formula is expressed as:

[0147] Φ = Sigmoid{conv{Gpool[F(G, T)]}};

[0148] where, Gpool[·] represents the global average pooling operation;

[0149] S322. Use the context attention coefficient Φ as the feature weight, copy it to the same dimension as the original feature and multiply to make the finally extracted features focus on the most relevant information; finally, obtain the predicted values of each category through the MLP operation and the Dropout operation.

[0150] S4. The facade topology extraction module, the enhanced sampling geometry extraction module, and the dual-feature attention fusion feature module are combined to form the topology-aware graph network FTG-Net; after dividing the building facade sample data into a training set, a validation set, and a test set and inputting them into FTG-Net, the overall network is trained and then tested and the results are evaluated to achieve the segmentation of the building facade point cloud with uneven categories.

[0151] In this embodiment, according to the designed total loss function, after obtaining the optimal models of the facade topology extraction module and the enhanced sampling geometry extraction module through the training set and the validation set, the test set is input, and the object-level topological features and hierarchical geometric features are extracted through the optimal models. After passing through the dual-feature attention fusion feature module, the classification of facade elements is finally achieved.

[0152] The effectiveness of the method proposed by the present invention will be verified below in combination with specific application examples.

[0153] In this application example, two sets of building facade point cloud datasets are created, as Figure 6 shown: a campus dataset and an urban street dataset. These two datasets are collected by different terrestrial laser scanners and have diverse facade styles. The datasets are labeled by combining an automatic segmentation algorithm with manual annotation. During the annotation process, in order to facilitate learning the topological information between facade elements, each facade element needs to be labeled separately, rather than as a whole. In the labeled facade dataset, 51 facades are used for training, 18 for validation, and 14 for testing.

[0154] To evaluate the facade segmentation, the following metrics are used in this application example: precision, recall, F1-score, intersection over union (IoU) for each category, and overall accuracy (OA) for all samples. The specific calculation formulas are as follows:

[0155]

[0156] where TP, FN, and FP are the true positive, false negative, and false positive counts, respectively.

[0157] After inputting the test set, the results of the building facade point cloud segmentation (or facade element classification) are as Figure 7 shown; Figure 7 (1)-(8) in are the point cloud segmentation result diagrams of 8 different building facades in the campus dataset, Figure 7 (9)-(16) in are the point cloud segmentation result diagrams of 8 different building facades in the urban street dataset; different colors in the figure represent different facade elements, including windows, walls, balconies, and doors; (it should be noted that the colors of each facade element in the appendix Figures 1 - 7 are only used for distinction from other facade elements and do not mean that each facade element is fixed to only one color).

[0158] The corresponding metrics for each building facade point cloud segmentation result are shown in Table 1 below:

[0159] Table 1 Metrics for each facade segmentation result

[0160]

[0161]

[0162]

[0163] From Table 1 and Figure 7As can be seen, for the uneven building facades of multiple selected categories, the method proposed by the present invention can provide detailed facade segmentation accuracy results, and has excellent performance in the accuracy of each facade small sample category and other categories. The facade segmentation result indicators in Table 1 demonstrate the accuracy, stability and precision of the method proposed by the present invention; in summary, the method proposed by the present invention can effectively solve the problem of uneven facade point cloud segmentation (i.e., element classification).

[0164] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0165] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for segmenting building facade point clouds with uneven categories based on a topological perception graph network, characterized in that: The specific steps include: S1, build a facade topology extraction module to capture the abstract object-level topological features between facade elements; S2, build an enhanced sampling geometry extraction module to extract hierarchical geometric features at the point level between facade elements; S3, construct a dual-feature attention fusion module to fuse the two features of different granularity and type obtained in steps S1 and S2 to achieve complementary advantages of the two features; S4, the facade topology extraction module, the enhanced sampling geometry extraction module, and the dual-feature attention fusion feature module are combined to form the topological perception graph network FTG-Net; the building facade sample data is divided into training set, validation set and test set and then input into FTG-Net. After the overall network is trained, it is tested and the results are evaluated to achieve the segmentation of uneven building facade point clouds.

2. The method for segmenting building facade point clouds with uneven categories according to claim 1, characterized in that: Step S1 specifically includes: S11. K-Means uniform sampling is used for each building facade sample data to obtain a complete facade point cloud. The sampled facade point cloud is used as an input unit, and a topological information encoder is used to encode the input unit to capture high-dimensional encoding features with strong discrimination. S12. Use the topological information decoder enhanced by topological attention to extract abstract object-level topological features between facade elements from high-dimensional encoded information.

3. The method for segmenting building facade point clouds with uneven categories according to claim 2, characterized in that: Step S11 is specifically as follows: The sampled facade point cloud is first subjected to two consecutive edge convolution operations to expand the receptive field in order to learn the global coarse-grained features containing topological information between facade elements; Then, the dimension is expanded through three layers of convolution operation with dimension D1 to make up for the limitation that edge convolution can only extract coarse-grained features, and obtain high-dimensional coding features with strong discrimination. The formula is: in, is the point in the sampled facade point cloud, a total of N; are its neighboring points, a total of K; and They represent the face input point and its neighboring points after the first edge convolution operation; EConv(·) and Conv (1) (·) are the edge convolution operation and the three-layer convolution operation with dimension D1; || is the feature concatenation operation.

4. The method for segmenting building facade point clouds with uneven categories according to claim 2, characterized in that: Step S12 specifically includes: S121. A two-dimensional Riemann square grid is introduced for the topological feature decoder; the two-dimensional Riemann square grid is set to a×a, then the number of elevation input points N is a 2 ; composed of high-dimensional encoding features F and two-dimensional Riemann square grid F Grid Learn the initial topology information f c , the formula is: Among them, Conv (2) (·) is a three-layer convolution operation with dimension D2; || is a feature concatenation operation; S122, add a topological attention enhancement module in the decoding process; based on f c Calculate the attention weight matrix W and perform linear transformation; then according to f c , V and the attention weight matrix W, generate the offset feature f sa , to reduce weakly correlated topology and retain strongly correlated information; f sa Then, the object-level topological feature f is obtained through linear transformation, batch normalization and ReLU activation function. s ' a ; The formula is: W=Softmax(Q·K T ); f sa =f c -V·W; f s ' a =ReLU{BN[Linear(f sa )]}; Among them, Q, K, V are the query, key, and value in the attention mechanism; Soft max(.) is the Softmax activation function; Linear(.) is the linear transformation; BN(.) is batch normalization.

5. The method for segmenting building facade point clouds with uneven categories according to claim 1, characterized in that: Step S2 specifically includes: S21, using adaptive weighted sampling technology to adaptively adjust the sampling weights of various facade elements to alleviate the category imbalance problem, thereby sampling the complete facade point cloud and dividing it into multiple overlapping facade point cloud blocks; S22. For the facade point cloud block obtained in step S21, multi-scale geometric features are learned by stacking three hierarchical geometry extraction modules to effectively capture hierarchical geometric features and enhance the description of facade elements.

6. The method for segmenting building facade point clouds with uneven categories according to claim 5, characterized in that: The calculation process of the sampling weight in step S21 is specifically as follows: For the input building facade sample data, i is used to represent a certain type of facade element, p i Indicates the proportion of the sample capacity of this type of facade elements to the sample data of the entire building facade; constructs an adaptive weighted sampling function w(p i ) to balance the sampling weights between small sample classes and large sample classes; the formula of the adaptive weighted sampling function is expressed as: Among them, q1, q2, q3; k1, k2, k3; v1, v2, v3 are all piecewise curve coefficients; s and l represent the boundary values ​​of the small sample class and the large sample class, and the boundary values ​​are preliminarily determined based on the approximate ratio of the small sample class to the large sample class; Then calculate the sampling number n of each type of facade element according to the sampling weight i : Among them, N represents the number of points in each facade point cloud block, and C is the number of categories.

7. The method for segmenting building facade point clouds with uneven categories according to claim 5, characterized in that: Step S22 is specifically as follows: S221, in each hierarchical geometry extraction module, for each point x after adaptive weighted sampling i , search its neighboring points through K-nearest neighbor search method Then the edge feature ε is obtained through the edge convolution operation i , the formula is: Among them, EConv(·) is the edge convolution operation; || is the feature concatenation operation; S222, use strip pooling operation to enhance long-distance geometric features; perform strip pooling operation, convolution operation and batch normalization in the horizontal and vertical directions respectively, and obtain the horizontal feature f i h and vertical feature f i v ; The formula is: f i h =Repeat(BN{conv[SP h (ε i )]}); f i v =Repeat(BN{conv[SP v (ε i )]}); Among them, SP h (·), SP v (·) are horizontal stripe pooling operation and vertical stripe pooling operation respectively; Repeat(.) is the copy operation; S223, after the horizontal features and vertical features are superimposed, they are successively subjected to the ReLU activation function, convolution operation, and Sigmoid normalization to obtain the fusion feature f i hv , the formula is: f i hv =Sigmoid{conv[ReLU(f i h +f i v )]}; Then calculate the edge feature ε i And the fusion feature f i hv The inner product between them is calculated and added point by point to obtain the extracted features of the hierarchical geometry extraction block; the extracted features of the previous hierarchical geometry extraction block are used as the input of the next hierarchical geometry extraction block; S224, aggregate the extracted features of the three hierarchical geometry extraction blocks, use convolution operation and perform maximum pooling, then copy them to the same dimension as the original features, aggregate and convolve again, and finally output the hierarchical geometry features at the point level 8. The method for segmenting building facade point clouds with uneven categories according to claim 1, characterized in that: Step S3 specifically includes: S31, build a mapping index mechanism to accurately associate and aggregate object-level topological features and hierarchical geometric features; S32. The channel attention mechanism is used to fuse topological and geometric features, retaining important topological and geometric features when segmenting facade elements, and achieving robust and accurate segmentation of facade elements.

9. The method for segmenting building facade point clouds with uneven categories according to claim 8, characterized in that: Step S31 specifically includes: S311, create a mapping index mechanism to associate each facade point cloud with its divided facade point cloud blocks; for an index {I m |0≤m≤d}, the corresponding facade point cloud block index is A mapping mechanism is constructed by taking the facade point cloud block index as the value and the facade point cloud index to which it belongs as the key; wherein d represents the number of facade point clouds, and j represents the number of facade point cloud blocks into which each facade point cloud is divided. For each facade point cloud block, the mapping mechanism indexes and searches for its corresponding facade point cloud; for each point in the facade point cloud block, search for its nearest point in the facade to form a pair; a pair of paired points is expressed as: in, represents the rth point in the point cloud block, represents the tth point of its corresponding facade; S312: hierarchical geometric features carried by all points in each point cloud block Object-level topological features carried by corresponding facade points Connect and get the dual feature fusion feature F(G,T): in, is an aggregation operation; K and L are the collections of the indices of the points in the point pairs constructed in the facade and point cloud blocks.

10. The method for segmenting building facade point clouds with uneven categories according to claim 9, characterized in that: Step S32 specifically includes: S321. Based on the dual-feature fusion feature F(G,T), calculate the context attention coefficient Φ, which is expressed as: Φ=Sigmoid{conv{Gpool[F(G,T)]}}; Where Gpool[·] represents the global average pooling operation; S322. Use the context attention coefficient Φ as the feature weight, copy it to the same dimension as the original feature and multiply it so that the final extracted feature focuses on the most relevant information; finally, obtain the predicted value of each category through MLP operation and Dropout operation.

Citation Information

Patent Citations

  • Building facade structure extraction method based on multi-scale dynamic graph convolution

    CN116977572A

  • Arterial aneurysm 3D point cloud automatic segmentation method combined with geometric topology analysis

    CN119379721A

  • Single building three-dimensional reconstruction method based on point cloud semantic segmentation and structure fitting

    WO2024077812A1