Automatic extraction of buildings from high-resolution remote sensing images in complex urban environments
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-08-11
AI Technical Summary
[0002]复杂城市环境下,高分辨率遥感影像包括丰富的光谱、空间结构及纹理信息,为建筑物提取提供了数据支撑,但城市区域建筑物分布密集、形态多样,且受地形遮挡、光谱混淆、尺度差异等因素影响,建筑物特征与背景信息的区分难度显著增加
[0015]有益效果:本发明提出复杂城市环境下高分辨率遥感影像建筑物自动提取方法,通过利用旋转不变性代价图聚合网络、边缘约束多任务分割模型、尺度一致性自适应调整算法及多源遥感特征融合分析平台,形成全流程协同的建筑物自动提取技术体系,针对特征处理与聚合能力不足的问题,通过多方向旋转特征代价矩阵构建与动态聚合,充分捕捉建筑物旋转不变特征,结合多尺度特征匹配与自适应调整,实现不同方向、不同尺度建筑物特征的全面精准表达,大幅提升建筑物与相似光谱背景地物的区分度;针对分割模型与特征融合机制不完善的问题,借助边缘梯度特征强化分割边界辨识度,通过层级化多源特征融合策略充分挖掘主遥感特征与多源辅助特征的互补价值,细化分割边界并优化特征空间结构,显著降低复杂城市环境中的漏检、误检概率,精准还原建筑物实际形态与空间分布;整套方法通过分步骤细化的特征处理、分割、调整与融合流程,实现高分辨率遥感影像建筑物提取的自动化与精准化,适配密集分布、形态多样的复杂城市环境,为城市规划、灾害应急、资源管理等领域提供高效可靠的空间信息支持。
Smart Images

Figure CN121962945B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of building extraction technology from remote sensing images, and in particular to an automatic method for extracting buildings from high-resolution remote sensing images in complex urban environments. Background Technology
[0002] In complex urban environments, high-resolution remote sensing imagery contains rich spectral, spatial structure, and texture information, providing data support for building extraction. However, urban areas are characterized by dense building distribution, diverse forms, and are significantly affected by factors such as terrain occlusion, spectral obfuscation, and scale differences, making it much more difficult to distinguish building features from background information. With the increasing demand for building spatial information in fields such as urban planning, disaster emergency response, and resource management, traditional extraction methods relying on manual processes or simple algorithms are no longer sufficient to meet the needs of efficient and accurate applications. There is an urgent need to construct an automated extraction technology system that integrates multi-source features and adapts to complex environments, enabling rapid identification and accurate segmentation of building areas in high-resolution remote sensing imagery, and providing reliable data support for decision-making in related fields.
[0003] Existing technologies for automatic building extraction from high-resolution remote sensing images have two prominent drawbacks: First, they lack sufficient feature processing and aggregation capabilities, failing to fully consider the rotational invariance and multi-scale consistency of building features. This results in poor adaptability to building features at different directions and scales, leading to incomplete feature representation and difficulty in effectively distinguishing buildings from background features with similar spectral characteristics. Second, the segmentation model and feature fusion mechanism are imperfect, lacking effective edge constraints and multi-source feature hierarchical fusion strategies. This results in blurred segmentation boundaries and failure to fully utilize the complementary information of multi-source auxiliary features. Consequently, building extraction results are prone to missed detections and false detections in complex urban environments, making it impossible to accurately reconstruct the actual shape and spatial distribution of buildings. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides a method for automatic extraction of buildings from high-resolution remote sensing images in complex urban environments.
[0005] The technical solution adopted in this invention is an automatic building extraction method for high-resolution remote sensing images in complex urban environments, comprising the following steps: S1, acquiring a dataset of spectral features, spatial structure features, and texture features of high-resolution remote sensing images, and selecting feature variables with building correlation higher than a preset threshold to construct an original feature set; S2, performing feature mapping on the original feature set through a rotation-invariant cost map aggregation network to generate a multi-directional rotation feature cost matrix, and dynamically aggregating the cost map by combining neighborhood feature correlation; S3, using an edge-constrained multi-task segmentation model to perform preliminary segmentation of building regions on the aggregated feature cost map, and enhancing the segmentation boundary identification through edge gradient features; S4, using a scale consistency adaptive adjustment algorithm to perform multi-scale feature matching and adjustment on the preliminary segmentation results, optimizing the building feature expression at different scales; S5, based on a multi-source remote sensing feature fusion analysis platform, hierarchically fusing the adjusted features with multi-source auxiliary remote sensing features to construct a fused feature space; S6, based on the unique distribution pattern of building features in the fused feature space, outputting the building extraction results through feature clustering and boundary optimization processing.
[0006] Furthermore, the characteristic aggregation expression of the rotation-invariant cost graph aggregation network is: ,in, This is the rotation-invariant feature map after aggregation. The number of rotation angles. For the first Weighting coefficients for each rotation angle, For angle The rotation transformation matrix, For the original feature set, For the first Convolution kernels corresponding to each rotation angle, The edge constraint coefficient, For gradient operators, For feature cost map, For element-wise multiplication, This is the neighborhood feature correlation matrix.
[0007] Furthermore, the segmentation expression of the edge-constrained multi-task segmentation model is as follows: ,in, This is the preliminary segmentation result. For activation function, For depthwise convolution operations, This is the rotation-invariant feature map after aggregation. For edge feature weights, For edge gradient features, For multi-task segmentation weight matrix, The boundary penalty coefficient, For boundary feature mapping function, These are the a priori features of the building boundary.
[0008] Furthermore, the adjustment expression of the scale consistency adaptive adjustment algorithm is: ,in, This is the scaled feature map. The number of scale levels. For the first A weight matrix of several scales, This is the preliminary segmentation result. For the first Adaptive coefficients at each scale For the first Feature mapping function at each scale The consistency constraint coefficient is... This is a scale consistency check function. This is the set of scale consistency parameters.
[0009] Furthermore, the feature fusion expression of the multi-source remote sensing feature fusion analysis platform is as follows: ,in, To merge feature spaces, For the number of multi-source auxiliary features, For the first Class principal feature weights, This is the scaled feature map. For the first Class auxiliary feature weights, For the first Multi-source assisted remote sensing features This is a feature correlation measurement function.
[0010] Furthermore, the boundary optimization expression for the building extraction results is as follows: ,in, For the final building extraction results, For the feature clustering matrix, For boundary optimization coefficients, It is a second-order gradient operator. To merge feature spaces, Optimize the weight matrix for the boundary.
[0011] Further, step S2 includes steps S21 to S23: S21, selecting calibrated feature variables from the original feature set of high-resolution remote sensing images, grouping them according to feature dimensions and building association characteristics to form multiple feature subsets; S22, inputting different feature subsets into a rotation-invariant cost map aggregation network, performing multi-directional rotation transformation on each feature subset through the network's hidden layer to generate feature cost components at corresponding angles; S23, calculating the association weights of different cost components based on neighborhood feature association, and aggregating all cost components through a dynamic weighted summation method to form a complete multi-directional rotation feature cost matrix.
[0012] Further, step S3 includes steps S31 to S34: S31, extracting edge gradient information from the aggregated feature cost map, constructing an edge feature vector set, and clarifying the boundary difference features between the building area and the background area; S32, inputting the feature cost map and the edge feature vector set into the edge-constrained multi-task segmentation model, starting the model's multi-task parallel computing process, and simultaneously segmenting the building area and the boundary area respectively; S33, integrating the edge feature vector set into the segmentation calculation process through the model's built-in constraint mechanism, enhancing the feature recognition of the boundary area; S34, integrating the multi-task segmentation output results to form a preliminary building area segmentation map, retaining the detailed features of the segmentation boundary.
[0013] Furthermore, S4 includes steps S41 to S43: S41, analyzing the differences in feature representation of buildings at different scales in the preliminary segmentation results, and determining multiple calibration scale levels and corresponding feature parameters; S42, activating the scale consistency adaptive adjustment algorithm to extract and quantize the features of buildings at each scale level separately; S43, based on the scale consistency criterion, mutually verifying and dynamically adjusting the feature quantization results of different scales, correcting feature deviations caused by scale mismatch, and outputting a building feature map of a unified scale.
[0014] Furthermore, S5 includes steps S51 to S53: S51, importing multi-source auxiliary feature data corresponding to the main remote sensing image through a multi-source remote sensing feature fusion analysis platform, including terrain features, spectral auxiliary features, and spatial correlation features; S52, performing dimensional alignment processing on the scale-adjusted feature map and the multi-source auxiliary feature data to ensure the fusionability of different types of features in the same feature space; S53, adopting a hierarchical fusion strategy, first performing internal aggregation of similar features, and then performing cross-level fusion of aggregated features of different categories, gradually constructing a fusion feature space including multi-dimensional information, while preserving the uniqueness and integrity of building features.
[0015] Beneficial Effects: This invention proposes an automatic building extraction method for high-resolution remote sensing images in complex urban environments. By utilizing a rotation-invariant cost map aggregation network, an edge-constrained multi-task segmentation model, a scale-consistent adaptive adjustment algorithm, and a multi-source remote sensing feature fusion analysis platform, a fully collaborative automatic building extraction technology system is formed. Addressing the issue of insufficient feature processing and aggregation capabilities, it fully captures rotation-invariant building features through the construction and dynamic aggregation of multi-directional rotational feature cost matrices. Combined with multi-scale feature matching and adaptive adjustment, it achieves comprehensive and accurate representation of building features at different directions and scales, significantly improving the distinguishability of buildings from similar spectral background features. To address the shortcomings of the segmentation model and feature fusion mechanism, this method enhances the segmentation boundary identification by leveraging edge gradient features. It fully exploits the complementary value of primary remote sensing features and multi-source auxiliary features through a hierarchical multi-source feature fusion strategy, refining segmentation boundaries and optimizing feature spatial structure. This significantly reduces the probability of missed and false detections in complex urban environments, accurately restoring the actual form and spatial distribution of buildings. The entire method, through a step-by-step, refined feature processing, segmentation, adjustment, and fusion process, automates and accurately extracts buildings from high-resolution remote sensing images. It is adaptable to densely distributed and morphologically diverse complex urban environments, providing efficient and reliable spatial information support for urban planning, disaster emergency response, resource management, and other fields. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the overall steps of the method of the present invention; Figure 2 This is a flowchart of method step S2 of the present invention; Figure 3 This is a flowchart of method step S3 of the present invention; Figure 4 This is a flowchart of method step S4 of the present invention; Figure 5 This is a flowchart of step S5 of the method of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] like Figure 1 As shown, a method for automatic building extraction from high-resolution remote sensing images in complex urban environments is characterized by the following steps: S1. Obtain the spectral features, spatial structure features, and texture features dataset of high-resolution remote sensing images, and select feature variables with building correlation higher than a preset threshold to construct the original feature set; Specifically, step S1 constructs a high-quality set of original features to lay the data foundation for subsequent extraction processes. During implementation, high-resolution remote sensing image data with a spatial resolution range of 0.5 to 2 meters is first collected, including multiple spectral bands such as visible light and near-infrared. Feature extraction tools are used to separate three main categories of core features from the images: spectral features, spatial structure features, and texture features. Spectral features include 12 to 15 variable features such as band reflectance and band combination index; spatial structure features include 8 to 10 variables such as building aspect ratio, area, and contour complexity; and texture features include 10 to 12 variables such as the mean of the gray-level co-occurrence matrix, entropy value, and contrast. These three categories of features together form 30 to 37 original feature variables. Subsequently, the feature correlation analysis method was used to calculate the Pearson correlation coefficient between each feature variable and the building target. The correlation threshold was set to 0.65, and feature variables with correlation coefficients higher than this threshold were selected. Redundant and irrelevant features were removed, and finally an original feature set including 18 to 25 core features was constructed. This set not only retains the key feature information of the building, but also reduces the complexity of subsequent calculations, providing reliable initial data support for accurate building extraction.
[0019] S2, the original feature set is mapped by a rotation-invariant cost graph aggregation network to generate a multi-directional rotated feature cost matrix, and the cost graph is dynamically aggregated by combining the neighborhood feature correlation. Specifically, step S2 uses a rotation-invariant cost graph aggregation network to perform deep processing and aggregation of features, improving the adaptability of features to changes in building rotation. In the implementation process, the original feature set constructed in S1 is first divided into 3 to 4 feature subsets according to feature type. Each subset includes 5 to 8 highly correlated feature variables. Each feature subset is input into the rotation-invariant cost graph aggregation network, which has 6 hidden layers: 4 convolutional layers and 2 pooling layers. The convolutional kernel sizes are 3×3 and 5×5, with numbers of 64, 128, 256, and 256 respectively. The pooling layers use max pooling with a 2×2 kernel size. The network performs rotation transformations on each feature subset in eight directions: 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315°, generating feature cost components for the corresponding angles. The dimension of each cost component is consistent with the input feature subset. Subsequently, the correlation of neighborhood features between each cost component is calculated. The correlation weight calculation window is set to 5×5. The weight coefficients of features at different locations are assigned based on the Gaussian weighting function. The weight values range from 0.1 to 0.9. The cost components in eight directions are aggregated by dynamic weighted summation to form a feature cost matrix with unchanged dimensions but including multi-directional rotation information. This matrix can effectively resist feature variation caused by building rotation and improve the stability of feature representation.
[0020] S3 uses an edge-constrained multi-task segmentation model to perform preliminary segmentation of building areas on the aggregated feature cost map, and enhances the segmentation boundary identification through edge gradient features. Specifically, step S3 uses an edge-constrained multi-task segmentation model to perform preliminary segmentation of building regions, enhancing the recognizability of segmentation boundaries. During implementation, edge gradients are first calculated on the feature cost matrix output from S2. The Sobel operator is used to extract the horizontal and vertical gradient values of pixels in the matrix. A gradient threshold of 0.3 is set, and pixels with gradient values higher than this threshold are selected as edge candidate points, constructing an edge feature vector set. The feature cost matrix and the edge feature vector set are then input into the edge-constrained multi-task segmentation model. The model adopts an encoder-decoder architecture. The encoder includes 5 convolutional blocks, each consisting of 2 convolutional layers and 1 batch normalization layer. The decoder performs feature upsampling through transposed convolution and sets 4 multi-task output branches, corresponding to building region segmentation, boundary segmentation, background suppression, and feature enhancement tasks, respectively. During model training, the batch size is set to 16, the number of iterations is 200 rounds, the initial learning rate is 0.001, decaying to 0.5 every 50 rounds, and the edge constraint weight coefficient is set to 0.8. The loss value is calculated by weighting the cross-entropy loss function and the edge loss function. During model execution, multi-task branches perform parallel computation, integrating edge feature vector sets into the computation process of each branch. The feature response of the boundary region is enhanced through the edge constraint mechanism. Finally, the output results of each branch are integrated to generate a preliminary building region segmentation map. The pixel classification accuracy of the building region and the background region in the segmentation map reaches more than 85%, and the boundary contours are initially clear.
[0021] S4. The scale consistency adaptive adjustment algorithm is used to perform multi-scale feature matching and adjustment on the preliminary segmentation results to optimize the feature representation of buildings at different scales. Specifically, step S4 utilizes a scale consistency adaptive adjustment algorithm to optimize the feature representation of buildings at multiple scales, addressing the issue of uneven extraction accuracy across different scales. During implementation, the preliminary segmentation results output from S3 are first analyzed for scale, dividing the data into three scale levels: small scale (area less than 100 pixels), medium scale (area between 100 and 500 pixels), and large scale (area greater than 500 pixels). Each scale level has a corresponding feature extraction window of 3×3, 7×7, and 11×11, respectively. The scale consistency adaptive adjustment algorithm is then activated, extracting features separately for each scale level's segmentation results, including shape features, texture features, and spatial distribution features. 12 to 15 feature parameters are extracted for each scale level. The algorithm sets a scale consistency verification threshold of 0.75. By calculating the similarity between features at different scale levels, the consistency of feature representation is determined. Features with similarity below the threshold are dynamically adjusted, with the adjustment coefficient set according to scale differences: 0.9 between small and medium scales, and 0.85 between medium and large scales. Meanwhile, the algorithm introduces a scale compensation mechanism to enhance the features of small-scale buildings with an enhancement coefficient of 1.2, and to smooth the features of large-scale buildings with a smoothing window of 5×5. Through matching, verification, adjustment and compensation of multi-scale features, a building feature map of a uniform scale is output. This feature map can balance the feature expression of buildings of different scales and improve the consistency of the overall extraction.
[0022] S5, based on the multi-source remote sensing feature fusion analysis platform, hierarchically fuses the adjusted features with multi-source auxiliary remote sensing features to construct a fused feature space; S6, based on the unique distribution pattern of building features in the fused feature space, outputs the building extraction results through feature clustering and boundary optimization.
[0023] Specifically, step S6 outputs the final building extraction results through feature clustering and boundary optimization, ensuring the accuracy and completeness of the extraction results. In implementation, feature clustering is first performed on the fused feature space constructed in S5 using the K-means clustering algorithm, with a cluster size of 2 (corresponding to the building area and background area respectively). The initial cluster centers are selected through random sampling, and the number of iterations is set to 100. The clustering termination condition is that the change in cluster centers is less than 0.001. After clustering, preliminary building pixel classification results are obtained. Subsequently, boundary optimization is performed using a combination of morphological dilation and erosion operations. Both the dilation kernel and the erosion kernel are set to 3×3. First, dilation is performed once to fill the internal holes of the area, and then erosion is performed once to eliminate boundary burrs. Simultaneously, the contour smoothness of the building areas in the classification results is calculated, with a smoothness threshold set to 0.8. Contours with smoothness below this threshold are smoothed using a multinomial fitting method with a fitting order of 3. Furthermore, an area filtering threshold is set to remove isolated areas with an area less than 30 pixels to avoid false detections. The final output of the building extraction result is a binary image, in which the pixel value of the building area is 1 and the pixel value of the background area is 0. The overall accuracy of the extraction result reaches more than 90%, and the boundary position error is controlled within 2 pixels, which can accurately restore the actual shape and spatial distribution of the building.
[0024] Preferably, the feature aggregation expression of the rotation-invariant cost graph aggregation network is: ,in, This is the rotation-invariant feature map after aggregation. The number of rotation angles. For the first Weighting coefficients for each rotation angle, For angle The rotation transformation matrix, For the original feature set, For the first Convolution kernels corresponding to each rotation angle, The edge constraint coefficient, For gradient operators, For feature cost map, For element-wise multiplication, This is the neighborhood feature correlation matrix.
[0025] Specifically, the feature aggregation logic of the rotation-invariant cost map aggregation network generates rotation-invariant feature maps through multi-directional rotation mapping and dynamic weighted aggregation. In implementation, the number of rotation angles is initially set to eight, covering a uniform distribution from 0° to 315°. The weight coefficient for each rotation angle is calculated based on the feature correlation, ranging from 0.1 to 0.9, with angles showing stronger building feature responses corresponding to weight coefficients higher than 0.7. The rotation transformation matrix is constructed based on the trigonometric function relationships of the corresponding angles to ensure that the features maintain spatial structural integrity during rotation. After rotation transformation, the original feature set is convolved with the corresponding convolution kernel. The kernel size is set to 3×3 or 5×5 according to the feature level, and the number increases from 64 to 256 depending on the feature dimension. The edge constraint coefficient is set to 0.3, and the gradient information of the feature cost map is calculated using the gradient operator to strengthen the feature weights in edge regions. The neighborhood feature association matrix is constructed based on the feature similarity within a 5×5 window. The similarity calculation adopts the cosine similarity method. The association matrix and the feature cost map are fused through element-wise multiplication. Finally, the rotation-invariant feature map is obtained by weighted summation. This process fully integrates multi-directional feature information, effectively offsets the feature distortion caused by building rotation, and improves the stability and robustness of feature representation.
[0026] Preferably, the segmentation expression of the edge-constrained multi-task segmentation model is: ,in, This is the preliminary segmentation result. For activation function, For depthwise convolution operations, This is the rotation-invariant feature map after aggregation. For edge feature weights, For edge gradient features, For multi-task segmentation weight matrix, The boundary penalty coefficient, For boundary feature mapping function, These are the a priori features of the building boundary.
[0027] Specifically, the segmentation mechanism of the edge-constrained multi-task segmentation model achieves accurate preliminary segmentation of building regions through deep convolution, edge constraints, and multi-task collaboration. In implementation, the Sigmoid function is used as the activation function, mapping the deep convolution output to the 0-1 range for easy pixel classification. The deep convolution operation employs an encoder-decoder architecture. The encoder extracts high-level features progressively through five convolutional blocks. Each convolutional block includes two convolutional layers and one batch normalization layer. The kernel size is 3×3, and the number of kernels gradually increases from 64 to 512. The batch normalization layer parameters are set to the default momentum of 0.9 and epsilon value of 1e-5. The edge feature weight is set to 0.8, and edge gradient features are extracted using the Sobel operator with a gradient threshold of 0.3. The selected edge features are used to construct an edge feature vector set. The multi-task segmentation weight matrix allocates weights according to the importance of each task: the building region segmentation task has a weight of 0.4, the boundary segmentation task has a weight of 0.3, and the background suppression and feature enhancement tasks each have a weight of 0.15. The boundary penalty coefficient is set to 0.5, and the boundary feature mapping function adopts a non-linear transformation to enhance the feature difference between the boundary region and the interior region. The prior features of the building boundary are derived from the sample data statistics, including the gradient distribution of boundary pixels, spatial spacing, and other information. Through the synergistic effect of the above parameters, the preliminary segmentation results output by the model can accurately distinguish between buildings and background regions while maintaining the clarity of the boundary contours.
[0028] Preferably, the adjustment expression of the scale consistency adaptive adjustment algorithm is: ,in, This is the scaled feature map. The number of scale levels. For the first A weight matrix of several scales, This is the preliminary segmentation result. For the first Adaptive coefficients at each scale For the first Feature mapping function at each scale The consistency constraint coefficient is... This is a scale consistency check function. This is the set of scale consistency parameters.
[0029] Specifically, the scale consistency adaptive adjustment algorithm optimizes the feature representation of buildings at different scales through multi-scale feature matching, verification, and compensation. In implementation, the number of scale levels is set to three, corresponding to small, medium, and large building scales. The weight matrix for each scale is set according to the importance of scale features. The weight matrix elements for small and large scales range from 0.6 to 0.9, while the weight matrix elements for medium scales range from 0.7 to 1.0. The adaptive coefficient is dynamically adjusted according to scale differences: 1.2 for small scales, 1.0 for medium scales, and 0.9 for large scales, ensuring balanced representation of features at different scales. The feature mapping function adopts a differentiated design for different scales: the small-scale feature mapping function strengthens the extraction of detailed features, while the large-scale feature mapping function focuses on capturing overall structural features. The consistency constraint coefficient is set to 0.4. The scale consistency verification function judges consistency by calculating the similarity between features at different scales. The similarity threshold is set to 0.75; features below this threshold are dynamically adjusted. The scale consistency parameter set includes parameters such as scale difference threshold, feature similarity threshold, and adjustment step size, with the adjustment step size set to 0.05. Iterative adjustments are made to ensure consistency among features at different scales. This algorithm effectively addresses the problems of neglecting small-scale building features and redundant large-scale building features through collaborative optimization of multi-scale features, improving the consistency and accuracy of building extraction at different scales.
[0030] Preferably, the feature fusion expression of the multi-source remote sensing feature fusion analysis platform is: ,in, To merge feature spaces, For the number of multi-source auxiliary features, For the first Class principal feature weights, This is the scaled feature map. For the first Class auxiliary feature weights, For the first Multi-source assisted remote sensing features This is a feature correlation measurement function.
[0031] Specifically, the feature fusion of the multi-source remote sensing feature fusion analysis platform constructs a comprehensive fused feature space through hierarchical weighted fusion of multi-source features. During implementation, the number of multi-source auxiliary features is set to three categories: digital elevation model features, hyperspectral auxiliary features, and radar image features. The weights of the main feature and auxiliary features are calculated based on their contribution. The weight of the main feature is set to 0.6, and the weights of the three auxiliary features are 0.15, 0.15, and 0.1, respectively, ensuring the core role of the main remote sensing feature while fully utilizing the complementary information of the auxiliary features. The feature correlation measurement function uses the Pearson correlation coefficient. A sliding window is used to traverse the feature space, with a window size of 3×3. The correlation coefficient between the main feature and each auxiliary feature within the window is calculated. Feature pairs with correlation coefficients higher than 0.7 will have their weights adjusted to reduce the impact of redundant features. The scale-adjusted feature map and the multi-source auxiliary features are first dimensionally aligned, mapping all features to a unified 256-dimensional feature space. Then, preliminary fusion is achieved through weighted summation. During the fusion process, features of the same type are first aggregated internally, and then cross-category features are fused. Each fusion level has a feature filtering step to remove invalid features with a variance of less than 0.01. The final fusion feature space includes 64 core dimensions, which comprehensively integrates multi-dimensional information such as spectrum, space, terrain, and radar, providing rich feature support for subsequent accurate extraction.
[0032] Preferably, the boundary optimization expression for the building extraction results is: ,in, For the final building extraction results, For the feature clustering matrix, For boundary optimization coefficients, It is a second-order gradient operator. To merge feature spaces, Optimize the weight matrix for the boundary.
[0033] Specifically, the boundary optimization mechanism for building extraction results improves the accuracy and completeness of the final extraction boundary through feature clustering and gradient optimization. In implementation, the feature clustering matrix is constructed based on K-means clustering results, with a cluster size of 2, corresponding to the building region and background region respectively. Clustering matrix elements take values of 0 or 1, indicating the pixel's category. The boundary optimization coefficient is set to 0.2. The second derivative of the fused feature space is calculated using a second-order gradient operator to enhance the feature differences in the boundary region. The second-order derivative threshold is set to 0.1; pixels exceeding this threshold are considered boundary candidate pixels. The boundary optimization weight matrix is set according to the gradient strength of the boundary pixels. Pixels with a gradient strength higher than 0.5 correspond to a weight coefficient of 0.9, pixels with a gradient strength between 0.3 and 0.5 correspond to a weight coefficient of 0.6, and pixels with a gradient strength lower than 0.3 correspond to a weight coefficient of 0.3. During implementation, the fused feature space and the feature clustering matrix are first multiplied to initially lock the building region. Then, boundary features are extracted using the second-order gradient operator and fused with the boundary optimization weight matrix, and the building region boundary is adjusted pixel-by-pixel. This process enhances boundary recognition through gradient optimization, corrects problems such as blurred and jagged boundaries, and preserves the internal integrity of the building area. The final output boundary position error is controlled within 2 pixels, accurately restoring the actual outline of the building.
[0034] Preferred, such as Figure 2 As shown, step S2 includes steps S21 to S23: S21, selecting calibrated feature variables from the original feature set of high-resolution remote sensing images, grouping them according to feature dimensions and building association characteristics to form multiple feature subsets; S22, inputting different feature subsets into a rotation-invariant cost map aggregation network, performing multi-directional rotation transformation on each feature subset through the network's hidden layer to generate feature cost components with corresponding angles; S23, calculating the association weights of different cost components based on neighborhood feature association, and aggregating all cost components through a dynamic weighted summation method to form a complete multi-directional rotation feature cost matrix.
[0035] Specifically, step S2 is implemented in stages, achieving accurate extraction of rotation-invariant features through feature grouping, multi-directional rotation mapping, and dynamic aggregation. In step S21, from the original feature set constructed in step S1, based on feature type (spectral, spatial structure, texture) and feature correlation ranking, 18 to 25 core feature variables are divided into 3 to 4 feature subsets. Each subset includes 5 to 8 feature variables, ensuring strong correlation and complementarity among features within each subset, laying the foundation for subsequent targeted processing. In step S22, each feature subset is sequentially input into the rotation-invariant cost map aggregation network. The network rotates each feature subset angle-by-angle according to 8 preset uniformly distributed rotation angles (0° to 315°). Through the network's built-in convolutional and pooling layers, feature response information at each angle is extracted, generating 8 feature cost components with the same dimension as the input feature subset, comprehensively capturing building features from different directions. When implementing step S23, a 5×5 neighborhood feature association calculation window is set, and the cosine similarity method is used to calculate the similarity between each cost component and other components. Based on the similarity results, an association weight of 0.1 to 0.9 is assigned. The higher the similarity, the larger the weight coefficient. Then, through a dynamic weighted summation algorithm, the feature cost components in the eight directions are aggregated to form a complete feature cost matrix that includes multi-directional rotation information. This matrix effectively integrates the advantages of features from different angles and improves the adaptability of features to changes in building rotation.
[0036] Preferred, such as Figure 3 As shown, step S3 includes steps S31 to S34: S31, extracting edge gradient information from the aggregated feature cost map, constructing an edge feature vector set, and clarifying the boundary difference features between the building area and the background area; S32, inputting the feature cost map and the edge feature vector set into the edge-constrained multi-task segmentation model, starting the model's multi-task parallel computing process, and simultaneously segmenting the building area and the boundary area respectively; S33, integrating the edge feature vector set into the segmentation calculation process through the model's built-in constraint mechanism, enhancing the feature recognition of the boundary area; S34, integrating the multi-task segmentation output results to form a preliminary building area segmentation map, retaining the detailed features of the segmentation boundary.
[0037] Specifically, step S3 is executed in steps, achieving preliminary accurate segmentation of the building region through edge feature extraction, multi-task parallel segmentation, and constraint integration. In step S31, the Sobel operator is used to calculate the edge gradient of the feature cost matrix output from step S2. A gradient threshold of 0.3 is set, and pixels with gradient values higher than this threshold are selected as edge candidates. The gradient direction and intensity of these candidate points are quantized to construct an edge feature vector set including edge position, intensity, and direction information, clarifying the boundary difference between the building and the background. During step S32, the feature cost matrix and the edge feature vector set are simultaneously input into the edge-constrained multi-task segmentation model. The model starts four parallel task branches to calculate building region segmentation, boundary segmentation, background suppression, and feature enhancement, respectively. The batch size is set to 16, and the number of iterations is 200 rounds to ensure efficient and coordinated progress of each task. In step S33, the model has a built-in edge constraint mechanism, setting the weight coefficient of the edge feature vector set to 0.8. Through weighted calculation of the loss function, the edge features play a dominant role in the segmentation process, strengthening the feature response of the boundary region and reducing background noise interference. In step S34, a weighted fusion algorithm is used to integrate the output results of the four task branches. The weight of the building region segmentation result is set to 0.4, the weight of the boundary segmentation result is set to 0.3, and the weights of the background suppression and feature enhancement results are each set to 0.15. Finally, a preliminary segmentation map is generated to ensure that the building region outline is clear and the background is completely separated.
[0038] Preferred, such as Figure 4 As shown, S4 includes steps S41 to S43: S41, analyze the differences in feature representation of buildings at different scales in the preliminary segmentation results, and determine multiple calibration scale levels and corresponding feature parameters; S42, start the scale consistency adaptive adjustment algorithm to extract and quantize the features of buildings at each scale level separately; S43, based on the scale consistency criterion, perform mutual verification and dynamic adjustment on the feature quantization results of different scales, correct the feature deviation caused by scale mismatch, and output a building feature map of a unified scale.
[0039] Specifically, step S4 is implemented in stages, optimizing the feature representation of buildings at multiple scales through scale hierarchy division, feature extraction quantization, and consistency adjustment. In step S41, the pixel area of the preliminary segmentation results output from step S3 is statistically analyzed. An area threshold is set to divide buildings into three levels: small scale (less than 100 pixels), medium scale (100 to 500 pixels), and large scale (greater than 500 pixels). At the same time, the feature distribution pattern of buildings at each level is analyzed to determine the key feature parameters corresponding to each level, including shape complexity, texture density, and spatial distribution density, providing a basis for differentiated processing. In step S42, the scale consistency adaptive adjustment algorithm is activated. Dedicated feature extraction windows are set for different scale levels (3×3 for small scale, 7×7 for medium scale, and 11×11 for large scale). Through the built-in feature extraction module of the algorithm, 12 to 15 feature parameters are extracted for each level. Each parameter is quantized, and the value range is uniformly mapped to the [0,1] interval to ensure the comparability of features. In step S43, the scale consistency verification threshold is set to 0.75. The cosine similarity method is used to calculate the similarity between features at different scale levels. For features with similarity below the threshold, adjustment coefficients are set according to scale differences (0.9 for small scale and medium scale, 0.85 for medium scale and large scale) for dynamic adjustment. At the same time, an enhancement coefficient of 1.2 is applied to small scale features, and a 5×5 window is used to smooth large scale features. Through multiple rounds of verification and adjustment, a building feature map of a unified scale is output to achieve a balanced expression of features at different scales.
[0040] Preferred, such as Figure 5 As shown, step S5 includes steps S51 to S53: S51, importing multi-source auxiliary feature data corresponding to the main remote sensing image through a multi-source remote sensing feature fusion analysis platform, including terrain features, spectral auxiliary features, and spatial correlation features; S52, performing dimensional alignment processing on the scale-adjusted feature map and multi-source auxiliary feature data to ensure the fusionability of different types of features in the same feature space; S53, adopting a hierarchical fusion strategy, first performing internal aggregation of similar features, and then performing cross-level fusion of aggregated features of different categories, gradually constructing a fusion feature space including multi-dimensional information, while preserving the uniqueness and integrity of building features.
[0041] Specifically, step S5 involves multi-source feature import, dimension alignment, and hierarchical fusion to construct a comprehensive fused feature space. In step S51, three types of multi-source auxiliary remote sensing features are imported through a multi-source remote sensing feature fusion analysis platform: digital elevation model features (altitude, slope, aspect), hyperspectral auxiliary features (20 spectral band reflectance), and radar image features (backscattering coefficient, polarization ratio, etc.). The platform verifies the imported auxiliary features, removing outliers and missing values to ensure their integrity and reliability. In step S52, the scale-adjusted feature map output from step S4 is dimensionally aligned with the verified multi-source auxiliary features. Feature standardization is used to map all feature values to the [0,1] interval, and the feature dimensions are adjusted to match the main feature map and each auxiliary feature to 256 dimensions, eliminating fusion obstacles caused by dimensional differences. In step S53, a hierarchical fusion strategy is adopted. In the first layer, the same-type feature aggregation stage, aggregation weights of 0.4, 0.15, 0.3, and 0.15 are applied to the main remote sensing features, digital elevation model features, hyperspectral auxiliary features, and radar image features, respectively, and the internal aggregation is completed by weighted averaging. In the second layer, the cross-level fusion stage, the four types of aggregated features are input into a fusion network including three fully connected layers (256, 128, and 64 neurons), and the feature correlation threshold is set to 0.6. Redundant features with correlation higher than the threshold are eliminated, and deep feature fusion is achieved through nonlinear transformation. Finally, a fusion feature space including 64 core dimensions is constructed, which comprehensively integrates multi-source information and provides strong support for subsequent accurate extraction.
[0042] The rotation-invariant cost graph aggregation network is a technical model used for feature preprocessing and aggregation in this invention. It generates rotation-invariant feature representations through multi-directional rotation transformation and dynamic weighted fusion. The implementation of this network consists of three steps: First, the original feature set is divided into 3 to 4 feature subsets based on type and correlation, with each subset containing 5 to 8 highly correlated feature variables. Then, the network performs 8 uniform angle rotation transformations from 0° to 315° on each subset, processing them through 4 built-in convolutional layers and 2 pooling layers (with kernel sizes of 3×3 or 5×5 and a number of 64 to 256) to extract feature responses at each angle, generating 8 feature cost components with consistent dimensions. Finally, a 5×5 neighborhood calculation window is set, and cosine similarity is used to calculate the correlation of each component, assigning dynamic weights from 0.1 to 0.9. These weighted sums are then aggregated into a complete feature cost matrix, comprehensively capturing building features in different rotation directions, offsetting feature distortion caused by rotation, improving the stability of feature representation, breaking through the dependence of traditional feature extraction on building pose, solving the feature adaptation problem caused by diverse building orientations in complex urban environments, providing a more robust feature foundation for subsequent segmentation, and significantly improving the adaptability of the extraction method to rotational changes.
[0043] The edge-constrained multi-task segmentation model is a model for achieving preliminary and accurate segmentation of building regions. It accurately distinguishes buildings from background regions and optimizes boundary representation through edge feature enhancement and multi-task collaborative computation. Its implementation process includes four key steps: First, the Sobel operator is used to calculate the edge gradient of the feature cost matrix, setting a gradient threshold of 0.3 to filter candidate edge points and constructing an edge feature vector set including location, intensity, and orientation information. Next, the feature cost matrix and edge vector set are input into the encoder-decoder architecture model, initiating four parallel task branches: building region segmentation, boundary segmentation, background suppression, and feature enhancement, with a batch size of 16 and 200 iterations. A built-in constraint mechanism assigns a weight of 0.8 to the edge features, incorporating a weighted loss function to enhance the feature response of the boundary region. Finally, the output results of the four tasks are fused with weights of 0.4, 0.3, 0.15, and 0.15 to generate a preliminary segmentation map. This model accurately separates buildings from the background while enhancing the distinctiveness of segmentation boundaries and reducing boundary ambiguity. It addresses the shortcomings of traditional segmentation models, such as coarse boundary extraction and incomplete background separation. Through the combination of multi-task collaboration and edge constraints, it provides a clear segmentation foundation for subsequent optimization steps, thereby improving the boundary accuracy of the overall extraction results.
[0044] The scale consistency adaptive adjustment algorithm is an algorithm for optimizing the feature representation of buildings at multiple scales. It achieves balanced representation of building features at different scales through scale hierarchy partitioning, differentiated feature processing, and consistency verification. Its implementation process mainly consists of three steps: First, the pixel area of the preliminary segmentation results is statistically analyzed, dividing the data into small, medium, and large scale levels based on thresholds of less than 100 pixels, 100 to 500 pixels, and greater than 500 pixels. The feature distribution patterns of each level are analyzed, and key parameters are determined. Then, specific feature extraction windows (3×3, 7×7, 11×11) are set for different scales, extracting 12 to 15 feature parameters and standardizing them to the [0,1] interval. Finally, a consistency verification threshold of 0.75 is set, and cosine similarity is used to calculate cross-scale feature similarity. For features below the threshold... The features are adjusted by a factor of 0.85 to 0.9 according to scale differences, an enhancement factor of 1.2 is applied to small-scale features, and a 5×5 window smoothing process is applied to large-scale features to output a uniform scale feature map. This balances the feature representation of buildings at different scales, corrects feature bias caused by scale mismatch, solves the problem of uneven extraction accuracy of buildings at different scales in traditional methods, effectively avoids the defects of missed detection of small-scale buildings and feature redundancy of large-scale buildings, improves the adaptability of the extraction method to buildings of various scales in complex urban environments, and ensures the consistency and comprehensiveness of the overall extraction results.
[0045] The multi-source remote sensing feature fusion analysis platform is a platform that integrates multi-dimensional information and constructs a comprehensive feature space. It fully explores the complementary value of different types of features through multi-source feature import, standardized processing and hierarchical fusion. The implementation process includes three steps: First, import three types of auxiliary features: digital elevation model (elevation, slope, aspect), hyperspectral imagery (20 band reflectivity), and radar imagery (backscattering coefficient, polarization ratio). Data verification is performed through the platform to remove outliers and missing values. Then, a feature standardization method is used to uniformly map the scale-adjusted main feature map and auxiliary features to the [0,1] interval, adjusting the dimension to 256 dimensions to achieve dimensional alignment. Finally, a hierarchical fusion strategy is implemented. The first layer aggregates the main feature and the three types of auxiliary features according to weights of 0.4, 0.15, 0.3, and 0.15 respectively. The second layer inputs the four aggregated features into a fusion network containing three fully connected layers (256, 128, and 64 neurons), sets a correlation threshold of 0.6 to remove redundant features, and generates a 64-dimensional fusion feature space through nonlinear transformation. This integrates complementary information from multiple sources, comprehensively depicting the essential features of the building and enhancing the richness of the feature space. This platform breaks through the bottleneck of limited single remote sensing feature information, makes full use of the advantages of multi-source data, solves the problem of building features being confused with the background in complex urban environments, provides comprehensive and reliable feature support for the final accurate extraction, and greatly improves the anti-interference ability and accuracy of the extraction method.
[0046] This paper presents an automatic building extraction method for high-resolution remote sensing images in complex urban environments. Addressing the shortcomings of existing technologies in feature processing and aggregation, this method utilizes a rotation-invariant cost graph aggregation network to perform multi-directional rotation mapping and dynamic aggregation of original features, comprehensively capturing rotation-invariant building features. Furthermore, a scale-consistency adaptive adjustment algorithm completes multi-scale feature matching and optimization, ensuring comprehensive and accurate representation of building features across different directions and scales. This completely resolves the issues of poor feature adaptability and insufficient discriminative power in traditional methods. To address the deficiencies in segmentation models and feature fusion mechanisms, an edge-constrained multi-task segmentation model enhances segmentation boundary identification. A hierarchical fusion strategy is implemented using a multi-source remote sensing feature fusion analysis platform to fully leverage the complementary value of primary remote sensing features and multi-source auxiliary features, effectively refining segmentation boundaries and optimizing feature spatial structure, significantly reducing missed and false detections in complex urban environments.
[0047] This method boasts strong technical synergy, deeply integrating core technologies such as rotation-invariant cost graph aggregation networks and edge-constrained multi-task segmentation models with a multi-source feature fusion platform. This forms a closed-loop process from feature extraction, aggregation, and segmentation to adjustment and fusion, with each stage of the technology precisely connected and working together efficiently. It exhibits outstanding environmental adaptability, effectively addressing challenges such as dense building distribution, diverse morphologies, and spectral obfuscation in complex urban environments through multi-directional feature capture, multi-scale adaptive adjustment, and edge constraint enhancement. The extraction results are accurate and reliable, achieving comprehensive mining and precise representation of building features through a step-by-step, refined technical process. This ensures both extraction efficiency and accurate reconstruction of the actual morphology and spatial distribution of buildings, providing high-quality spatial information support for related applications.
[0048] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for automatic building extraction from high-resolution remote sensing images in complex urban environments, characterized in that, Includes the following steps: S1. Acquire the spectral features, spatial structure features, and texture features of high-resolution remote sensing images, and select feature variables with building correlation higher than a preset threshold to construct the original feature set; S2. Perform feature mapping on the original feature set through a rotation-invariant cost map aggregation network to generate a multi-directional rotation feature cost matrix, and dynamically aggregate the cost map by combining neighborhood feature correlation; S3. Use an edge-constrained multi-task segmentation model to perform preliminary segmentation of building regions on the aggregated feature cost map, and enhance the segmentation boundary recognition through edge gradient features; S4. Use a scale consistency adaptive adjustment algorithm to perform multi-scale feature matching and adjustment on the preliminary segmentation results, and optimize the building feature representation at different scales; S5, Based on the multi-source remote sensing feature fusion analysis platform, the adjusted features are hierarchically fused with multi-source auxiliary remote sensing features to construct a fused feature space; S6, Based on the unique distribution pattern of building features in the fused feature space, the building extraction results are output through feature clustering and boundary optimization processing; The characteristic aggregation expression of the rotation-invariant cost graph aggregation network is: ,in, This is the rotation-invariant feature map after aggregation. The number of rotation angles. For the first Weighting coefficients for each rotation angle, For angle The rotation transformation matrix, For the original feature set, For the first Convolution kernels corresponding to each rotation angle, The edge constraint coefficient, For gradient operators, For feature cost map, For element-wise multiplication, The neighborhood feature correlation matrix; The adjustment expression for the scale consistency adaptive adjustment algorithm is: ,in, This is the scaled feature map. The number of scale levels. For the first Weight matrices of various scales This is the preliminary segmentation result. For the first Adaptive coefficients at each scale For the first Feature mapping function at each scale The consistency constraint coefficient is... This is a scale consistency check function. This is the set of scale consistency parameters.
2. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, The segmentation expression of the edge-constrained multi-task segmentation model is as follows: ,in, This is the preliminary segmentation result. For activation function, For depthwise convolution operations, This is the rotation-invariant feature map after aggregation. For edge feature weights, For edge gradient features, For multi-task segmentation weight matrix, The boundary penalty coefficient, For boundary feature mapping function, These are the a priori features of the building boundary.
3. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, The feature fusion expression of the multi-source remote sensing feature fusion analysis platform is: ,in, To merge feature spaces, For the number of multi-source auxiliary features, For the first Class principal feature weights, This is the scaled feature map. For the first Class auxiliary feature weights, For the first Multi-source assisted remote sensing features This is a feature correlation measurement function.
4. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, The boundary optimization expression for the building extraction results is: ,in, For the final building extraction results, For the feature clustering matrix, For boundary optimization coefficients, It is a second-order gradient operator. To merge feature spaces, Optimize the weight matrix for the boundary.
5. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, S2 includes steps S21 to S23: S21, selecting calibrated feature variables from the original feature set of high-resolution remote sensing images, grouping them according to feature dimensions and building association characteristics to form multiple feature subsets; S22, inputting different feature subsets into a rotation-invariant cost map aggregation network, performing multi-directional rotation transformation on each feature subset through the network's hidden layer to generate feature cost components at corresponding angles; S23, calculating the association weights of different cost components based on neighborhood feature association, and aggregating all cost components through a dynamic weighted summation method to form a complete multi-directional rotation feature cost matrix.
6. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, S3 includes steps S31 to S34: S31, extracting edge gradient information from the aggregated feature cost map, constructing an edge feature vector set, and clarifying the boundary difference features between the building area and the background area; S32, inputting the feature cost map and the edge feature vector set into the edge-constrained multi-task segmentation model, starting the model's multi-task parallel computing process, and simultaneously segmenting the building area and the boundary area respectively; S33, integrating the edge feature vector set into the segmentation calculation process through the model's built-in constraint mechanism, enhancing the feature recognition of the boundary area; S34, integrating the multi-task segmentation output results to form a preliminary building area segmentation map, retaining the detailed features of the segmentation boundaries.
7. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, S4 includes sub-steps S41 to S43: S41, analyze the differences in the feature expression of buildings of different scales in the preliminary segmentation results, and determine multiple calibration scale levels and corresponding feature parameters; S42, initiate the scale consistency adaptive adjustment algorithm to extract and quantize the building features at each scale level separately; S43, based on the scale consistency criterion, performs mutual verification and dynamic adjustment of feature quantification results at different scales, corrects feature deviations caused by scale mismatch, and outputs building feature maps at a unified scale.
8. The method for automatic building extraction from high-resolution remote sensing images in complex urban environments according to claim 1, characterized in that, S5 includes steps S51 to S53: S51, importing multi-source auxiliary feature data corresponding to the main remote sensing image through a multi-source remote sensing feature fusion analysis platform, including topographic features, spectral auxiliary features, and spatial correlation features; S52, performing dimensional alignment processing on the scale-adjusted feature map and the multi-source auxiliary feature data to ensure the fusionability of different types of features in the same feature space; S53, adopting a hierarchical fusion strategy, first aggregating similar features internally, and then fusing aggregated features of different categories across levels, gradually constructing a fusion feature space including multi-dimensional information. Preserve the uniqueness and integrity of the building's features.
Citation Information
Patent Citations
High-resolution remote sensing image building extraction method and system
CN112633070A
Remote sensing image building semantic segmentation edge optimization method based on multitask CNN+GCN
CN113449640A