Multi-scale building extraction method and system

By processing remote sensing images using depthwise separable convolution and attention mechanisms, multi-scale building features are extracted, solving the problem of inaccurate multi-scale building extraction in existing technologies and achieving high accuracy in building distribution maps.

CN121280909APending Publication Date: 2026-01-06JIANGXI COLLEGE OF APPLIED TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511531082.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing technologies cannot be compatible with semantic information, matching coefficients, and second-level feature maps in building extraction, resulting in inaccurate multi-scale feature extraction methods and making it impossible to achieve accurate extraction of buildings at multiple scales.

Method used

By using depth-separable convolution processing on remote sensing images of the target building area, local features are extracted, and a second feature map is determined using attention processing channels and spatial attention channels. By combining semantic information, matching coefficients, and the second feature map, the extraction method of multi-scale features is determined, the scale relationship and location of building features are marked, the building group features are optimized, and finally, a building distribution map is generated.

Benefits of technology

It achieves accurate extraction of buildings at multiple scales, improves the accuracy of building distribution maps, takes into account semantic information, matching coefficients, and the overall consideration of the second feature map, and optimizes the transformation process of building cluster features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280909A_ABST
    Figure CN121280909A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale building extraction method and system, and relates to the technical field of building extraction, corresponding matching coefficients are determined based on features of an application channel and features of a space channel, and semantic information is determined according to detection of a second feature map. The multi-scale feature extraction mode is determined based on the semantic information, the matching coefficient and the second feature map, the accuracy of the multi-scale feature extraction mode is improved, and accurate extraction of the multi-scale building is achieved. Therefore, the optimized building group features are determined according to the positions and scale relationships of the plurality of building features and the form of the second feature map; the corresponding extraction loss part is determined according to the recognition of the feature extraction event of the second feature map, and the final building distribution map is determined based on the extraction loss part, the building distribution condition of the optimized building group features and the actual image of the target building area, so that the accuracy of the final building distribution map is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of building extraction, and more particularly to a method and system for extracting multi-scale buildings. Background Technology

[0002] With the development of technology, remote sensing imagery is increasingly being used in urban planning, disaster monitoring, and resource management. As an important component of urban space, the automated extraction and analysis of buildings is crucial for achieving precise urban management and smart city construction. Currently, building extraction methods based on remote sensing imagery are mainly divided into traditional methods and deep learning methods. Traditional methods rely on manually designed features, such as spectral information, shape, texture, and context, and combine them with classifiers such as support vector machines (SVM) and random forests for extraction. However, they do not take into account semantic information, matching coefficients, and the overall consideration of second-level feature maps, which affects the extraction of multi-scale features and makes it impossible to achieve accurate extraction of buildings at multiple scales. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and system for extracting multi-scale buildings.

[0004] This invention provides a method for extracting multi-scale buildings, including: Local features are extracted by depth-separable convolution processing of remote sensing images of the target building area to output the corresponding first-layer feature map; The second feature map is determined based on the first feature map, the attention processing channel, and the spatial attention channel, and the features of the application channel and the spatial channel are labeled. The matching coefficients are determined based on the features of the application channel and the spatial channel. Semantic information is determined based on the detection of the second feature map. Based on the semantic information, the matching coefficients, and the second feature map, the extraction method of multi-scale features is determined. Multiple building features are determined based on the extraction methods of the second feature map and multi-scale features, and the scale relationship of the multiple building features is marked. The optimized building cluster features are determined based on the location, scale relationship, and shape of the second feature map of the multiple building features. The corresponding extraction loss is determined based on the identification of feature extraction events in the second feature map. The final building distribution map is determined based on the extraction loss, the optimized building cluster features, and the actual image of the target building area.

[0005] This invention provides a multi-scale building extraction system, which is applied to the above-described multi-scale building extraction method. The multi-scale building extraction system includes: The first feature map module is used to extract local features based on the depth-separable convolution processing of the remote sensing image of the target building area, so as to output the corresponding first feature map. The feature module is used to determine the second feature map based on the first feature map, the attention processing channel, and the spatial attention channel, and to label the features of the application channel and the features of the spatial channel. The extraction method module is used to determine the corresponding matching coefficients based on the features of the application channel and the features of the spatial channel, determine the semantic information based on the detection of the second feature map, and determine the extraction method of multi-scale features based on the semantic information, the matching coefficients and the second feature map. The building cluster feature module is used to determine multiple building features based on the extraction methods of the second feature map and multi-scale features, and to mark the scale relationship of multiple building features. Based on the location, scale relationship, and shape of the second feature map of multiple building features, the optimized building cluster features are determined. The building distribution map module is used to determine the corresponding extraction loss part based on the recognition of feature extraction events in the second feature map, and to determine the final building distribution map based on the extraction loss part, the optimized building cluster features, and the actual image of the target building area.

[0006] Compared with the prior art, the beneficial effects of the present invention are: In this embodiment of the invention, the method extracts local features based on depth-separable convolution processing of remote sensing images of the target building area to output a corresponding first feature map. A second feature map is determined based on the first feature map, the attention processing channel, and the spatial attention channel, and the features of the application channel and the spatial channel are labeled. Matching coefficients are determined based on the features of the application channel and the spatial channel. Semantic information is determined based on the detection of the second feature map. Based on this semantic information, the matching coefficients, and the second feature map, a multi-scale feature extraction method is determined. The introduction of the second feature map incorporates the overall considerations of the semantic information, the matching coefficients, and the second feature map, improving the accuracy of the multi-scale feature extraction method and achieving accurate extraction of multi-scale buildings.

[0007] Therefore, multiple building features are determined based on the extraction methods of the second-level feature map and multi-scale features, and the scale relationship of the multiple building features is marked. The optimized building cluster features are determined based on the position, scale relationship, and morphology of the second-level feature map of the multiple building features. The corresponding extraction loss part is determined based on the recognition of feature extraction events in the second-level feature map. The final building distribution map is determined based on the extraction loss part, the building distribution of the optimized building cluster features, and the actual image of the target building area. This achieves further transformation of the optimized building cluster features and introduces the extraction loss part, realizing the overall consideration of the extraction loss part, the building distribution of the optimized building cluster features, and the actual image of the target building area, thus improving the accuracy of the final building distribution map. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the method for extracting multi-scale buildings in an embodiment of the present invention; Figure 2 This is a flowchart illustrating step S11 in the multi-scale building extraction method of this invention. Figure 3 This is a flowchart illustrating step S12 in the multi-scale building extraction method of this invention. Figure 4 This is a flowchart illustrating step S13 in the multi-scale building extraction method of this invention. Figure 5 This is a flowchart illustrating step S14 of the multi-scale building extraction method in this embodiment of the invention. Figure 6 This is a flowchart illustrating step S15 in the multi-scale building extraction method of this invention. Figure 7 This is a schematic diagram of the structural composition of the multi-scale building extraction system in an embodiment of the present invention; Figure 8 This is a schematic diagram of building cluster features in the multi-scale building extraction method of this invention. Detailed Implementation

[0009] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0010] Please see Figures 1 to 8 A method for extracting multi-scale buildings is proposed and applied to multi-scale building extraction scenarios. The method includes: Step S11: Extract local features based on the depth-separable convolution processing of the remote sensing image of the target building area to output the corresponding first-layer feature map; Step S12: Determine the second feature map based on the first feature map, the attention processing channel, and the spatial attention channel, and label the features of the application channel and the features of the spatial channel; Step S13: Determine the corresponding matching coefficients based on the features of the application channel and the spatial channel, determine the semantic information based on the detection of the second feature map, and determine the extraction method of multi-scale features based on the semantic information, the matching coefficients and the second feature map; Step S14: Determine multiple building features based on the extraction method of the second feature map and multi-scale features, and mark the scale relationship of multiple building features. Determine the optimized building cluster features based on the location, scale relationship, and shape of the second feature map of multiple building features. Step S15: Determine the corresponding extraction loss part based on the recognition of the feature extraction event of the second feature map, and determine the final building distribution map based on the extraction loss part, the building distribution of the optimized building group features and the actual image of the target building area; refer to Figure 2 In step S11, the specific steps are as follows: S111: Acquire the location of the target building area, acquire remote sensing images of the target building area based on the location detection of the target building area, and trigger depth-separable convolution processing of the remote sensing images of the target building area based on the remote sensing images of the target building area and the corresponding convolution model. S112: During the depth-separable convolutional processing of the remote sensing image of the target building area, multiple convolutional layers are acquired and unfolded sequentially, and the corresponding binary segmentation map is determined based on the processing of the multiple convolutional layers. S113: Determine local features based on the extraction of the binary segmentation map; determine the first-level feature map based on the synthesis of the position and shape of each local feature, and output the corresponding first-level feature map; In the embodiments of this application, the location of the target building area is collected, and the specific geographical range for building extraction is determined. This is usually a bounding box, polygon, or a specific geographical area name defined by geographical coordinates (longitude and latitude). At the same time, high-resolution remote sensing image data covering the defined geographical area is acquired. A suitable remote sensing data source is selected according to the requirements (resolution, time, cost).

[0011] The first step in feature extraction is to process the acquired remote sensing images using a pre-trained deep learning model (based on depthwise separable convolutions). This involves acquiring a pre-trained convolutional neural network model whose core layer is a depthwise separable convolution. The model has been trained on a large-scale remote sensing dataset and possesses preliminary image understanding capabilities, or it has been fine-tuned for specific building extraction tasks. The model file is then loaded into the computing environment (e.g., using frameworks such as PyTorch or TensorFlow). The remote sensing image files are loaded into memory. Preprocessing is required according to the model input requirements. For example: if the image is too large, it needs to be cropped into image tiles suitable for the model input size; the pixel values ​​are scaled to the range used during model training (such as [0,1] or [-1,1]); if the model only accepts specific bands, the corresponding bands need to be selected; the prepared image data is input into the loaded convolutional model; the model's depthwise separable convolutional layers will automatically start processing the input image and perform feature extraction operations; the model starts processing the input remote sensing image and prepares to generate the first feature map (input of S112); at this time, depthwise separable convolution calculations are being performed inside the model.

[0012] Furthermore, depthwise separable convolution typically consists of two parts: depthwise convolution and pointwise convolution. A "convolutional layer" here refers to a depthwise convolutional layer, or a "block" of depthwise + pointwise convolution, or a coarser-grained stage in the model. Capturing multiple "sequentially unfolded" layers means focusing on different stages in the model's processing and capturing features at different levels and with different degrees of abstraction.

[0013] First, it's necessary to understand the specific structure of the convolutional model used. For example, a typical backbone network contains multiple stages, each stage consisting of multiple convolutional blocks, and each block contains depthwise and pointwise convolutional layers. During model training or inference, it's necessary to "insert hooks" after specific, selected convolutional layers or directly access their outputs, which usually requires modifying the model code. Several key layers are selected, such as: shallower layers (closer to the input): capturing low-level features like edges and textures; intermediate layers: capturing more complex shapes and parts; and deeper layers (closer to the output): capturing more abstract, global semantic information. For each selected convolutional layer, its processed output is a "feature map," the size of which usually decreases with increasing network depth (through stride or pooling), while the number of channels increases. For example, suppose the outputs of the 3rd, 7th, and 12th convolutional blocks (each containing depthwise + pointwise) are selected as the "multiple convolutional layers" to be sampled.

[0014] The features extracted by the model at a specific layer are converted into a preliminary, binary building region prediction map. This "binary segmentation map" is not the final high-precision segmentation result, but an intermediate representation reflecting the specific network layer. At this point, for each acquired feature map, a method is needed to convert it into a probability map, where the value of each pixel represents the "confidence" or "probability" that the pixel belongs to a building. Optionally, a very simple fully connected layer or a global average pooling + Softmax layer can be added after the selected convolutional layer to convert the feature map into a single-channel probability map. This classification head only needs to distinguish between "building" and "non-building". Applying the above method to each feature map yields a binary segmentation map (0 and 1) with the same size as the feature map. Since the feature map size is usually smaller than the original image (after downsampling), this binary map is also downsampled. It needs to be upsampled (e.g., by nearest neighbor interpolation or bilinear interpolation) to the same size as the original input image or some intermediate uniform size for subsequent processing.

[0015] Therefore, three binary segmentation maps with different resolutions, reflecting the model's "preliminary judgment" of buildings at different depth levels, were obtained: binary_map_3 (128x128), binary_map_7 (64x64, upsampled to 128x128), and binary_map_12 (32x32, upsampled to 128x128). These maps will be used as input to S113 to extract more refined local features. At this point, morphological operations (such as dilation, erosion, opening, and closing operations) are applied to remove noise, fill small holes, and connect broken building outlines. On the processed binary maps, all interconnected pixel regions (i.e., connected regions) are identified. Each connected region is considered a potential local building instance or part; for each identified connected region, its features are calculated, including: location: the center coordinates (x, y) of the region; shape: the area (number of pixels), bounding box, minimum bounding rectangle, aspect ratio, perimeter, compactness, principal axis direction, etc.; relationship to the original image: on which resolution of the binary image was the region detected (this implies scale information); the pixel values ​​of the region in the original binary image (although the binary image only has 0 and 1, its intensity or which probability image it comes from can be recorded).

[0016] All the local features extracted in the previous step are integrated into a unified and richer feature map. This "first-layer feature map" should contain more information than the original binary segmentation map, providing a more solid foundation for subsequent steps (such as attention mechanisms and semantic understanding). At this point, the attributes of each local feature (location, shape, resolution, etc.) are encoded into a vector or a set of values. For example, center coordinates, area, aspect ratio, resolution index, etc. can be combined. On a predefined, usually high-resolution grid (e.g., consistent with the resolution of the remote sensing image input to S111, or consistent with the resolution of a certain intermediate layer feature map, such as 128x128), the encoded features are placed at the positions of the corresponding local features.

[0017] Since a location corresponds to multiple local features at different resolutions or levels (for example, a building may be detected in parts of binarymap3, binarymap7, and binarymap12), a fusion strategy needs to be designed. Common strategies include: stacking feature vectors along their dimensions to create a long vector containing all relevant features for each grid point; weighted summation based on feature confidence, importance, or similarity to other features; taking the maximum value of all relevant features along a certain dimension for each grid point; encoding spatial relationships (such as distance and direction) between local features into the feature map; and filling grid points that do not detect any local features with zero vectors or special markers.

[0018] refer to Figure 3 In step S12, the specific steps are as follows: S121: Collect the first layer of feature map, determine the first feature region based on the first layer of feature map and the attention processing channel; determine the second feature region based on the first layer of feature map and the spatial attention channel; S122: Determine the corresponding feature map construction method based on the mapping relationship between the first feature region, the second feature region, and the feature map construction method, and determine the second feature map based on the feature map construction method, the first feature region, and the second feature region; S123: Multiple channel features are determined based on the detection of the second feature map, and the features of the application channel and the features of the spatial channel are determined based on the filtering of multiple channel features; In the embodiments of this application, the feature map output from the previous step S113, which already contains multi-scale local feature information, is obtained as the input for the attention mechanism. Here, we assume the first feature map output by S113 is a 4D tensor with shape [BatchSize, Channels, Height, Width]. Simplifying, let's assume BatchSize = 1 (processing only one image), Channels = 5 (representing the extraction of 5 different local features, corresponding to different scales or types of features), and Height = Width = 4 (a very small feature map for demonstration purposes). This tensor can then be represented as a 5x4x4 cube; each channel represents a type of local feature, for example... Channel 0: Represents the morphologically processed location information of the building area center point extracted from coarsescale (binarymap3); Channel 1: Represents the morphologically processed location information of the building area center point extracted from mediumscale (binarymap7); Channel 2: Represents the morphologically processed location information of the building area center point extracted from finescale (binarymap12); Channel 3: Represents the weighted composite information of building morphology (such as aspect ratio) extracted from multiple scales; Channel 4: Represents the weighted composite confidence score information extracted from multiple scales.

[0019] By using the channel attention mechanism, we can determine which channels (channel dimensions) in the first feature map contain more global and important information. Here, the "first feature region" can be understood as the set of channels or features on these channels that the channel attention mechanism considers important. At this time, channel attention usually generates a channel weight vector by calculating the global information of each channel (such as global average pooling or global max pooling), and then enhances or suppresses the channels of the original feature map based on this weight vector.

[0020] Assuming a simple channel attention module is used, it first performs global average pooling (GAP) on the input feature map, compressing the HxW dimension to 1. For the 5x2x2 patch above: GAPonChannel0: (0.1+0.2+0.3+0.4) / 4=0.25; GAPonChannel1: (0.2+0.1+0.4+0.3) / 4=0.25; GAPonChannel2: (0.0+0.0+0.0+0.0) / 4=0.0; GAPonChannel3: (2.5+2.7+2.6+2.4) / 4≈2.55; GAPonChannel4: (1.1+1.0+1.2+1.1) / 4≈1.10; The GAP result is obtained as [0.25, 0.25, 0.0, 2.55, 1.10] (shape [Channels]). The GAP result is then passed through a small fully connected network (usually containing ReLU and Sigmoid layers) to generate channel weights. Assuming that after network processing, the resulting weight vector is [0.1, 0.1, 0.0, 0.8, 0.5] (the weight sum is not equal to 1, or it has been L2 normalized), this weight vector indicates that Channel3 (ShapeRatio) and Channel4 (Confidence) are considered more important, while Channel2 (FinePosition) has an importance of 0.

[0021] The first feature region can be understood as the channel emphasized by high weights; in this example, it mainly refers to Channel3 and Channel4. The model can adjust the original feature map using these weights, for example, by multiplying the weights by the original feature map (Element-wise Multiply) to obtain a weighted feature map. This weighted feature map, or directly referring to the channels (Channel3,4) marked by high weights, is the "first feature region". It represents the information that the model considers more important in the channel dimension and is more relevant to application-related features (such as the shape of buildings and category confidence).

[0022] Using the spatial attention mechanism, we determine which spatial locations (Height x Width dimension) in the first feature map contain more important information. Here, the "second feature region" can be understood as the spatial region or features in these regions that are considered important by the spatial attention mechanism. At this time, spatial attention usually generates a spatial weight matrix by comparing the information of each spatial location with its surrounding locations, and then enhances or suppresses the spatial locations of the original feature map based on this weight matrix.

[0023] Suppose we use a simple spatial attention module, which typically first computes the channel information for each spatial location (e.g., through global max pooling and global average pooling, and then concatenates them), and then generates spatial weights through a small network.

[0024] For the input 5x2x2 patch, calculate the maximum and average channel values ​​at each spatial location; Position (0,0): Max=max(0.1,0.2,0.0,2.5,1.1)=2.5;Avg=(0.1+0.2+0.0+2.5+1.1) / 5=0.76; Position (0,1): Max=max(0.2,0.1,0.0,2.7,1.0)=2.7;Avg=(0.2+0.1+0.0+2.7+1.0) / 5=0.72; Position (1,0): Max=max(0.3,0.4,0.0,2.6,1.2)=2.6;Avg=(0.3+0.4+0.0+2.6+1.2) / 5=0.76; Position (1,1): Max=max(0.4,0.3,0.0,2.4,1.1)=2.4;Avg=(0.4+0.3+0.0+2.4+1.1) / 5=0.72; The channel statistics (Max, Avg) are obtained as follows: [(2.5, 0.76), (2.7, 0.72), (2.6, 0.76), (2.4, 0.72)] (with shape [Height, Width, 2]). The channel statistics are then processed through a small convolutional network (e.g., a 3x3 convolution followed by a sigmoid function) to generate a spatial weight matrix. Assuming that the weight matrix obtained after network processing is: [[0.9, 0.95], [0.92, 0.88]] (with shape [Height, Width]), this weight matrix indicates that all spatial locations are relatively important, but the location (0, 1) is the most important.

[0025] The second feature region can be understood as the spatial location emphasized by high weights; in this example, it is mainly the location (0,1). The model can obtain a weighted feature map by performing Element-wise Multiply on this weight matrix and the original feature map. This weighted feature map, or directly referring to the spatial location marked by high weights (such as a small area centered at (0,1), is the "second feature region". It represents information that the model considers more important in the spatial dimension and is more related to the actual location, boundaries and other spatial structures of the building.

[0026] Furthermore, the corresponding feature map construction method is determined based on the mapping relationship between the first feature region, the second feature region, and the feature map construction method. The second feature map is determined based on the feature map construction method, the first feature region, and the second feature region. This approach takes into account the overall considerations of the feature map construction method, the first feature region, and the second feature region, ensuring the accuracy of the second feature map.

[0027] At this point, there is a first feature region description from S121 (e.g., important channel indices {3,4} and their weights {0.8,0.5}, and their approximate distribution on the feature map) and a second feature region description (e.g., important spatial locations (0,1) and their weights 0.95); there is also a “feature map construction mapping relationship”, which can be regarded as a strategy library.

[0028] Feature map construction methods and mapping relationships: Strategy A: "If the first feature region mainly contains morphological and confidence information (channels 3, 4), and the second feature region is concentrated in a small area, then the 'channel weighting + spatial alignment' method is used for construction"; Strategy B: "If the first feature region contains location information at multiple scales (channels 0, 1, 2), and the second feature region is scattered, then the 'multi-scale fusion' method is used for construction"; Strategy C: "If the first and second feature regions highly overlap, then the 'local enhancement' method is used for construction".

[0029] Based on the information from the first feature region (channels 3 and 4 are important) and the second feature region (position (0, 1) is important), the model queries the mapping relationship and finds that it matches the description of "strategy A" (morphology / confidence channels are important and spatially concentrated); therefore, the model determines that the "feature map construction method" used in this case is "channel weighting + spatial alignment".

[0030] The first layer of feature maps is manipulated according to the determined "construction method," which typically involves operations such as feature map weighting, concatenation, convolution, and attention-weighted summation. Here, the first layer of feature maps is a 5x4x4 cube, constructed using the determined method ("channel weighting + spatial alignment"), with a first feature region description (channels {3,4}, weights {0.8,0.5}) and a second feature region description (position (0,1), weight 0.95). The channels of the first layer of feature maps are weighted. Based on the information from the first feature region, channel 3 is assigned a weight of 0.8, channel 4 a weight of 0.5, and other channels have lower weights (e.g., 0.1). The weighted channel features are calculated as: weighted_channels = 0.8*channel3 + 0.5*channel4 + 0.1*channel0 + 0.1*channel1 + 0.1*channel2, which generates a new feature representation along the channel dimension.

[0031] Using the information from the second feature region, find the spatial location (0,1) and use its high weight of 0.95. A spatial weighting can be applied to the entire weighted feature map, emphasizing information near the location (0,1). For example, a spatial weight matrix can be created, centered at (0,1), decreasing around it, with a maximum value of 0.95. Then, this spatial weight matrix is ​​element-wise multiplied with weighted_channels; spatially aligned_features = spatial_weight_matrix * weighted_channels.

[0032] spatially_aligned_features is the generated second-level feature map, which is a new feature representation that integrates important channel information (morphology and confidence) and spatially aligns important regions (positions (0,1)). Assuming the number of channels remains unchanged, the second-level feature map is still a 5x4x4 tensor, but its content has been reorganized and emphasized.

[0033] Therefore, multiple channel features are determined based on the detection of the second feature map, and the features of the application channel and the spatial channel are determined based on the selection of multiple channel features. This approach takes into account the overall consideration of selecting multiple channel features and ensures the accuracy of the features of the application channel and the spatial channel.

[0034] At this point, the specific information contained in each channel is extracted from the second feature map generated by S122. This is equivalent to "unpacking" the feature map to see what each channel actually represents. The average value of each channel in the feature map is taken in the spatial dimension (Height, Width) to obtain a vector or scalar representing the overall information of that channel. The feature map is input into a small linear layer, and a certain representation or score of each channel is output. Using unsupervised or supervised learning methods, similar channels are grouped, or a category label (e.g., "edge", "texture", "color", "shape", etc.) is directly assigned to each channel.

[0035] The multiple channel features extracted in the previous step are divided into two categories based on their content or importance: one category is related to the specific "application" or "attribute" of the building (such as type, material, integrity), and the other category is related to the "space" or "structure" of the building (such as edges, corners, orientation). This is to prepare for subsequent steps (such as multi-scale feature fusion). At this point, if channel i is designed to extract texture information, it belongs to the application channel; if channel j is designed to extract edge information, it belongs to the spatial channel; or, channels with high activation values ​​and responding to texture patterns belong to the application channel, and channels with high activation values ​​and responding to edge patterns belong to the spatial channel.

[0036] Optionally, the resulting channel feature vectors [avgC1, avgC2, ..., avgCn] are used. Assuming predefined rules are applied to the activation values, based on the model design, we know that: Channel C1 primarily captures the texture information of the building (e.g., roof tile texture); Channel C2 primarily captures the color information of the building (e.g., red, gray); Channel C3 primarily captures the overall shape or integrity of the building; Channel C4 primarily captures the vertical edges of the building; Channel C5 primarily captures the horizontal edges of the building. Based on these design objectives, the channels can be directly divided. Application channel features: channels C1 (texture), C2 (color), C3 (shape / integrity); corresponding average activation values ​​are [0.8, 0.7, 0.6]; Spatial channel characteristics: channels C4 (vertical edge), C5 (horizontal edge); corresponding average activation values ​​are [0.5, 0.4]; Two subsets were obtained: application channel feature set: {C1:0.8, C2:0.7, C3:0.6}; spatial channel feature set: {C4:0.5, C5:0.4}. This division clarifies which channel information is related to building attributes and which is related to spatial structure.

[0037] refer to Figure 4 In step S13, the specific steps are as follows: S131: Collect the features of the application channel and the spatial channel, match the features of the application channel and the spatial channel, and determine the corresponding matching coefficient based on the matching of the features of the application channel and the spatial channel. S132: Based on the detection of the second feature map, determine multiple feature information, determine semantic information based on the multiple feature information and the morphology of the second feature map, and determine the first sub-extraction method based on the semantic information and the second feature map; S133: Determine the second sub-extraction method based on the matching coefficient and the second feature map; determine the multi-scale feature extraction method based on the synthesis of the first and second sub-extraction methods; In the embodiments of this application, two sets of pre-classified features are obtained from the output of S123, which are the basis for subsequent matching. The two sets output by S123 are: Application Channel Features: These features contain information about what the building "looks like", such as roof color, texture, window distribution, etc., which are usually extracted by a specially designed convolutional kernel or attention mechanism. Assuming there are 3 application channel features, they can be represented as vectors: C1 (Application Channel 1): [0.72, 0.35, 0.28] (assuming these numbers represent the response intensity of the channel at different locations in the image, which is related to the material); C2 (Application Channel 2): ​​[0.15, 0.80, 0.42] (related to the roof shape); C3 (Application Channel 3): [0.30, 0.25, 0.90] (related to the window distribution).

[0038] Spatial Channel Features: These features typically contain information about where a building is and how it is shaped, such as edges, corners, and outlines. Suppose there are two spatial channel features: C4 (spatial channel 1): [0.50, 0.60] (representing edge information); C5 (spatial channel 2): ​​[0.10, 0.30] (representing corner information).

[0039] The model needs to have access to both sets; in neural networks, this typically means extracting tensors representing the application channel features (such as shape [Batch, NumAppChannels, Height, Width]) and tensors representing the spatial channel features (such as shape [Batch, NumSpatialChannels, Height, Width]) in preparation for the next step of processing; they need to be flattened or their dimensions adjusted for matching computation.

[0040] Calculating the similarity or correlation between applied channel features and spatial channel features helps the model understand which applied features (such as material) are relevant to which spatial features (such as edges) when describing the same building; there are many ways to calculate the matching, common ones include: Dot Product Similarity: This is one of the simplest and most commonly used methods; for each applied channel feature vector and a spatial channel feature vector, calculate their dot product. Cosine Similarity: Calculates the cosine of the angle between two vectors to measure their similarity in direction. It is not sensitive to the length of the vectors. Attention Mechanism: A more complex method that can dynamically learn the weight relationship between applied features and spatial features. For example, applied channel features can be used as queries and spatial channel features as keys. The similarity between the query and the key is calculated to obtain attention weights, which are then multiplied by the value (which can be the spatial channel feature itself or its transformation) to obtain a weighted spatial feature representation. Here, the dot product is mainly used as an example.

[0041] The matching result calculated in the previous step (the original similarity value) is converted into the final "matching coefficient." This coefficient is typically a value between 0 and 1, representing the strength or importance of the match. Method: Normalization: A common method is to use the Softmax function; for each feature in the application channel set, calculate the Softmax of its matching values ​​with all spatial channel features to obtain a set of probability distributions, which serve as the matching coefficients between the application channel and each spatial channel; similarly, a similar operation can be performed on the spatial channels, or bidirectional matching coefficients can be calculated. Fixed threshold / rule: A threshold is set based on experience or predefined rules, and matches that exceed the threshold are considered significant; The learned transformation: A small neural network layer learns how to transform the original matching values ​​into the final matching coefficients.

[0042] Specifically, suppose we are processing an image of a square building with a flat roof and regularly arranged windows; in S123, we identify: Application channel C1: high response to flat roof areas; Application channel C2: high response to regular window arrangement patterns; Application channel C3: high response to overall structural symmetry; Spatial channel C4: high response to edges in the image, with clear building outlines; Spatial channel C5: high response to corners in the image, with prominent building corners.

[0043] C1, C2, C3, C4, and C5 were collected; their matching was calculated, and it was found that C2 (regular window) and C4 (edge) had a high match, because regular windows are usually accompanied by clear edges; C3 (symmetry) and C4 (edge) also had a high match, because symmetrical structures often have clear edge contours; C1 (flat roof) and C4 (edge) also had a certain match, because flat roofs have clear edges; through Softmax calculation, the matching coefficient matrix M was obtained. This matrix tells the model: "When describing this building, the regular window feature (C2) and the edge feature (C4) are most closely related, and the symmetry feature (C3) and the edge feature (C4) are also closely related." This matching coefficient matrix M will play a role in the subsequent S133 step, helping the model to adjust or guide the building extraction method according to the synergistic relationship between applied features and spatial features. For example, more attention should be paid to regular windows and symmetry features in areas with rich edge information.

[0044] Furthermore, multiple feature information is determined based on the detection of the second feature map, semantic information is determined based on the multiple feature information and the morphology of the second feature map, and the first sub-extraction method is determined based on the semantic information and the second feature map. This approach takes into account both semantic information and the second feature map as a whole, ensuring the accuracy of the first sub-extraction method.

[0045] At this point, specific and measurable feature representations are extracted from the second feature map. Here, "detection" can be understood as feature extraction or feature response recognition. The second feature map is further processed to capture the different types of low-level features it contains. At the same time, global average pooling or max pooling is performed on the feature map to obtain the global feature representation of each channel. The self-attention mechanism is used to calculate the correlation between different positions and channels within the feature map, thereby obtaining the weighted feature representation.

[0046] By combining the multiple feature information extracted in the previous step with the overall structure (morphology) of the feature map itself, a higher-level semantic meaning is inferred. Here, "morphology" can refer to the spatial distribution of the feature map, the correlation between channels, the overall activation pattern, etc. At this point, recurring patterns in the feature information are identified and combined with the morphology of the feature map (such as the location and shape of high-response regions) to infer semantics. For example, if multiple channels related to "edges" form a closed ring-shaped high-response region on the feature map, it means that there is an object boundary. Predefined rules or templates are used to match feature information and morphology. For example, "if the responses of channels A and B are both high and their spatial response regions overlap, it is inferred to be a building." A specially designed network layer or model is used, with the input being feature information and a certain morphological description of the feature map (such as gradient, curvature), and the output being semantic labels or confidence scores. The output semantic information is usually some labels, category probabilities, or descriptive text / symbols, which represent the understanding of the scene content encoded by the feature map.

[0047] Based on the inferred semantic information, a preliminary feature extraction or processing strategy for specific semantic content is formulated. This "first sub-extraction method" is part of the subsequent multi-scale feature extraction method. At this point, different post-processing modules or operations are selected according to the semantic information. For example, if the semantic information is "buildings," a module that focuses more on edges and closed regions is selected; if the semantic information is "vegetation," a module that focuses more on texture and color is selected. The parameters of subsequent processing steps are adjusted according to the semantic information. For example, if it is determined to be a building, the threshold for edge detection can be increased, or the attention weight for high-texture regions can be increased. Different channels or regions of the second feature map are weighted according to the semantic information to highlight features related to the semantics. Output: a specific operation instruction or configuration, i.e., the "first sub-extraction method," which can be a function, a set of parameters, a module selection result, etc.

[0048] Specifically, suppose we are processing a satellite image containing multiple different types of buildings. Multiple features are detected from the second feature map F2: channel 5 (regular texture) has a high response, channel 10 (vertical edges) has a high response, and channel 15 (symmetry) has a moderate response. Simultaneously, the self-attention mechanism shows that channels 10 and 15 are strongly correlated in building regions. Combining these features with the shape of F2 (e.g., the high-response areas of channel 10 form several clear rectangular outlines), the semantic information is determined: "There are multiple rectangular structures with regular textures and vertical edges in the image; these are buildings." Based on this semantic information, the first sub-extraction method is determined: "Enhance vertical edge features, find closed rectangular regions, and focus on regions with regular textures." This method guides subsequent processing; for example, in the segmentation network, edge information is particularly emphasized, and attempts are made to identify regions with regular textures and clear edges as buildings.

[0049] Therefore, the second sub-extraction method is determined based on the matching coefficient and the second feature map; the multi-scale feature extraction method is determined based on the synthesis of the first and second sub-extraction methods, which takes into account the overall consideration of the synthesis of the first and second sub-extraction methods, ensuring the accuracy of the multi-scale feature extraction method. At the same time, the introduction of the second feature map takes into account the semantic information, the matching coefficient and the second feature map, improving the accuracy of the multi-scale feature extraction method and realizing the accurate extraction of multi-scale buildings.

[0050] At this point, the matching relationship between the application channel features and spatial channel features calculated in S131, and the second feature map generated in S122, are used to formulate an extraction method that focuses on feature collaboration. Only when application features (such as material and roof type) and spatial features (such as edges and corners) work together can a building be defined more accurately. The matching coefficient quantifies this collaborative effect.

[0051] Simultaneously, a matching coefficient matrix M (from S131) ​​is used to guide how to fuse applied channel features and spatial channel features. For example, for a certain region in the second feature map F2, if the matching coefficients show a high match between applied feature C2 (regular window) and spatial feature C4 (edge), then when processing this region, higher weights can be given to the feature map channels corresponding to C2 and C4, or they can be fused in a specific way (such as multiplying or concatenating them before feeding them into a small network). Based on the matching coefficients, which applied channel and spatial channel features are more important to the current region (or the entire image) are dynamically selected. For example, if the sum of the rows / columns of matrix M shows... Since C2 and C4 have the highest overall matching degree, when generating the second sub-extraction method, priority can be given to determining the building boundary or judging the internal attributes based on the features of C2 and C4. The matching coefficient can be used as a weight and applied to certain processing steps of the second feature map F2. For example, when performing upsampling or feature fusion, the contribution of features from different sources can be adjusted according to the matching coefficient. The second sub-extraction method usually manifests as a more refined strategy that considers feature collaboration. For example, "in areas with rich edge information, regular window features are used first to confirm the existence of buildings; in areas with uniform texture, symmetry features are used to assist in judging the shape."

[0052] The first sub-extraction method based on overall semantics obtained in S132 is combined with the second sub-extraction method based on feature collaboration obtained in S133.1 to form a final extraction strategy that can adapt to buildings of different scales and effectively utilize multiple features. At this point, the "synthesis" here is usually not a simple splicing, but a more complex fusion or decision-making process.

[0053] The advantages of the two sub-methods are combined. For example, the first sub-method indicates "finding large rectangular areas", while the second sub-method indicates "focusing on regular windows at the edges". The synthesized multi-scale feature extraction method is: "finding large rectangular areas, but prioritizing the check for regular window features at their edges and inside to confirm that they are buildings and refine the boundaries".

[0054] The first sub-method provides macro-level guidance (e.g., what types of buildings are in the current image and at what scale), while the second sub-method provides micro-level adjustments (e.g., how to use feature collaboration to achieve precise segmentation in a specific region). For example, if the first sub-method determines that there are tall buildings in the image (semantic information), then the weight of "using regular window features" in the second sub-method will be increased because tall buildings usually have more regular window arrangements.

[0055] By combining the two sub-methods, a network structure or processing flow that can handle features at different scales can be designed. For example, the combination of the two sub-methods can be applied to feature maps of different resolutions, or a module that can fuse multi-scale information can be designed, where the first sub-method provides global constraints and the second sub-method provides local detail guidance.

[0056] The final extraction method is a comprehensive strategy that considers both the overall shape and attributes of the building (first sub-method) and the synergistic relationship between the building's internal features and structural features (second sub-method). It can also adapt to buildings of different sizes. For example: "First, the building area is initially located based on the overall shape and texture (first sub-method). Then, within these areas, the synergistic relationship between regular window features and edge features is used to accurately delineate the building boundary and confirm its attributes (second sub-method). Different levels of detail are applied to candidate areas of different sizes (multi-scale)."

[0057] Optionally, assume there is a medium-sized flat-roofed building in the image; there is a matching coefficient matrix M, which shows that the regular window feature (C2) and the edge feature (C4) have a high matching degree; the second feature map F2 shows that the edge of the building area is clear (C4 response is high), but the window texture is not very regular (C2 response is moderate); based on the matching coefficient, the second sub-extraction method is determined: "In places with clear edges, the regular window feature is used first to assist in confirmation, but the edge information is mainly relied on to outline the contour."

[0058] The first sub-extraction method is to "find rectangular regions with regular textures and vertical edges." This is now combined with the second sub-extraction method. The combined multi-scale feature extraction method is as follows: "First, find rectangular regions with regular textures and vertical edges (first sub-method). For each candidate region found, focus on using edge features (C4) to accurately delineate the boundary at its edges. Simultaneously, check for the presence of regular window features (C2) within the region to further confirm it is a building, and perform more detailed segmentation inside (e.g., distinguishing between roofs and walls). For buildings smaller than this medium size, reduce reliance on regular window features and rely more on edges and overall shape. For larger buildings, increase the check on internal texture and symmetry features." This final multi-scale feature extraction method utilizes both the judgment of the building's overall shape and attributes, as well as the synergistic relationship between internal and structural features. Furthermore, it can adjust the strategy according to different building sizes, thus extracting buildings more robustly and accurately.

[0059] refer to Figure 5 In step S14, the specific steps are as follows: S141: The extraction method of multi-scale features is collected. The second feature map extracts features along the multi-scale feature extraction method and outputs multiple building features. The scale relationship of multiple building features is determined based on the matching of multiple building features, and the scale relationship of multiple building features is marked. S142: Collect the locations of multiple building features, and determine the features of the first building complex based on the location and scale relationship of the multiple building features; S143: Determine the second building group features based on the location of multiple building features and the shape of the second feature map; determine the optimized building group features based on the synthesis of the first building group features and the second building group features; In the embodiments of this application, a multi-scale feature extraction method is adopted. The second feature map is used to extract features along the multi-scale feature extraction method and output multiple building features. The scale relationship of multiple building features is determined based on the matching of multiple building features, and the scale relationship of multiple building features is marked. This takes into account the overall consideration of matching multiple building features and ensures the accuracy of the scale relationship of multiple building features.

[0060] At this point, the finalized multi-scale feature extraction method is obtained from S13 (especially S133). This method is the culmination of all previous steps (feature extraction, channel partitioning, matching, and sub-method synthesis). It tells the model "how" and "at what scale" to find and extract building features. The model needs to obtain a representation of this extraction method from the previous stage. This is usually a computational graph or a set of parameters / weights that defines how to process the second-level feature map to find buildings. For example, this method includes: Scale focus: Indicates at which resolutions / scales the model finds buildings of different sizes (e.g., focusing on roof texture at small scales and on building outlines at large scales). Feature focus: Indicates which types of features the model should prioritize (e.g., edges, corners, specific texture patterns such as regular windows, symmetry, etc.). Spatial layout focus: Indicates what kind of spatial layout pattern the model should look for (e.g., independent points, linear arrangement, clustering, etc.). It's like having a detailed treasure map. The map not only marks the area (scale) where the treasure exists, but also points out the characteristics of the treasure (such as shape, color, and material), as well as the ways in which they appear.

[0061] Based on scale and spatial layout considerations in the extraction method, a series of candidate regions are generated on the second feature map. This is achieved through a sliding window, a region proposal network (RPN), or a feature saliency-based method. For example, the model slides across the feature map, searching for regions that exhibit strong edges and regular textures across multiple scales. For each candidate region, features within the region are aggregated using the feature considerations specified in the extraction method (such as edges, textures, symmetry, etc.). This typically involves operations such as convolution and pooling to generate a fixed-length vector representing the candidate region's attributes as a building. A classifier (a fully connected layer or a more complex structure) is then used to determine whether each candidate region is truly a building. If it is a building, output the feature vector of the building (including its size, shape, texture, etc.). This step is also accompanied by instance segmentation to accurately delineate the boundaries of the building. Due to the overlap of candidate regions, redundant detections need to be removed, and only the most reliable one is kept. With a treasure map, carefully search on the actual area covered by the map (second-level feature map). According to the map, you will look for a stone of a specific shape and color on a certain hillside (specific scale) (feature focus). After you find several stones (candidate regions), you will carefully examine them (feature aggregation) to determine which ones are indeed treasures (classification / instance segmentation) and ignore those that just look like stones (deduplication).

[0062] The system analyzes and labels the size relationships (scale relationships) between multiple identified building instances for use in subsequent steps (such as S142); it compares the feature vectors of multiple buildings output from S141, which contain information about building size, shape, etc.; matching can be achieved by calculating the similarity between feature vectors (such as cosine similarity, Euclidean distance) or by directly comparing certain dimensions of the vectors (such as the dimension representing area); and it determines the relative size between buildings based on the matching results, for example: Master-slave relationship: If the feature vector of one building is significantly larger than that of another in the "size" dimension, and they are spatially close, it can be determined that they are a main building and an auxiliary building. Similarity: If the feature vectors of two buildings are very close in the "size" dimension, they can be judged to be buildings of similar size. Containment relationship: If the bounding box of one building completely contains another, and the feature vectors also show a size difference, a containment relationship can be determined. The determined scale relationships (such as "master-slave", "similar", "containment") are attached to the corresponding building features or the relationships between them. This can be a label, a value (such as size ratio), or an edge in a relationship graph. After the treasure hunt, you have multiple treasures you found (multiple building features). You start comparing their size, material, etc. (feature matching). You find that one treasure is particularly large and of superior material, while there is a smaller one next to it with slightly inferior material (scale relationship judgment: master-slave). Another treasure is about the same size as one you found earlier (scale relationship judgment: similarity). You record these relationships (relationship labeling).

[0063] Furthermore, the locations of multiple building features are collected, and the features of the first building group are determined based on the location and scale relationship of the multiple building features. This takes into account the overall consideration of the location and scale relationship of multiple building features, ensuring the accuracy of the features of the first building group.

[0064] At this point, the specific coordinates or region information of each individual building instance previously identified in S141 on the original image or feature map are obtained, which provides a basis for understanding their spatial layout; the specific coordinates or region information of each individual building instance previously identified in S141 on the original image or feature map are obtained, which provides a basis for understanding their spatial layout.

[0065] By utilizing the spatial proximity between buildings (determined by location) and previously established scale relationships (such as hierarchical relationships, size ratios, etc.), related buildings are grouped into a preliminary building complex. Here, the "characteristics of the first building complex" include a list of buildings within the complex, their spatial layout patterns (such as compact or loose), and the hierarchical structure within the complex (such as the relationship between the main building and auxiliary buildings).

[0066] Calculate the distance or overlap between building bounding boxes. Common methods include calculating the Euclidean distance between the center points of the bounding boxes or calculating the Intersection over Union (IoU). Set a threshold; if the distance is less than the threshold or the IoU is greater than the threshold, the two buildings are considered spatially adjacent. Combine this with the scale relationship determined in S141. For example, if two buildings are marked as master-slave, they are classified into the same building group even if they are slightly far apart (e.g., separated by a small courtyard). Conversely, if two buildings are similar in size and far apart, they do not belong to the same building group. Based on proximity and scale relationship, use clustering algorithms (such as DBSCAN, which can automatically form clusters based on distance) or rule-based grouping methods to group the buildings. For each group, record the list of buildings it contains, the overall bounding box of the group (the union of all member bounding boxes can be calculated), the average size of the buildings in the group, and the main scale relationship (e.g., whether there is a clear master building in the group).

[0067] Therefore, the second building group features are determined based on the location of multiple building features and the shape of the second feature map; the optimized building group features are determined based on the synthesis of the first and second building group features, which takes into account the overall consideration of the synthesis of the first and second building group features and ensures the accuracy of the optimized building group features.

[0068] At this point, we identify building groups that are spatially clustered and whose surrounding local environments (reflected by the morphology of the second feature map) are similar. Here, the "second building group features" focus more on describing the relationship between the building group and its external environment, such as whether they together form a block or are located in similar geographical environments (e.g., they are all buildings on the edge of a park).

[0069] Based on the initial building clusters (e.g., Group 1) formed in S142, or by directly considering the location of all building features, a spatial clustering threshold is set (e.g., the center point distance is less than a certain value, or the bounding boxes overlap). Sets of buildings that are spatially close to each other are found. For each candidate group, an area representing its surrounding environment is determined, typically a slightly larger area than the group itself (e.g., the group bounding box extends outward by a certain number of pixels, such as 20-50 pixels). Then, features are extracted from the second feature map (C2) within this area. These features include: texture features: using Local Binary Pattern (LBP), Histogram of Gradients (HOG), etc., to describe the texture complexity or pattern within the area; edge density / direction: calculating the intensity and main direction of edge feature channels within the area (such as the spatial channel features determined in S123); color / spectral features: if the feature map retains some spectral information, the average color or color distribution can be extracted; features of non-building areas: inferring the environment type through the features of the background area (non-building area), for example, high vegetation texture indicates proximity to a park, and high road edge features indicate proximity to a road.

[0070] The local environmental features of the candidate groups are compared. If the local environmental features of two or more candidate groups are very similar (e.g., using cosine similarity, Euclidean distance, etc.), they are considered to belong to the same environmental class. The set of buildings with similar environments and spatial clusters is divided into a "second building group". This feature can be represented as: Member list: [Building A, Building B, ...]; Group bounding box: Bounding box containing all members; Group environmental feature description: e.g., "high vegetation texture, low road density" or "high road edge features, low vegetation texture"; Group environment type inference: e.g., "park edge building group" or "street-side building group".

[0071] Merging two feature sets is necessary. For example, the first feature set includes a list of members, internal structure (such as master-slave relationships), average size, etc.; the second feature set includes a list of members (overlapping or containing more), environmental features, environmental type, etc. When merging, it is necessary to handle the overlapping or inclusion relationships of the member lists.

[0072] If the member lists of the first and second building group features are not completely consistent, they need to be merged or split based on which information is more reliable or according to preset rules. For example, if the first building group feature shows two buildings are closely related, but the second building group feature separates them due to environmental differences, further analysis is needed. The environmental information of the second building group feature can be used to enrich the description of the first building group feature, for example, by adding an "environment type" field to the first building group feature. The internal structural information of the first building group feature can be used to verify or refine the members of the second building group feature. For example, if the second building group feature contains a building that is obviously isolated, but the first building group feature shows that it is far removed from the group core, it needs to be excluded or treated separately.

[0073] The final optimized building cluster feature is a representation that integrates internal structure and external environment. It includes: a unified and corrected list of members; internal structural information of the cluster (from S142); environmental characteristics and type information of the cluster (from S143.1); the relationship between the cluster and other clusters (such as adjacency, isolation, etc.); and a comprehensive confidence or quality score.

[0074] refer to Figure 6 In step S15, the specific steps are as follows: S151: Monitor the feature extraction of the second feature map in real time, collect the feature extraction events of the second feature map, determine the corresponding abnormal extraction events based on the identification of the feature extraction events of the second feature map, and determine the corresponding extraction loss part based on the analysis of the abnormal extraction events. S152: Collect the building distribution information of the optimized building cluster features, determine the first building distribution map based on the optimized building cluster features and the extracted loss part, and determine the final building distribution map by synthesizing the first building distribution map and the actual image of the target building area.

[0075] In the embodiments of this application, feature extraction of the second feature map is monitored in real time, and feature extraction events of the second feature map are collected. Based on the identification of feature extraction events of the second feature map, corresponding abnormal extraction events are determined, and the corresponding extraction loss part is determined according to the analysis of the abnormal extraction events. This approach is compatible with the overall consideration of feature extraction event identification of the second feature map and ensures the accuracy of the corresponding abnormal extraction events.

[0076] At this point, when the feature extraction algorithm processes the second feature map, the system does not operate "blindly" but continuously observes the process. It records the key operations performed by the algorithm on the feature map and the intermediate results generated. These recorded operations and results are the "feature extraction events." For example, the algorithm scans every region of the feature map, records which regions are initially determined to contain buildings (e.g., by detecting strong edge features, specific texture patterns, etc.), records which types of features are extracted (such as edge intensity, corner positions, color histograms, etc.), and the positions and intensities of these features. These records are all made in real time.

[0077] The system analyzes the collected feature extraction events according to preset rules or patterns to determine which events or combinations of events indicate problems or errors in the extraction process. These problems include: an area that should be a building is not marked (omission), a non-building area is incorrectly marked (false detection), the feature extraction result of a certain area is seriously inconsistent with the expected pattern, or the feature extraction process gets stuck or fails in a certain area. Identifying abnormal events is the first step in self-correction.

[0078] Once an anomaly is identified, the system needs to analyze in depth what specific information loss or error was caused by the anomaly. For false positives, the loss is "incorrectly identifying non-building areas as buildings"; for omissions, the loss is "failure to identify actual buildings" or "failure to extract the complete boundaries / features of buildings". Analyzing anomalies is to quantify or explicitly describe what information that should have been there will be missing or what information that should not have been there will be added in the final result because of this anomaly.

[0079] The system achieves dynamic monitoring and self-diagnosis of the feature extraction process. It first records key events in the extraction process, then uses rules or patterns to identify the erroneous links, and finally analyzes in depth what information deviations or omissions were caused by these errors. In this way, the system can clearly know where the problems occurred and what the specific problems are, providing a basis for subsequent corrections (such as adjusting the final distribution map in S152). In the example above, the system knows that there is a falsely detected "building" in the park (the loss part is the erroneous information) and an omitted "building" on the edge of the residential area (the loss part is the missing information).

[0080] Furthermore, the building distribution of the optimized building cluster features is collected. Based on the optimized building cluster features and the extracted loss component, a first building distribution map is determined. The final building distribution map is determined by synthesizing the first building distribution map and the actual image of the target building area. This approach takes into account the overall consideration of synthesizing the first building distribution map and the actual image of the target building area, ensuring the accuracy of the final building distribution map. At the same time, it realizes the further transformation of the optimized building cluster features and introduces the extracted loss component. This approach takes into account the extracted loss component, the optimized building cluster features, and the actual image of the target building area, thus improving the accuracy of the final building distribution map.

[0081] At this point, information about how the building complexes are spatially distributed is extracted from the results obtained in the previous steps (especially S142 and S143); "optimized building complex features" means that these features have been initially corrected, for example, the boundaries or members of the building complexes have been adjusted based on some of the anomalies identified in S151; "building distribution" specifically refers to: the location of the building complexes: the coordinate range or center point of each building complex in the image; the size and shape of the building complexes: the area covered by each building complex and its approximate outline; the relative positions between building complexes: whether the building complexes are densely or sparsely distributed, the distance between them, etc.; the internal composition of the building complexes: for example, the relative positional relationship between the main building and the auxiliary buildings (if S143 has been refined to this extent).

[0082] The data is typically a structured representation, such as a list, where each element represents a group of buildings, containing information such as the group's bounding box, center point, area, and list of members (if applicable).

[0083] Using the collected building distribution information and the identified "extraction loss components" (i.e., errors and omissions), a preliminary but corrected building distribution map is generated. This "first building distribution map" is usually a binary mask or a list with bounding boxes. If S151 finds that a certain area is incorrectly identified as a building (e.g., trees in a park are misidentified), and the optimization in S143 has partially corrected the boundary of building cluster A so that it no longer completely includes the tree area, then this step further ensures that this area is excluded from the first building distribution map. If S151 finds that a small building (such as a hut on the edge of a residential area) is omitted, and this information is not reflected in the optimized building cluster features (or is reflected as no building cluster in the area), then it is necessary to manually or semi-automatically add the marker of this omitted building to the first building distribution map based on the information of the loss components (e.g., setting 1 at the corresponding position in the mask, or adding a new bounding box).

[0084] The first building distribution map can be generated directly based on the optimized building cluster boundaries, and then locally modified according to the loss function. For example, if the bounding box of building cluster A is (x1,y1,x2,y2), but the loss function indicates that the region from (xa,ya) to (xb,yb) is a false detection, then the first building distribution map will mark this region as 0 (non-building); if the loss function indicates that there is a missing hut at the position from (xc,yc) to (xd,yd), then a region or bounding box marked as 1 will be added to the first building distribution map.

[0085] The "first building distribution map" (usually an abstract mask or bounding box) generated in the previous step is combined with the original "actual image" which contains rich visual information to generate the final, more accurate and reliable building distribution map. The "composite" here is not a simple image overlay, but uses the information of the actual image to further verify, refine or improve the first building distribution map.

[0086] At this point, the boundary of the first building distribution map is compared with the building edges in the actual image; if a boundary in the first distribution map is found to deviate significantly from the building outline in the actual image, it can be adjusted to make it more realistic.

[0087] By utilizing contextual information from the actual images, for example, if the first distribution map marks an area as a building, but the actual image shows that the area is water or a large grassy area, it can be corrected to be non-building; conversely, if the first distribution map misses an area, but the actual image shows that there is clearly a building roof or structure there, it can be added (although we have already tried to fill in the omissions, we can use image information again to confirm this).

[0088] Some morphological operations (such as dilation and erosion) are applied to smooth the boundaries, remove small noise points (if they are not buildings but are marked), or connect different parts of the same building that have been incorrectly segmented; the final building distribution map is a high-quality result that contains both structured information based on feature extraction and optimization and has been visually verified and corrected by actual images, so it is closer to the real situation. Its form can be an accurate boundary mask, polygonal outline, or a list of bounding boxes with confidence.

[0089] Optionally, the first building distribution map is a corrected mask; now it is composited with the original aerial image: the mask is overlaid on the image, and the boundary of building group A is observed; it is found that the mask boundary covers a small patch of grass slightly more than the actual roof edge in a certain corner; based on the actual image, the mask boundary is slightly shrunken inward to exclude that patch of grass; the newly added small building marker area (xc, yc) to (xd, yd) is checked; the actual image shows that there is indeed a small house roof there, with a regular shape, which is clearly different from the surrounding environment (grass or path); this marker is confirmed to be correct; the boundary of building group B is checked; the actual image shows that it is a standalone bungalow, and the mask boundary matches its outline very well, requiring no adjustment; some morphological operations are also performed, such as slight erosion, to remove small isolated white areas (non-buildings) in the mask caused by noise; the final "final building distribution map" is an accurate mask (or outline) that accurately delineates the boundaries of all buildings in the image, eliminates false detections, fills in omissions, and the boundaries are highly consistent with the shape of the actual buildings.

[0090] Please see Figure 7 , Figure 7 This is a schematic diagram of the structural composition of the multi-scale building extraction system in an embodiment of the present invention; the multi-scale building extraction system includes: The first feature map module 21 is used to extract local features based on the depth-separable convolution processing of the remote sensing image of the target building area, so as to output the corresponding first feature map. Feature module 22 is used to determine a second feature map based on the first feature map, the attention processing channel, and the spatial attention channel, and to label the features of the application channel and the features of the spatial channel; The extraction method module 23 is used to determine the corresponding matching coefficient based on the features of the application channel and the features of the spatial channel, determine the semantic information based on the detection of the second feature map, and determine the extraction method of multi-scale features based on the semantic information, the matching coefficient and the second feature map. Building cluster feature module 24 is used to determine multiple building features based on the extraction method of the second feature map and multi-scale features, and to mark the scale relationship of multiple building features. Based on the location, scale relationship and shape of the second feature map of multiple building features, the optimized building cluster features are determined. The building distribution map module 25 is used to determine the corresponding extraction loss part based on the recognition of feature extraction events in the second feature map, and to determine the final building distribution map based on the extraction loss part, the building distribution of the optimized building group features, and the actual image of the target building area. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method of extraction of a multi-scale building, characterized in that, The method comprises the following steps: extracting local features based on deep separable convolution processing of remote sensing images of a target building area to output a corresponding first heavy feature map; determining a second heavy feature map according to the first heavy feature map, an attention processing channel and a spatial attention channel, and marking the features of the application channel and the features of the spatial channel; determining a corresponding matching coefficient based on the features of the application channel and the features of the spatial channel, determining semantic information according to the detection of the second heavy feature map, and determining the extraction mode of multi-scale features based on the semantic information, the matching coefficient and the second heavy feature map; determining a plurality of building features according to the second heavy feature map and the extraction mode of multi-scale features, and marking the scale relationship of the plurality of building features, and determining the optimized building group features according to the position of the plurality of building features, the scale relationship and the shape of the second heavy feature map; determining a corresponding extraction loss part according to the recognition of the feature extraction event of the second heavy feature map, and determining a final building distribution map based on the extraction loss part, the building distribution of the optimized building group features and the actual image of the target building area.

2. The method for extracting a building at multiple scales according to claim 1, wherein, The method comprises the following steps: acquiring the position of the target building area, acquiring the remote sensing image of the target building area according to the positioning detection of the position of the target building area, and triggering deep separable convolution processing of the remote sensing image of the target building area based on the remote sensing image of the target building area and the corresponding convolution model; during the deep separable convolution processing of the remote sensing image of the target building area, acquiring a plurality of convolution layers unfolded in sequence, and determining a corresponding binary segmentation map according to the processing of the plurality of convolution layers; determining local features based on the extraction of the binary segmentation map, determining the first heavy feature map according to the synthesis of the position and shape of each local feature, and outputting the corresponding first heavy feature map.

3. The method for extracting a building at multiple scales of claim 1, wherein, The method comprises the following steps: acquiring the first heavy feature map, determining a first feature region according to the first heavy feature map and the attention processing channel, and determining a second feature region according to the first heavy feature map and the spatial attention channel; determining a corresponding feature map construction mode according to the first feature region, the second feature region and the feature map construction mode mapping relationship, and determining the second heavy feature map according to the feature map construction mode, the first feature region and the second feature region; determining a plurality of channel features based on the detection of the second heavy feature map, and determining the features of the application channel and the features of the spatial channel according to the screening of the plurality of channel features.

4. The method for extracting a building at multiple scales of claim 1, wherein, The method comprises the following steps: acquiring the features of the application channel and the features of the spatial channel, matching the features of the application channel and the features of the spatial channel, and determining a corresponding matching coefficient according to the matching of the features of the application channel and the features of the spatial channel.

5. The method for extracting a building from a multiscale image according to claim 4, wherein, The application channel-based feature and the spatial channel-based feature determine a corresponding matching coefficient, semantic information is determined according to detection of the second heavy feature map, the extraction mode of the multi-scale feature is determined based on the semantic information, the matching coefficient and the second heavy feature map, and the method further comprises: A plurality of feature information is determined based on detection of the second heavy feature map, the semantic information is determined according to the plurality of feature information and the form of the second heavy feature map, and the first sub-extraction mode is determined according to the semantic information and the second heavy feature map. The second sub-extraction mode is determined according to the matching coefficient and the second heavy feature map, and the extraction mode of the multi-scale feature is determined according to a combination of the first sub-extraction mode and the second sub-extraction mode.

6. The method for extracting a building from a multiscale image according to claim 1, wherein, The plurality of building features is determined according to the second heavy feature map and the extraction mode of the multi-scale feature, and the scale relationship of the plurality of building features is marked, the optimized building group feature is determined according to the position of the plurality of building features, the scale relationship and the form of the second heavy feature map, and the method comprises: The extraction mode of the multi-scale feature is collected, the second heavy feature map performs feature extraction along the extraction mode of the multi-scale feature, and a plurality of building features is output, the scale relationship of the plurality of building features is determined according to matching of the plurality of building features, and the scale relationship of the plurality of building features is marked. The position of the plurality of building features is collected, and the first building group feature is determined according to the position of the plurality of building features and the scale relationship.

7. The method for extracting a building from a multiscale image according to claim 6, wherein, The plurality of building features is determined according to the second heavy feature map and the extraction mode of the multi-scale feature, and the scale relationship of the plurality of building features is marked, the optimized building group feature is determined according to the position of the plurality of building features, the scale relationship and the form of the second heavy feature map, and the method further comprises: The second building group feature is determined according to the position of the plurality of building features and the form of the second heavy feature map, and the optimized building group feature is determined based on a combination of the first building group feature and the second building group feature.

8. The method for extracting a building from a multiscale image according to claim 1, wherein, The corresponding extraction loss part is determined according to recognition of the feature extraction event of the second heavy feature map, the final building distribution map is determined based on the extraction loss part, the building distribution of the optimized building group feature and the actual image of the target building area, and the method comprises: The feature extraction of the second heavy feature map is monitored in real time, the feature extraction event of the second heavy feature map is collected, the corresponding abnormal extraction event is determined based on recognition of the feature extraction event of the second heavy feature map, and the corresponding extraction loss part is determined according to analysis of the abnormal extraction event.

9. The method for extracting a building from a multiscale image according to claim 8, wherein, The corresponding extraction loss part is determined according to recognition of the feature extraction event of the second heavy feature map, the final building distribution map is determined based on the extraction loss part, the building distribution of the optimized building group feature and the actual image of the target building area, and the method further comprises: The building distribution of the optimized building group feature is collected, the first building distribution map is determined according to the building distribution of the optimized building group feature and the extraction loss part, and the final building distribution map is determined according to a combination of the first building distribution map and the actual image of the target building area.

10. A multi-scale building extraction system characterized by, The multiscale building extraction system is applied to the multiscale building extraction method as claimed in any one of claims 1-9, and the multiscale building extraction system comprises: a first heavy feature map module, configured to extract local features based on deep separable convolution processing of remote sensing images of a target building area, to output corresponding first heavy feature maps; a feature module, configured to determine second heavy feature maps according to the first heavy feature maps, attention processing channels, and spatial attention channels, and mark features of the application channels and features of the spatial channels; an extraction manner module, configured to determine corresponding matching coefficients based on the features of the application channels and the features of the spatial channels, determine semantic information according to detection of the second heavy feature maps, and determine an extraction manner of multiscale features based on the semantic information, the matching coefficients, and the second heavy feature maps; a building group feature module, configured to determine a plurality of building features according to the second heavy feature maps and the extraction manner of the multiscale features, mark scale relationships of the plurality of building features, and determine optimized building group features according to positions of the plurality of building features, the scale relationships, and a shape of the second heavy feature maps; and a building distribution map module, configured to determine corresponding extraction loss parts according to recognition of feature extraction events of the second heavy feature maps, determine a final building distribution map based on the extraction loss parts, building distribution situations of the optimized building group features, and actual images of the target building area.