Method for extracting rural residential area and judging building density by using tile data

By combining map tile data and deep learning technology, rural residential areas and determining the density of buildings is solved, and the problems of high workload and low efficiency in traditional methods are achieved, and high-precision and high-efficiency extraction and judgment are achieved.

CN120147884AActive Publication Date: 2025-06-13CHANGGUANG SATELLITE TECH CO LTD

Patent Information

Application Number
CN202510215697.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The traditional rural residential extraction method has problems such as large workload, high cost, long cycle and untimely information acquisition, which is difficult to meet the requirements of high efficiency.

Method used

Combining map tile data and deep learning technology, rural residential areas are extracted and building density is determined by preprocessing remote sensing images, building data sets, training deep learning models, cutting and combining tile images.

Benefits of technology

It improves the accuracy and efficiency of extraction of rural residential areas and determination of building density, reduces hardware resource consumption, and is suitable for large-scale extraction of rural residential areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147884A_ABST
    Figure CN120147884A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of optical remote sensing image application, and provides a method for extracting a rural residential area and judging the closeness of a building by using tile data, which is used for extracting the rural residential area and judging the closeness of the building by combining map tile data and a deep learning technology, and judging the closeness of the building through tile resolution characteristics and actual characteristics of the rural residential area. 14-level tile images are selected for rural residential area distribution extraction, and after the 14-level tile images are obtained, a model obtained through SegFormer-based training is used for extracting rural residential area distribution; the method comprises the following steps: selecting 18-level tile images to extract rural building distribution, extracting rural buildings by using a model obtained by training based on DeepLabV3 +, and selecting medium-resolution hierarchical images to interpret, thereby improving the precision and efficiency of rural residential area extraction and building density determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical remote sensing image applications, and particularly to a method for extracting rural residential areas and determining building density using tile data. Background Art

[0002] Rural residential areas are an important part of rural geography. How to quickly and accurately extract the distribution of rural residential areas and determine their building density has important application significance in fields such as rural development planning, village consolidation, and population census. Traditional methods for extracting residential areas mainly rely on manual ground surveys, which have problems such as large workload, high cost, long cycle, and untimely information acquisition, and are difficult to meet the requirements of high-efficiency work.

[0003] Using the method of remote sensing image interpretation can greatly reduce the number of staff and workload. However, in the process of large-scale interpretation, using raster images has high requirements for electronic hardware and is inconvenient to operate. At present, remote sensing image services based on the tile form are becoming more and more common, which is not only convenient for online viewing, but also has data of different resolution levels available. For example, the tile map service is an online service that follows the pyramid model. As Figure 1 shown, the tile map pyramid model is a multi-resolution hierarchical model. From the bottom layer to the top layer of the tile pyramid, the resolution becomes lower, but the represented geographical range remains unchanged; for the same geographical area, the higher the layer in the tile pyramid, the lower the resolution and the smaller the data volume; the lower the layer in the tile pyramid, the higher the resolution and the larger the data volume.

[0004] In recent years, a new generation of machine learning methods represented by deep learning has made breakthrough progress in related fields of computer vision and has gradually penetrated into the research in the cross-field of computer vision and satellite remote sensing technology. Its application makes it possible to automate and efficiently process satellite image data, greatly improving the accuracy and efficiency of information extraction.

[0005] Therefore, how to combine map tile data and deep learning technology to extract rural residential areas and determine building density is a key technical problem that needs to be solved urgently at present. Summary of the Invention

[0006] In order to combine map tile data and deep learning technology to improve the accuracy and efficiency of rural residential area extraction and building density determination, the present invention proposes the following technical solutions:

[0007] A method for extracting rural residential areas and determining building density using tile data, as Figure 2 shown, includes the following steps:

[0008] Step 1. Preprocess the remote sensing image: Preprocess the sub-meter satellite remote sensing image used, and set the corresponding resolutions for 14-level tiles and 18-level tiles respectively, so as to obtain the remote sensing images that can be used for two types of ground objects, namely rural residential areas and rural buildings;

[0009] Step 2. Construct the rural residential area dataset and rural building dataset for deep learning training;

[0010] Step 3. Train the deep learning extraction model, including the rural residential area extraction model and the rural building extraction model;

[0011] Step 4. Crawl the 14-level tile images within the prediction area. Through tile combination, cut the rural residential area remote sensing image into image blocks of the same size as the data in the rural residential area dataset, record the position in the original image, and generate the remote sensing image that can be predicted;

[0012] Step 5. Input the image blocks in Step 4 into the rural residential area extraction model to obtain the rural residential area prediction result, and obtain the contour vector patch that conforms to the actual boundary of the rural residential area;

[0013] Step 6. According to the contour vector patch of the actual boundary of the rural residential area in Step 5, crawl the 18-level tile images within the prediction area respectively. Through tile combination, cut the rural residential area remote sensing image into image blocks of the same size as the data in the rural building dataset, record the position in the original image, and generate the remote sensing image that can be predicted;

[0014] Step 7. Input the image blocks in Step 6 into the rural building extraction model to obtain the rural building prediction result, and obtain the contour vector patch that conforms to the actual boundary of the rural building;

[0015] Step 8. Determine the rural building density level according to the rural building prediction result, assign the obtained rural building density level to each rural residential area, and obtain a vector file containing the distribution of rural residential areas and the building density attributes of each residential area.

[0016] Technical effect:

[0017] The present invention combines map tile data and deep learning technology to extract rural residential areas and determine the building density. Through the tile resolution characteristics and the actual characteristics of rural residential areas, 14-level tile images are selected to extract the distribution of rural residential areas. After obtaining the 14-level tile images, the model trained based on SegFormer is used to extract the distribution of rural residential areas and convert it into patch data.

[0018] Meanwhile, considering the characteristic of relatively small floor area of rural buildings, the present invention selects high-resolution remote sensing images as the data source. Combining the tile resolution characteristics and the actual characteristics of rural buildings, 18-level tile images are selected for the extraction of rural building distribution. According to the scope of rural residential area patches, 18-level tile images within this area are specifically obtained. This method can exclude areas without residential areas such as farmland and forest land, reduce the workload of prediction, reduce the consumption of hardware resources, greatly improve work efficiency, and use the model trained based on DeepLabV3+ to extract rural buildings, and calculate the analysis result of the building density in rural residential areas.

[0019] During the extraction process of large-scale rural residential areas, due to the large scope, if high-resolution image data is used, it will consume more computing resources and the prediction time will also be longer. Considering the obvious characteristics of rural residential areas, the present invention innovatively selects medium-resolution hierarchical images for interpretation, so as to improve the accuracy and efficiency of the extraction of rural residential areas and the determination of building density. The method proposed by the present invention for the extraction details of the distribution of some rural residential areas is as follows Figure 5 and Figure 6 shown.

[0020] In order to verify the effect of the present invention in using Jilin-1 sub-meter satellite images and their sliced data for residential area extraction, an area is selected for experimental verification. Manually delineate all residential areas in this area as accuracy verification sample squares, and select two methods, HRNet-OCR and DeepLabV3+, for comparison, which proves that the method proposed by the present invention has a good effect. The comparison methods select producer accuracy (PA) and user accuracy (UA), and the comparison accuracies are shown in the following table.

[0021] Method PA UA HRNet-OCR 86.87% 83.30% DeepLabV3+ 84.21% 82.60% The present invention SegFormer 88.62% 85.74%

[0022] Although HRNet-OCR and DeepLabV3+ have certain advantages in segmentation accuracy, their complex structures and high computational complexity may not be suitable for real-time applications or resource-constrained scenarios. While SegFormer is based on the encoder structure of Transformer, which can effectively model long-range dependencies, which is particularly important for the segmentation of residential areas because the boundaries and structures of residential areas may span a large spatial range. SegFormer can handle residential areas of different scales while maintaining computational efficiency. Therefore, the method proposed by the present invention has the highest value in the accuracy verification and has practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of the tile map pyramid model.

[0024] Figure 2 is the overall flow block diagram of the present invention.

[0025] Figure 3 It is a block diagram of the network structure of the rural residential area extraction model.

[0026] Figure 4 It is a block diagram of the network structure of the rural building extraction model.

[0027] Figure 5 It is the original map of the distribution of rural residential areas in a certain prediction area.

[0028] Figure 6 It is a schematic diagram of the extraction result of rural residential areas after the present invention. Specific implementation manners

[0029] For a better understanding of the technical solution of the present invention, the embodiments provided by the present invention will be described in detail below with reference to the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0030] Step 1, preprocess the remote sensing image: Preprocess the sub-meter satellite remote sensing image used, and set the corresponding resolutions of 14-level tiles and 18-level tiles respectively, so as to obtain the remote sensing images that can be used for two types of ground objects, namely rural residential areas and rural buildings; wherein the preprocessing process includes atmospheric correction, geometric correction, image fusion and image registration.

[0031] Step 2, construct a rural residential area dataset and a rural building dataset for deep learning training;

[0032] Manually draw the contours of rural residential areas on the low-resolution remote sensing images of rural residential areas obtained in Step 1. The distribution of rural residential areas in the images should include the following characteristics: intensive type, normal distribution type, sparse type and scattered type. During the manual drawing process, try to fit the edge of the residential area contour as much as possible, and divide the image data and annotation data into blocks to obtain the initial rural residential area dataset;

[0033] Manually draw the contours of rural buildings on the high-resolution remote sensing images of rural buildings obtained in Step 1. Select the data of areas with a relatively complete range of building types during the selection of remote sensing images to ensure that the dataset includes basic rural building types such as single-story, multi-story, intensive, and sparse. During the manual drawing process, try to fit the edge of the external contour of the building as much as possible; and divide the image data and annotation data into blocks to obtain the initial rural building dataset;

[0034] Perform data augmentation on the initial rural residential area dataset and the initial rural residential area dataset to obtain the rural residential area dataset and the rural building dataset; further, the data augmentation includes a random combination of horizontal flipping, vertical flipping, random rotation and affine transformation methods.

[0035] Step 3: Train deep learning extraction models, including rural residential area extraction model and rural building extraction model:

[0036] The network structure diagram of the rural residential area extraction model is as Figure 3 shown. In the extraction of rural residential areas, the SegFormer model structure is selected as the basis for training the extraction model and prediction. SegFormer has a relatively large effective receptive field, which enables the model to capture more context information, thereby improving the accuracy of segmentation and being more suitable for the extraction task of rural residential areas. In the process, a hierarchical Transformer Block that does not require positional encoding is used, which can capture higher-resolution features and thus obtain richer detailed information. At the same time, a lightweight decoder with all MLP (Multi-Layer Perceptron) layers is used, which does not require complex calculations and high computing resources and can obtain effective feature expressions.

[0037] The specific process is as follows:

[0038] 3.1a. Input remote sensing images of rural residential areas where H and W are the height and width of the image respectively, and 3 represents the three RGB channels;

[0039] The input image is segmented into non-overlapping Patches, and the size of each Patch is P×P, then the image is segmented into multiple Each Patch is flattened into a vector and passed through a linear projection to obtain Patch Embedding:

[0040]

[0041] where C is the dimension of Patch Embedding, Flatten represents flattening the Patch, and Linear represents linear projection;

[0042] 3.1b. Based on the SegFormer model structure, the Patch Embedding passes through multiple stages of Transformer Blocks in the encoder. The resolution of the output feature map gradually decreases and the number of channels gradually increases in each stage. The encoder can capture feature information from low level to high level through the Transformer Block and generate multi-scale feature maps;

[0043] The encoder is the core part of SegFormer. It adopts a hierarchical Transformer structure, can capture multi-scale feature information, consists of multiple stages, and each stage contains several Transformer Blocks.

[0044] In this embodiment, it is divided into 4 stages (Block1 to Block4), as Figure 3 shown. The resolution of the output feature map of each stage gradually decreases, and the number of channels gradually increases.

[0045] Output of Stage 1

[0046] Output of Stage 2

[0047] Output of Stage 3

[0048] Output of Stage 4

[0049] Through the Transformer Blocks of multiple stages, the encoder can capture feature information from low levels to high levels and generate multi-scale feature maps {F 1 , F 2 , F 3 , F 4}, which respectively correspond to the outputs of the above 4 stages.

[0050] Furthermore, the Transformer Block is composed of the MHSA self-attention mechanism and the FFN feed-forward neural network.

[0051] MHSA self-attention mechanism:

[0052]

[0053] Among them, Q, K, and V respectively represent Query, Key, and Value, and d k is the dimension of the Key;

[0054] FFN feed-forward neural network:

[0055] FFN(x) = Linear(GELU(Linear(x))), GELU is the activation function, and Linear represents linear projection.

[0056] 3.1c. The decoder is composed of an MLP layer, an upsampling layer, and a 1x1 convolutional layer, and can fuse the multi-scale feature maps extracted by the encoder, and perform channel adjustment and feature enhancement on the multi-scale feature maps {F 1 , F 2 , F 3 , F 4}:

[0057] F i ' = MLP(F i ), i ∈ {1, 2, 3, 4}

[0058] Then, upsample the adjusted feature map to the same resolution (usually 1 / 4 the size of the input image) and perform concatenation:

[0059] F fused = Concat(Upsample(F 1 '), Upsample(F 2 '), Upsample(F 3 '), Upsample(F 4 '))

[0060] Finally, generate the final segmentation result Y through a lightweight MLP layer and a 1x1 convolutional layer:

[0061]

[0062] where F fused is the multi-scale fusion feature map and N is the number of classes;

[0063] Each pixel corresponds to a class probability distribution, and the predicted class of each pixel is obtained through the argmax operation:

[0064] Prediction = argmax(Y),

[0065] Apply the trained rural residential area extraction model to the prediction of the rural residential area range.

[0066] The network structure diagram of the rural building extraction model is as Figure 4 shown, and the specific process is as follows:

[0067] 3.2a. Input the remote sensing image of the rural residential area where H and W are the height and width of the image respectively, 3 represents the RGB three channels, and normalize the input image, scale the pixel values to the range [0, 1] for easy model training.

[0068] 3.2b. Based on the DeepLabV3+ model structure, in the encoder, pass through the SENet154 backbone network. Further, its core feature is the introduction of the SE (Squeeze-and-Excitation) module, which can adaptively adjust the importance of channel features, extract multi-level feature information from the input image, including the Squeeze module, which compresses the spatial information of each channel into a scalar through global average pooling (GlobalAverage Pooling, GAP):

[0069]

[0070] Among them, x c (i, j) represents the eigenvalue of the c-th channel at the position (i, j), and z c is the compression value of the c-th channel.

[0071] The Excitation module generates channel attention weights through a fully connected layer (ReLU) and an activation function (Sigmoid):

[0072] s = σ(W 2 ·ReLU(W 1 ·z))

[0073] Among them, W 1 and W 2 are the weights of the fully connected layer, σ is the Sigmoid function, and s is the channel attention weight.

[0074] Multiply the channel attention weight by the original feature to obtain the recalibrated feature:

[0075] SENet154 outputs feature maps of 6 different scales. Among them, the feature map with a scale of 1 / 16 and 2048 channels is used as the main feature and input into the ASPP module for information processing;

[0076] Furthermore, the ASPP module uses 4 parallel dilated convolution branches, which respectively adopt different dilation rates and a global average pooling branch for multi-scale feature extraction. In this embodiment, the different dilation rates are 1, 6, 12, and 18 respectively. The dilation rate of 1 captures local detail information; the dilation rates of 6, 12, and 18 capture context information in a larger range; the global average pooling captures global context information;

[0077] The formula for dilated convolution is

[0078] Among them, r is the dilation rate, and w(m, n) is the weight of the convolution kernel;

[0079] Stitch the extracted multi-scale feature maps together and perform channel adjustment through a 1x1 convolution:

[0080] F ASPP = Conv 1×1 (Concat(F 1 , F 8 , F 12 , F 18 , F GAP ))

[0081] Among them, F 1 , F 6 , F 12 , F 18are feature maps with different porosity rates, F GAP is the global average pooling feature map.

[0082] 3.2c. In the decoder, fuse high-level features with low-level features. In this embodiment, the low-level features are low-level feature maps F with a scale of 1 / 4 and 256 channels extracted from the SENet154 backbone network low ; the high-level features are the feature maps F output by the ASPP module ASPP upsampled to a scale of 1 / 4; concatenate the low-level features and the high-level features, and fuse them through a convolutional layer:

[0083] F fused′ = Conv 3×3 (Concat(F low , Upsample(F ASPP ))

[0084] Map the fused features to the class space through an upsampling layer and a 1x1 convolutional layer, and output the final segmentation result Z,

[0085]

[0086] where F fused′ is the fused feature, N is the number of classes, and the predicted class of each pixel is obtained through the argmax operation:

[0087] Prediction = argmax(Z),

[0088] Apply the trained rural building extraction model to rural building prediction.

[0089] Step 4: Crawl 14-level tile images:

[0090] Parse the geographical range boundary included in the vector file of the area to be predicted. According to the range, use the tile crawling technology to crawl the 14-level tile images in the area in the order from left to right and from top to bottom, and record the position of each tile; Reassemble the raster image file of the predicted area according to the recorded tile positions, and assign geographical information according to the prediction range; Crop the remote sensing image into image blocks of the same size as the dataset data, record the position in the original image, and the overlap rate is set to 0.05 in this embodiment.

[0091] Step 5: Input the image blocks in Step 4 into the rural residential area extraction model to obtain the prediction results of the rural residential area and obtain the contour vector patches that conform to the actual boundary of the rural residential area.

[0092] Further, use the raster-to-vector function in GDAL to convert the raster extraction results of residential areas into vector results, retain the residential area part, and delete the background value part; optimize the obtained residential area results morphologically, and use the Douglas-Peucker algorithm to reduce the amount of point data. Smooth the burrs in the contour to obtain the final contour vector patch that conforms to the actual boundary of rural residential areas.

[0093] Step 6: Crawl 18-level tile images:

[0094] According to the contour vector patches of the actual boundaries of rural residential areas in Step 5, parse the geographical range boundaries they contain. Use the tile crawling technology to crawl 18-level tile images within the prediction area in the order from left to right and from top to bottom, record the position of each tile, and reassemble the raster image file of the prediction area according to the recorded tile positions. Crop the rural residential area remote sensing image into image blocks of the same size as the rural building dataset data through tile combination, record the positions in the original image, and generate the remote sensing image that can be predicted. In this embodiment, the overlap rate is set to 0.05.

[0095] Step 7: Input the image blocks in Step 6 into the rural building extraction model to obtain the rural building prediction results. All the prediction results are reassembled into a result image of the same size as the original remote sensing image according to the recorded position information to obtain the final building raster prediction result. Convert the building raster prediction result into a vector file and obtain the contour vector patches that conform to the actual boundaries of rural buildings.

[0096] Further, optimization processing can be performed on each building patch, including boundary optimization, removing redundant points, filling concave line segments, smoothing burrs, and performing right-angle edge optimization, etc., to make the building patch closer to the actual building form.

[0097] Step 8: Determine the rural building density level according to the rural building prediction results, assign the obtained rural building density level to each rural residential area, and obtain a vector file containing the distribution of rural residential areas and the building density attributes of each residential area.

[0098] In this embodiment, the rural building density is defined as five sparsity types: scattered, sparse, normal, dense, and extremely dense. Calculate the area of each rural residential area; record the number of buildings in each rural residential area and calculate the area of each building; calculate the ratio of the building area in each rural residential area to the residential area.

[0099] Classify the residential areas with a rural residential area area less than 2000 square meters, or the number of buildings less than 10, or the area ratio less than 0.1 as scattered sparsity.

[0100] Rural residential areas with an area proportion greater than or equal to 0.1 and less than 0.3 are classified as sparse density levels.

[0101] Rural residential areas with an area proportion greater than or equal to 0.3 and less than 0.55 are classified as normal density levels.

[0102] Rural residential areas with an area proportion greater than or equal to 0.55 and less than 0.75 are classified as dense density levels.

[0103] Rural residential areas with an area proportion greater than or equal to 0.75 are classified as extremely dense density levels.

[0104] Contents not described in detail in this specification belong to the prior art well-known to those skilled in the art. At the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manners and application scopes according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for extracting rural residential areas and determining building density using tile data, characterized in that: The steps include: Step 1: Preprocess remote sensing images: Preprocess the sub-meter satellite remote sensing images used, and set them to the corresponding resolutions of 14-level tiles and 18-level tiles, so as to obtain remote sensing images that can be used for two types of land features: rural residential areas and rural buildings; Step 2: Construct a rural residential area dataset and a rural building dataset for deep learning training; Step 3: Train deep learning extraction models, including rural residential area extraction models and rural building extraction models; Step 4: crawl the 14-level tile images in the prediction area, and cut the rural residential area remote sensing images into image blocks of the same size as the rural residential area dataset data through tile combination, record the positions in the original image, and generate remote sensing images that can be predicted; Step 5: Input the image block in step 4 into the rural residential area extraction model to obtain the rural residential area prediction result and obtain the contour vector patch that conforms to the actual boundary of the rural residential area; Step 6: crawl 18-level tile images in the prediction area according to the contour vector map of the actual boundary of the rural residential area in step 5, cut the rural residential area remote sensing image into image blocks of the same size as the rural building data set through tile combination, record the position in the original image, and generate a remote sensing image that can be predicted; Step 7: Input the image block in step 6 into the rural building extraction model to obtain the rural building prediction result and obtain the contour vector patch that conforms to the actual boundary of the rural building; Step 8: Determine the rural building density level based on the rural building prediction results, assign a value to each rural residential area to obtain the rural building density level, and obtain a vector file containing the distribution of rural residential areas and the building density attributes of each residential area.

2. The method for extracting rural residential areas and determining building density using tile data according to claim 1, characterized in that: The specific process of extracting the rural residential area model in step 3 is as follows: 3.1a. Input remote sensing image of rural residential areas Where H and W are the height and width of the image respectively, and 3 represents the RGB three channels; The input image is divided into non-overlapping patches, each of which is of size P×P. The image is then divided into multiple Each patch is flattened into a vector and the patch embedding is obtained by linear projection: Where C is the dimension of Patch Embedding, Flatten means flattening the Patch, and Linear means linear projection; 3.1b. Based on the SegFormer model structure, the Patch Embedding is passed through multiple stages of Transformer Block in the encoder. The resolution of the output feature map of each stage gradually decreases, and the number of channels gradually increases. The encoder can capture the feature information from low level to high level through the Transformer Block and generate multi-scale feature maps. 3.1c. The decoder consists of an MLP layer, an upsampling layer, and a 1x1 convolutional layer. It can fuse the multi-scale feature maps extracted by the encoder and output the final segmentation result Y: Among them, F fused is a multi-scale fusion feature map, N is the number of categories; Each pixel corresponds to a category probability distribution, and the predicted category of each pixel is obtained through the argmax operation: Prediction = argmax(Y), The trained rural residential area extraction model is used to predict the scope of rural residential areas.

3. The method for extracting rural residential areas and determining building density using tile data according to claim 2, characterized in that: The TransformerBlock is composed of the MHSA self-attention mechanism and the FFN feed-forward neural network. MHSA self-attention mechanism: Among them, Q, K, and V represent Query, Key, and Value respectively. k is the dimension of the Key key; FFN Feedforward Neural Network: FFN(x)=Linear(GELU(Linear(x))), GELU is the activation function, and Linear represents linear projection.

4. The method for extracting rural residential areas and determining building density using tile data according to claim 1, characterized in that: The specific process of extracting the rural building model in step 3 is as follows: 3.2a. Input remote sensing image of rural residential areas Where H and W are the height and width of the image respectively, 3 represents the RGB three channels, and the input image is normalized; 3.2b. Based on the DeepLabV3+ model structure, the encoder passes through the SENet154 backbone network to extract multi-level feature information from the input image. SENet154 outputs 6 feature maps of different scales, among which the feature map with a scale of 1 / 16 and a channel number of 2048 is used as the main feature and input into the ASPP module for information processing. 3.2c. In the decoder, the high-level features are fused with the low-level features. The low-level features are the low-level feature maps F with a scale of 1 / 4 and a number of channels of 256 extracted from the SENet154 backbone network. low ; The high-level feature is the feature map F output by the ASPP module ASPP Upsample to 1 / 4 scale; concatenate low-level features and high-level features and fuse them through convolutional layers: F fuced′ =Conv 3×3 (Concat(F low ,Upsample(F ASPP ))) The fused features are mapped to the category space through the upsampling layer and the 1×1 convolution layer, and the final segmentation result Z is output. Among them, F fused ′ is the fusion feature, N is the number of categories, and the predicted category of each pixel is obtained through the argmax operation: Prediction = argmax(Z), The trained rural building extraction model is used for rural building prediction.

5. The method for extracting rural residential areas and determining building density using tile data according to claim 4, characterized in that: The information processing in the ASPP module is specifically as follows: The ASPP module uses four parallel dilated convolution branches with different dilation rates of 1, 6, 12, and 18, and a global average pooling branch for multi-scale feature extraction. The dilated convolution formula is: Among them, r is the void rate, w(m,n) is the convolution kernel weight; The extracted multi-scale feature maps are concatenated together and channel adjusted through 1x1 convolution: F ASPP =Conv 1×1 (Concat(F1,F8,F 12 ,F 18 ,F GAP )) Among them, F1, F6, F 12 , F 18 is the feature map of different void ratios, F GAP is the global average pooling feature map.

6. The method for extracting rural residential areas and determining building density using tile data according to claim 1, characterized in that: In step 5, the raster extraction result of the residential area is converted into a vector result using the raster-to-vector function, the residential area portion is retained, and the background value portion is deleted; the extracted residential area result is morphologically optimized, the Douglas-Peucker algorithm is used to reduce the amount of point data, and the burr phenomenon in the contour is smoothed to obtain the final contour vector patch that conforms to the actual boundary of the rural residential area.

7. The method for extracting rural residential areas and determining building density using tile data according to claim 1, characterized in that: In step 7, each building outline vector image patch is optimized, including boundary optimization, removal of redundant points, filling of sunken line segments, smoothing of burrs, and optimization of right-angle edges, so that the building patch is closer to the actual building form.

8. The method for extracting rural residential areas and determining building density using tile data according to claim 1, characterized in that: In step 8, the rural building density is defined as five types of sparseness: scattered, sparse, normal, dense, and extremely dense, and the area of ​​each rural residential area is calculated; Record the number of buildings in each rural settlement and calculate the area of ​​each building; Calculate the ratio of building area to residential area in each rural residential area; Rural residential areas with an area of ​​less than 2,000 square meters, or with less than 10 buildings, or with an area ratio of less than 0.1, are classified as scattered density; Rural residential areas with an area ratio greater than or equal to 0.1 and less than 0.3 are classified as sparse density; Rural residential areas with an area ratio greater than or equal to 0.3 and less than 0.55 are classified as normal density; Rural residential areas with an area ratio greater than or equal to 0.55 and less than 0.75 are classified as densely depopulated; Rural residential areas with an area ratio greater than or equal to 0.75 are classified as extremely densely populated.

Citation Information

Patent Citations

  • Urban building automatic extraction method suitable for large-scale regional remote sensing image

    CN113505842A

  • Remote sensing image building extraction method based on attention mechanism and boundary loss

    CN114387521A

  • Human activity detection method and system based on deep learning

    CN117253155A

  • Rural homestead remote sensing intelligent detection method and system based on improved SegFormer network

    CN118351441A

  • High-resolution image-based ecological patch extraction method

    WO2024020744A1

Cited By

  • Virtual natural environment construction method, device, equipment, medium and program

    CN121810969A