A Remote Sensing Image Building Extraction Method Based on Deep Learning and Contour Regularization

By constructing a neural network with multi-scale parallel connection and regularization of contours, the remote sensing image building extraction results are optimized, and the problem of low extraction accuracy is solved, and high-precision and high generalization building extraction is achieved.

CN115984693BActive Publication Date: 2025-07-22CHANGGUANG SATELLITE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211693340.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2025-07-22
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

In the extraction of remote sensing image building, the extraction results are quite different from the actual situation and the optimization is not ideal, resulting in low extraction accuracy and cannot meet the needs of large areas and high frequency updates.

Method used

Using a method based on deep learning and contour regularization, the architectural extraction results are optimized by constructing a neural network and attention mechanism with multi-scale parallel connections, combined with morphological contour regularization processing.

Benefits of technology

Under different image scales, building distribution vectors with regular outlines and strong application can be obtained, which improves extraction accuracy and generalization, and the result is closer to the actual edge of the building.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984693B_ABST
    Figure CN115984693B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for extracting buildings from remote sensing images based on deep learning and contour regularization, belonging to the field of application of optical remote sensing technology, and includes the steps of: artificially constructing a building training sample data set; constructing and training a building extraction model; inputting the remote sensing image of the area to be extracted into the building extraction model to obtain a corresponding initial building distribution grid result; performing morphological contour regularization processing on the initial building distribution grid result to obtain a building distribution vector file including building contour vector patches. The method for extracting buildings from remote sensing images based on deep learning and contour regularization proposed by the present invention can realize obtaining building distribution vectors with regular contours and strong applicability at different image scales, the model has stronger generalization ability, the extracted building contours are closer to the actual edges of the buildings, and the extraction accuracy is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of optical remote sensing technology applications, and specifically relates to a method for extracting buildings from remote sensing images based on deep learning and contour regularization. Background Art

[0002] Building extraction is an important part of work such as extracting human activities in nature reserves, delimiting ecological protection red lines, and non-agriculturalization of cultivated land. Traditional building extraction methods mostly use manual on-site measurement or manual delineation based on drone images, which rely heavily on the number of staff and are inefficient, and cannot meet the requirements of large-area and high-frequency updates of building information in many scenarios such as urban construction development and national land space planning.

[0003] With the continuous improvement of image resolution, the features and contour information of buildings in images are becoming more and more obvious, and satellite remote sensing data is commonly used in building analysis work. The development of deep learning technology provides a technical basis for automatic building extraction. Convolutional neural networks can efficiently extract and learn the features of ground objects in remote sensing images and obtain an automatic building extraction model through training. However, the building contours extracted by such automatic methods are quite different from the actual ones, often showing over-segmentation or under-segmentation phenomena. To meet the practical needs, the results need to be optimized. However, the optimization results are not ideal, and there are still large differences between the optimized building contours and the actual building contour edges. Summary of the Invention

[0004] Aiming at the problem that the building contours extracted by the existing remote sensing image building extraction method based on deep learning technology are quite different from the actual ones, and the optimization of the extraction results is not ideal, resulting in low extraction accuracy, the present invention provides a method for extracting buildings from remote sensing images based on deep learning and contour regularization.

[0005] To solve the above problems, the present invention adopts the following technical solutions:

[0006] A method for extracting buildings from remote sensing images based on deep learning and contour regularization, comprising the following steps:

[0007] Step 1: Manually construct a building training sample data set;

[0008] Step 2: Construct and train a building extraction model, including the following steps:

[0009] Step 21: Construct a neural network with multi-scale parallel connections, and the neural network with multi-scale parallel connections is used to parallelly extract feature maps of four different scales and perform feature cross-fusion to obtain an initial network fusion feature;

[0010] Step 22: Calculate the class probability map and class features respectively according to the initial network fusion feature;

[0011] Step 23: Calculate the regional context features by using the conversion method of the attention mechanism;

[0012] Step 24: Fuse the initial network fusion features and the regional context features to obtain the segmentation features;

[0013] Step 25: The sample images in the building training sample dataset pass through two 3×3 convolutions, two BN layers, and the ReLU activation function, and the initial feature output is a feature map of the original image size The obtained initial feature map is input into the multi-scale parallel-connected neural network. The cross-entropy loss function is used as the loss function, and the stochastic gradient descent algorithm is used for optimization training to obtain the building extraction model after training;

[0014] Step 3: Input the remote sensing image of the area to be extracted into the building extraction model to obtain the corresponding initial building distribution grid result;

[0015] Step 4: Perform morphological contour regularization processing on the initial building distribution grid result. The specific contour regularization processing process is as follows:

[0016] Step 41: Perform opening operation on the initial building distribution grid result;

[0017] Step 42: Based on the morphological forms of individual buildings obtained after the opening operation, extract the initial building contour information and record it in the form of a point set;

[0018] Step 43: Use the Douglas-Peucker algorithm to optimize the point set to obtain the contour L of each building after preliminary simplification;

[0019] Step 44: Determine the minimum bounding rectangle R corresponding to the contour L, and calculate the ratio of the area of the contour L to the area of the minimum bounding rectangle R. If the ratio is greater than the threshold, use the contour of the minimum bounding rectangle R as the final building contour, and then execute Step 46; if the ratio is less than or equal to the threshold, execute Step 45;

[0020] Step 45: Fill the depressions of the contour L and optimize the right-angled sides;

[0021] Step 46: Generate a building distribution vector file including the building contour vector patches.

[0022] The beneficial effects of the present invention are:

[0023] The present invention proposes a method for extracting buildings from remote sensing images based on deep learning and contour regularization, which can obtain building distribution vectors with regular contours and strong applicability at different image scales. A neural network structure with multi-scale parallel connection is adopted to ensure the effective utilization of building features at different resolutions. Context information is associated during the process to improve the accuracy of building extraction. The generalization of the model under multi-resolution images is improved by using multi-scale training and multi-scale prediction. The contour regularization method is used to optimize the building segmentation results so that the contours are closer to the actual edges of the buildings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of a method for extracting buildings from remote sensing images based on deep learning and contour regularization according to the present invention;

[0025] Figure 2 is a schematic diagram of a sample in an artificially constructed building training sample dataset;

[0026] Figure 3 is a schematic diagram of the network structure of a neural network with multi-scale parallel connection;

[0027] Figure 4 is an extraction effect diagram of remote sensing images of three different groups of different regions and different building types;

[0028] Figure 5 is an extraction effect diagram of large-area buildings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The method for extracting buildings from remote sensing images proposed by the present invention combines a convolutional neural network with high generalization and morphological contour regularization. An expert knowledge-based sample set suitable for deep learning training is artificially constructed. In the parallel multi-scale feature extraction neural network, the main structure is the parallel extraction of four feature maps of different scales and the cross-fusion of features. The preliminary building extraction results of high-resolution remote sensing images are obtained through multi-scale training and prediction. On this basis, morphological contour optimization is performed to obtain the regularized building vectors. The technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0030] Figure 1 As shown, it is a flowchart of a method for extracting buildings from remote sensing images based on deep learning and contour regularization proposed by the present invention, which specifically includes the following steps:

[0031] Step 1: Artificially construct a building training sample dataset

[0032] Step 11: Obtain a number of high-resolution satellite remote sensing images with a resolution better than 0.8 meters.

[0033] To make the building extraction model more applicable to high-resolution remote sensing images and more accurate in the process of manually delineating building contours, high-resolution satellite remote sensing images with a resolution better than 0.8 meters are selected to construct the sample set. In image selection, areas with a relatively complete distribution of building types are preferentially selected for delineation, such as Figure 2 (a) shows the selected true-color image, ensuring that basic building types such as residential buildings, factories, stadiums, commercial buildings, and rural buildings are included in the sample set.

[0034] Step 12: Manually delineate the building contours for each high-resolution satellite remote sensing image using expert knowledge. After delineation, generate building binary labels.

[0035] During the delineation process, the lines fit the outer contour edges of the building, satisfying the contour lines being simple and covering individual buildings to the greatest extent, such as Figure 2 (b) shows the result of the manually delineated vector. After delineation, generate building binary labels, such as Figure 2 (c) shows. In the figure, black is the background and white is the building.

[0036] Step 13: Divide the high-resolution satellite remote sensing image and its corresponding building binary label into blocks to obtain the initial building sample set.

[0037] This step divides the image data and annotation data (i.e., building binary labels) into blocks to obtain the initial building sample set.

[0038] Step 14: Perform data augmentation on the initial building sample set to obtain the building training sample data set.

[0039] This step performs data augmentation on the images and annotation files, including any random combination of methods such as horizontal flipping, vertical flipping, random rotation, and affine transformation. The augmented data set is used for deep learning network training.

[0040] In particular, when manually constructing the building training sample data set, the construction process follows the following principles: (1) Include various building types, and ensure that the sample quantities of each type of building are relatively balanced during the sample delineation process; (2) The minimum building delineation area is maintained at 30 square meters or more than 60 pixels; (3) During the delineation process, delineate the roof contour of the building to maintain a regular shape.

[0041] Step 2: Construct and train the building extraction model

[0042] Construct a neural network with multi-scale parallel connections, including a multi-scale feature extraction backbone network, and perform feature fusion in each level of feature extraction. Obtain a building extraction model with strong generalization ability using the multi-scale training method. Use the building extraction model obtained through training to extract the remote sensing images within the range to be processed by means of multi-scale prediction, and obtain the initial building distribution grid result.

[0043] Constructing and training the building extraction model specifically includes the following steps:

[0044] Step 21: Construct a neural network with multi-scale parallel connections, which is used to parallelly extract feature maps of four different scales and perform feature cross-fusion to obtain the initial fusion features of the network.

[0045] The neural network with multi-scale parallel connections includes a parallel convolution stream composed of parallel feature extraction networks of four scales. As Figure 3 shown, after the initial feature map corresponding to the sample image in the building training sample dataset is input into this neural network with multi-scale parallel connections, it will pass through a parallel convolution stream composed of parallel feature extraction networks of four scales. Its feature is to perform feature fusion on the pre-convolution features and the features after convolution with a series of different scales to ensure the richness of the feature levels. Taking high-resolution remote sensing satellite images as an example, in the network, the size of the input sample image is 650×650×3 (i.e., image length × image width × number of bands).

[0046] The sample image first passes through two 3×3 convolutions, two BN (Batch Normalization) layers, and the ReLU activation function, and the initial feature output is a feature map of the original image size 162×162×64. Next, the initial feature map passes through four convolution stream branches and performs low-resolution and high-resolution fusion. The specific process is as follows:

[0047] In the first convolution stream branch, the initial feature map passes through four Bottleneck Block residual blocks, and the output feature map is S 1 : 162×162×256; the size of the feature map is

[0048] In the second convolution stream branch, scale transformation is performed on the feature map S 1 obtained from the first convolution stream branch, the dimensionality of the channel features is reduced and downsampled to expand and generate: For and S b 2 perform feature fusion of different resolutions, and output a feature map of the size of the original image ;

[0049] In the third convolutional stream branch, the feature map obtained from the second convolutional stream branch is expanded in parallel to obtain the feature map of the third scale For feature fusion with different resolutions is performed, and the feature map of the size of the original image map is output ;

[0050] In the fourth convolutional stream branch, the feature map obtained from the third convolutional stream branch is expanded in parallel to obtain the feature map of the fourth scale For feature fusion with different resolutions is performed, and the feature map of the size of the original image map is output ;

[0051] The feature maps obtained from each convolutional stream branch are uniformly sampled to 162×162×720 and fused to obtain the initial fusion feature F1 of the network.

[0052] Step 22: Calculate the class probability map P m and the class feature F c .

[0053] The present invention considers the features of the surrounding neighborhood pixels of the pixel to be classified for context feature fusion. First, the class probability map P of the image is calculated by the following formula (1) through the feature of the initial fusion feature F1 of the network m .

[0054] P m = softmax(F1)(1)

[0055] Then, according to formula (2), the transpose of the initial fusion feature F1 of the network is multiplied by the class probability map P m to obtain the class feature F c .

[0056] F C = F1 T × P m (2)

[0057] Step 23: Calculate the regional context feature by using the method of attention mechanism transformation.

[0058] In this step, the pixel feature and the class feature are calculated by using the method of attention mechanism transformation, and the calculation method of the regional context feature F q is shown in the following formula (3):

[0059]

[0060] Among them, C is the number of categories in the given image (in the present invention, C = 2, and the categories are buildings and non-buildings), w ij represents the similarity measure between the pixel p i and the feature Fc of category j j , and δ(·) and ρ(·) are two feature transformation equations of the attention mechanism.

[0061] Step 24: Fuse the network initial fusion feature and the regional context feature to obtain the segmentation feature.

[0062] Finally, according to the obtained regional context feature F q , fuse the network initial fusion feature F1 and the regional context feature F q . As shown in formula (4), assuming the fusion weight λ = 0.4, the finally obtained segmentation feature F2 has a size of 162×162×512.

[0063] F2 = λ·F1+(1 - λ)·F q (4)

[0064] where λ is the fusion weight.

[0065] Step 25: Input the initial feature map generated from the sample images in the building training sample dataset into the neural network with multi-scale parallel connections, use the cross-entropy loss function as the loss function, and adopt the stochastic gradient descent algorithm for optimization training to obtain the building extraction model after training.

[0066] The overall loss function of the network is selected to use the cross-entropy loss function, as shown in formula (5):

[0067]

[0068] where C represents the number of categories, y i is the true label value, is the predicted value.

[0069] Adopt the stochastic gradient descent (SGD) algorithm for optimization training, set the learning rate to decrease linearly, and obtain the building extraction model after training.

[0070] Train the building extraction model in a multi-scale manner. When the sample images are input into the network, randomly select from the scale set and perform scale transformation based on this scale. The scale set is [0.5, 0.75, 1, 1.25, 1.5, 1.75, 2], which contains seven different scales. Using this method can maximize the generalization of the model.

[0071] During the prediction process, a multi-scale approach is also adopted for prediction. During the prediction process, for the input image, predictions are made separately with scale sets of [0.5, 0.75, 1, 1.25, 1.5, 1.75, 2]. The intersection of the building extraction results at different scales is taken to obtain the final initial building distribution raster result.

[0072] Step 3: Input the remote sensing image of the area to be extracted into the building extraction model to obtain the corresponding initial building distribution raster result.

[0073] Step 4: Regularize the building outline

[0074] The initial building distribution raster result obtained by extraction is optimized morphologically. The boundary is extracted based on the general shape of each building. The Douglas–Peucker algorithm is used to reduce the amount of point data, and the burr phenomenon in the contour is smoothed, and the approximate right-angled sides are regularized to obtain the final contour vector patch that conforms to the actual boundary of the building.

[0075] In this step, the specific process of morphologically regularizing the contour of the initial building distribution raster result is as follows:

[0076] Step 41: Perform an opening operation on the initial building distribution raster result. For the obtained initial building distribution raster result, perform an opening operation with a scale of 3 to remove isolated small dots, burrs, and small bridges while ensuring the position and shape remain unchanged.

[0077] Step 42: Based on the morphology of each individual building obtained after the opening operation, extract the initial building contour information and record it in the form of a point set P.

[0078] Step 43: Use the Douglas–Peucker algorithm to optimize the point set P, remove redundant points, smooth the burr line segments, and ensure that the shape of the trajectory curve remains roughly unchanged to obtain the contour L of each building after preliminary simplification;

[0079] Step 44: Determine the minimum bounding rectangle R corresponding to the contour L and calculate the ratio of the area of the contour L to the area of the minimum bounding rectangle R. If the ratio is greater than the threshold (e.g., 0.8), then use the contour of the minimum bounding rectangle R as the final building contour, and then execute Step 46; otherwise, if the ratio is less than or equal to the threshold, consider the building to have a complex shape and further optimization is required, and execute Step 45;

[0080] Step 45: Fill the depressions in the contour L and optimize the right-angled sides;

[0081] When filling the concavities of the contour L, calculate the angles between all pairs of adjacent sides in the contour L. If the formed triangle is a concave triangle, calculate the area of the triangle. If the area of the triangle is less than 5% of the area of the entire contour, then fill the triangle.

[0082] When optimizing the right-angled sides of the contour L, it includes two cases:

[0083] (1) For the case where the included angle between two sides is approximately 90 degrees, adjust the position of the middle point so that the two sides form a right-angled shape;

[0084] (2) Calculate the included angle between two sides separated by one side. If the included angle is approximately 90 degrees, calculate the area of the triangle formed by the intersection point of the extensions of the two sides and the two endpoints of the middle side. If the area of the triangle is less than 5% of the area of the entire contour, then delete the two endpoints of the middle side from the point set, and then insert the intersection point of the extensions of the two sides into the point set.

[0085] Step 46: Generate a building distribution vector file including building contour vector patches. Based on the obtained point set P1, generate a building distribution vector file. The contours of each building contour vector patch in the building distribution vector file are optimized, and each side presents a regularized characteristic with right angles as the main body. The finally obtained building contour vector patches are closest to the actual building edges, and the data volume is the smallest, improving the practical applicability of the remote sensing image building extraction method.

[0086] The present invention extracts an initial building raster result in the area to be predicted by means of deep learning semantic segmentation, and then uses a morphological-based method to regularize the result to obtain vector patches that fit the actual building edges. An artificial drawing is used to create a building sample set, and the prediction accuracy is improved by means of multi-scale training and multi-scale prediction. In sub-meter high-resolution images, the finally optimized vector patches have high practicality.

[0087] The technical effects of the present invention will be described below in conjunction with specific examples.

[0088] The present invention selects domestic cities such as Shanghai, Wuhan, Changchun, and Harbin, and foreign cities such as Las Vegas and Paris to make a training sample set. The resolution of the satellite remote sensing image is better than 0.8 meters. During the model training process, the data is as follows:

[0089] Table 1 Building Extraction Model Training Parameter Table

[0090] Sample quantity Training / validation quantity Training time 8800 6100 / 2700 36 hours

[0091] As shown in Table 1, a total of 8800 images are constructed as samples and trained. During the testing process, the average prediction time for each test sample is about 1.5 seconds.

[0092] To accurately verify the practical applicability of the present invention, 10 sample plots with a size of 2 km × 2 km were outlined in 7 cities including Changchun, Harbin, Shenyang, Baishan, Jinzhou, Siping, and Songyuan. The total area was 40 km2 and was used for the accuracy verification of building extraction. The verification accuracy data is shown in the following table:

[0093] Table 2 Accuracy Verification of Remote Sensing Image Building Extraction Based on Deep Learning and Contour Regularization

[0094] Original model extraction accuracy Regularized extraction accuracy Overall accuracy 90% 93%

[0095] As shown in Table 2, the overall accuracy of buildings in all sample plots reached 93%, indicating a good extraction effect. Among them, the overall extraction accuracy of the extraction result after regularization increased by 3% compared with the extraction result of the original model that only used deep learning, indicating that the contour optimization technology provided by the present invention has a good effect on improving the extraction accuracy.

[0096] As Figure 4 shown, three groups of remote sensing images with different regions and different building types were selected for the display of extraction effects. Among them, (a1), (a2), and (a3) are true color images; (b1), (b2), and (b3) are ground truth labels; (c1), (c2), and (c3) are the extraction effects of only using the deep learning method; (d1), (d2), and (d3) are the extraction effects of the present invention. The black in the figure is the background and the white is the building. Compared with the ground truth labels, the present invention can extract more accurate building results. Compared with the method that only uses deep learning, the present invention can obtain results that are closer to the actual building contours.

[0097] Figure 5 Figure is the rendering of large-area building extraction, where (a) is the true color image; (b) is the extraction result of the present invention. It can be seen from the results that the present invention can obtain a relatively accurate building distribution, and good extraction results can be obtained for various building types, meeting the requirements of actual use.

[0098] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0099] The above-described embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent should be subject to the appended claims.

Claims

1. A method for extracting buildings from remote sensing images based on deep learning and contour regularization, characterized in that, It includes the following steps: Step 1: Manually construct a building training sample dataset; Step 2: Construct and train a building extraction model, including the following steps: Step 21: Construct a neural network with multi-scale parallel connections, which is used to parallelly extract feature maps of four different scales and perform feature cross-fusion to obtain an initial network fusion feature; Step 22: Calculate the class probability map and class features respectively based on the initial network fusion feature; The calculation formula of the class probability map is: P m = softmax(F1) (1) Among them, P m is the class probability map, and F1 is the initial fusion feature of the network; The calculation formula of the class features is: F C = F1 T × P m (2) Among them, F c is the category feature, and F1 T is the transpose of the initial network fusion feature F1; Step 23: Calculate the regional context feature by means of the attention mechanism transformation; Step 24: Fuse the initial network fusion feature and the regional context feature to obtain a segmentation feature; The calculation formula of the segmentation feature is: F2 = λ·F1+(1 - λ)·F q (4) Among them, F2 is the segmentation feature, F q is the regional context feature, and λ is the fusion weight; Step 25: The sample images in the building training sample dataset pass through two 3×3 convolutions, two BN layers, and the ReLU activation function, and the initial feature output is a feature map of the original image size sized feature map. The obtained initial feature map is input into the multi-scale parallel-connected neural network, the cross-entropy loss function is used as the loss function, and the stochastic gradient descent algorithm is used for optimization training to obtain a building extraction model after training; Step 3: Input the remote sensing image of the area to be extracted into the building extraction model to obtain the corresponding initial building distribution grid result; Step 4: Perform morphological contour regularization processing on the initial building distribution grid result. The specific contour regularization processing process is as follows: Step 41: Perform an opening operation on the initial building distribution grid result; Step 42: Based on the morphology of each single building obtained after the opening operation, extract the initial building contour information and record it in the form of a point set; Step 43: Use the Douglas-Peucker algorithm to optimize the point set to obtain the contour L of each building after preliminary simplification; Step 44: Determine the minimum bounding rectangle R corresponding to the contour L, and calculate the ratio of the area of the contour L to the area of the minimum bounding rectangle R. If the ratio is greater than the threshold, use the contour of the minimum bounding rectangle R as the final building contour, and then execute Step 46; if the ratio is less than or equal to the threshold, execute Step 45; Step 45: Fill the depressions in the contour L and optimize the right-angle sides; When filling the depressions in the contour L, calculate the included angle of all two sides in the contour L. If the formed triangle is an inner concave triangle, calculate the area of the triangle. If the area of the triangle is less than 5% of the area of the entire contour, fill the triangle; When optimizing the right-angle sides of the contour L, it includes two cases: (1) For the case where the included angle between two sides is approximately 90 degrees, adjust the position of the middle point so that the two sides are in a right-angle shape; (2) Calculate the included angle between two sides separated by one side. If the included angle is approximately 90 degrees, calculate the area of the triangle formed by the intersection point of the extended lines of the two sides and the two endpoints of the middle side. If the area of the triangle is less than 5% of the area of the entire contour, delete the two endpoints of the middle side from the point set, and then insert the intersection point into the point set; Step 46: Generate a building distribution vector file including building contour vector patches.

2. A method for extracting buildings from remote sensing images based on deep learning and contour regularization according to claim 1, characterized in that The multi-scale parallel-connected neural network includes a parallel convolutional stream composed of parallel feature extraction networks of four scales. After the sample image is input into the multi-scale parallel-connected neural network, it first passes through two 3×3 convolutions, two BN layers, and a ReLU activation function, and the initial feature output is a feature map of the original image size with a size of 162×162×64. Next, it passes through four convolutional stream branches, and the specific process is as follows: In the first convolutional stream branch, the initial feature map passes through four Bottleneck Block residual blocks, and the output feature map is S 1 : 162×162×256; the size of the feature map is In the second convolutional stream branch, the feature map S obtained from the first convolutional stream branch 1 is subjected to scale transformation to reduce the dimensionality of the channel features and perform downsampling, thereby expanding and generating: 162×162×48; 81×81×96; For and feature fusion at different resolutions is performed, and a feature map of the size of the original image map is output ; In the third convolutional stream branch, the feature map obtained from the second convolutional stream branch is expanded in parallel to obtain the feature map of the third scale 40×40×192; for feature fusion at different resolutions is performed, and the feature map of the size of the original image map is output; In the fourth convolutional stream branch, the feature map obtained from the third convolutional stream branch is expanded in parallel to obtain the feature map of the fourth scale 28×28×384; for feature fusion at different resolutions is performed, and the feature map of the size of the original image is output; The feature maps obtained from each convolutional stream branch are uniformly sampled to 162×162×720 and fused to obtain the initial network fusion feature F1.

3. A method for extracting buildings from remote sensing images based on deep learning and contour regularization according to claim 1, characterized in that, Step 1 includes the following steps: Step 11: Obtain several high-resolution satellite remote sensing images with a resolution better than 0.8 meters; Step 12: Manually draw the building contours on each of the high-resolution satellite remote sensing images using expert knowledge. After the drawing is completed, generate a building binary label; Step 13: Divide the high-resolution satellite remote sensing images and their corresponding building binary labels into blocks to obtain an initial building sample set; Step 14: Perform data augmentation processing on the initial building sample set to obtain a building training sample data set.

4. A method for extracting buildings from remote sensing images based on deep learning and contour regularization according to claim 3, characterized in that, When manually constructing a building training sample data set, the following principles are followed: (1) Include multiple building types, and ensure that the sample quantities of each type of building are relatively balanced during the process of delineating samples; (2) The minimum building delineation area is maintained at 30 square meters or more than 60 pixels; (3) During the delineation process, delineate the roof outline of the building and maintain a regular shape.

5. A method for extracting buildings from remote sensing images based on deep learning and contour regularization according to claim 3 or 4, characterized in that The operations of data augmentation processing include any random combination of horizontal flipping, vertical flipping, random rotation, and affine transformation.

6. The method for extracting buildings from remote sensing images based on deep learning and contour regularization according to claim 1, characterized in that, The threshold is 0.8.

Citation Information

Patent Citations

  • Image segmentation method of lung tissue on CT chest radiography based on level set

    CN109064476A

  • Contour detection method under complex background, terminal device and storage medium

    CN110570442A