Building contour extraction system

By integrating SeResNeXt and Unet++ networks, the building outline extraction system solves the problems of data dependence and high computational resources in building outline extraction from remote sensing images, achieving efficient and accurate building identification and extraction, and providing a user-friendly interface.

CN120976729APending Publication Date: 2025-11-18SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510890435.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for extracting building outlines from remote sensing images suffer from several drawbacks, including insufficient data dependence and generalization ability, misidentification, high computational resource requirements, sensitivity to image quality, and poor detection performance for small targets and irregularly shaped buildings.

Method used

A building contour extraction system integrating SeResNeXt and Unet++ networks is adopted. Combining data preprocessing, building recognition, contour extraction and graphical user interface modules, it achieves accurate building recognition and extraction through data augmentation, CBAM attention module and morphological filtering.

Benefits of technology

It improves the accuracy and efficiency of building outline extraction, reduces computational costs, enhances the ability to handle occluded buildings, and provides user interaction functions, simplifying the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976729A_ABST
    Figure CN120976729A_ABST
Patent Text Reader

Abstract

The invention discloses a building contour extraction system, and the system comprises a data preprocessing module which carries out the normalization, enhancement and size unification of an image; the building identification module adopts an improved Unet + + network, combines SeResNeXt as a backbone network to extract multi-scale features, embeds a CBAM attention module, and outputs a binary mask image; the contour extraction module is used for simplifying a contour through morphological filtering and denoising, contour detection and an adaptive Douglas-Peucker algorithm, and generating a geographic coordinate file through affine transformation; and the graphical interactive interface GUI module supports automatic processing and manual annotation. According to the method, the building recognition robustness is improved through multi-scale feature fusion, the CBAM attention module accelerates convergence, manual intervention is reduced through adaptive contour simplification, a shielding target is effectively processed, and the precision and efficiency of applications such as urban planning and disaster management are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image building contour extraction, and particularly to a building contour extraction system. BACKGROUND

[0002] Automatic building contour extraction from remote sensing images is a key and complex technical challenge in the field of remote sensing. The goal is to accurately identify and delineate building contours from high-resolution images. This technology is crucial for applications such as urban planning, disaster management, and environmental monitoring. Accurate building contour data can significantly improve the efficiency and effectiveness of urban infrastructure planning, traffic navigation, and economic development.

[0003] However, building contour extraction faces severe challenges. Buildings vary greatly in type, size, shape, and architectural style, and there are widespread occlusions, shadows, lighting changes, and background clutter in images, making it extremely difficult to maintain robustness, accuracy, and identify and delineate buildings. To address these challenges, researchers have developed various methods:

[0004] 1. Pixel-based methods: Directly operate at the pixel level, using threshold segmentation, edge detection, mathematical morphology, etc. This method is susceptible to noise, complex scenes, and building appearance variations, with limited accuracy.

[0005] 2. Object-based methods: Consider the spatial relationship between pixels, first obtain candidate regions through image segmentation, and then use shape, texture, context, etc. features to extract building contours with the help of machine learning algorithms. This method is an improvement over pixel-based methods, but still has difficulty completely overcoming the problems caused by complex scenes.

[0006] Traditional image processing methods often struggle when dealing with lighting fluctuations and irregularly shaped structures, resulting in inaccurate extraction results. In recent years, deep learning methods represented by convolutional neural networks (CNN) have shown greater potential and can better learn complex features. However, existing methods still have significant limitations:

[0007] 1. Data dependency and generalization: Model performance is limited by the diversity and size of the training data set, and the generalization ability for unseen scenes is insufficient.

[0008] 2. False identification: It is difficult to avoid false positives, i.e. identifying non-buildings as buildings and false negatives, i.e. missing buildings.

[0009] 3. Small targets and irregular shapes: Small buildings and irregularly shaped buildings such as shantytowns and industrial buildings are not well detected.

[0010] 4. Computing resource requirements: Complex deep learning models require high computational costs for training and inference.

[0011] 5. Sensitivity to image quality: factors such as image resolution, noise, cloud cover, etc. directly affect the extraction accuracy.

[0012] Therefore, although after years of research, to develop a more robust, scalable, efficient and as far as possible to overcome the above challenges of building contour extraction method to meet the actual application of high-precision, automated solutions to the needs of the field, there is still a research gap to be filled. SUMMARY

[0013] The purpose of the present application is to overcome the shortcomings and deficiencies of the prior art, and to provide a building contour extraction system that integrates the advantages of SeResNeXt and Unet++ networks to achieve accurate identification of buildings and improve the accuracy of automatic building contour extraction.

[0014] To achieve the above purpose, the technical scheme provided by the present application is as follows: a building contour extraction system, comprising:

[0015] A data preprocessing module is used to load and preprocess images, divide the preprocessed images into a training set and a test set, and perform data augmentation on the training set.

[0016] A building recognition module is used to input the augmented training set into an improved Unet++ network for training, and after training, input the test set into the trained improved Unet++ network for recognition. The improved Unet++ network outputs a binary mask image for each input image. The improved Unet++ network uses a SeResNeXt network as a backbone network and removes the classification layer at the end of the SeResNeXt network. For input image data, the backbone network generates different scale feature maps s1, s2, s3, s4 from shallow to deep, which are input into the improved Unet++ network in sequence. 1,1 The node is represented by X 4,1 The node is represented by X i,j where X represents the i-th layer and the j-th node in the improved Unet++ network. In addition, CBAM attention modules are added to the skip connections of the first node and the last node of each layer in the improved Unet++ network.

[0017] A contour extraction module is used to process the binary mask image output by the building recognition module: first, remove noise through morphological filtering; then use a contour detection algorithm to extract the building contour from the denoised image; then use a regression algorithm to simplify the extracted building contour; finally, convert the simplified contour into a standard geographic coordinate format file, called a building contour coordinate format file.

[0018] A graphical user interface (GUI) module is configured to provide user interaction functions, including: a first function configured to receive a user input image, invoke the contour extraction module to process the image, and output a corresponding building contour coordinate format file; and a second function configured to provide a user manual marking tool, so that the user can mark a building contour on the image and save the marking result.

[0019] Further, the data preprocessing module performs the following steps:

[0020] 1) Normalizing the input image, dividing each pixel value in the image by 255, converting all pixels to floating-point numbers between 0 and 1, and then dividing the image into a training set and a test set, wherein each image in the training set has a corresponding binary mask image;

[0021] 2) Data augmentation is performed on the training set obtained in step 1), specifically, the same random horizontal flip, random rotation, and random cropping are performed on each image in the training set and its corresponding binary mask image to obtain an augmented training set;

[0022] 3) Adjusting the image size of the test set in step 1) and the augmented training set in step 2), using bilinear interpolation to adjust all images to 512*512 pixels in size.

[0023] Further, the improved Unet++ network is specifically introduced as follows:

[0024] ① Removing the global average pooling layer and the fully connected layer at the end of the SeResNeXt network, retaining the feature extraction part as the backbone network; for the input image data, the backbone network outputs different scale feature maps s1, s2, s3, s4 from shallow to deep;

[0025] ② The Unet++ network is a major improved version of the classic U-Net architecture, and the core innovation lies in the introduction of densely nested skip connections, a total of 4 layers, of which the i-th layer has 5-i nodes, and the j-th node of the i-th layer is denoted as X i,j The densely nested skip connections of the node need to come from different stages of the encoder, i.e., different depth and different resolution feature maps, so the output feature maps of certain stages in the backbone network, i.e., the above-mentioned different scale feature maps s1, s2, s3, s4, are intercepted as the input of the first node X i,1 of each layer of the improved Unet++ network, for each node of the i-th layer of the improved Unet++ network, it receives the up-sampling result from the node of the i+1-th layer, receives the feature map from the corresponding level of the backbone network, and receives the feature map of the interpolated resolution from the previous node of the improved Unet++ network, and splices them together to obtain the initial feature map of the current node, and then performs convolution operation to fuse the information;

[0026] ③In the improved Unet++ network, for the first node and the last node of each layer, the feature maps received from the backbone network through the skip connection need to be processed by a CBAM attention module, which is used to weight the spatial information and channel information of the features.

[0027] Further, the contour extraction module performs the following steps:

[0028] 1) Perform morphological filtering on the binary mask image output by the building recognition module: first perform a closing operation to fill small holes inside the building, and then perform an opening operation to remove small noise points; wherein the kernel size of the morphological operation is 3x3;

[0029] 2) Using the Suzuki-Abe algorithm to obtain pixel-based building contour lines from the binary mask image processed in step 1), filter out the contour lines that are not within the preset range and contact the image edge, and for the case where multiple contour lines are detected for the same building, calculate the minimum bounding rectangle of each contour line, if the positions of the minimum bounding rectangles of two contour lines are similar, it is determined that they belong to the same building, and the contour lines are merged accordingly;

[0030] 3) Perform polygon simplification on the contour lines processed in step 2) to obtain simplified contour line coordinates, wherein the Douglas-Peucker algorithm is used for simplification, select the points Pmin and Pmax with the smallest and largest horizontal coordinates on the contour line, divide the contour line into two halves, and perform the Douglas-Peucker algorithm on each half, the tolerance parameter of the algorithm is selected automatically according to the perimeter of the contour line;

[0031] 4) Convert the simplified contour line coordinates obtained in step 3) from the image pixel coordinate system to the geographic coordinate system through affine transformation, and save it as a standard geographic coordinate format file, called a building contour coordinate format file.

[0032] Further, the functions of the graphical user interface GUI module include:

[0033] First function: load and display the image specified by the user; in response to user operation, call the contour extraction module to process the image, automatically extract the building contour and display the result, and output the corresponding building contour coordinate format file;

[0034] Second function: load and display the image specified by the user; provide a drawing tool to enable the user to manually draw and edit the building contour polygon on the image; save the contour polygon data drawn by the user.

[0035] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0036] 1、The network architecture finally obtained has multiple levels of characteristics and can process buildings of different sizes, and is more reliable than the original Unet++ network.

[0037] 2、The neural network used in the present application has a faster training convergence speed when training, and converges better under the same training batch.

[0038] 3、The present application proposes an adaptive contour extraction method when performing contour line extraction, which reduces the use threshold of the user.

[0039] 4、According to experimental verification, the method of the present application is simpler, more efficient, faster in operation, and has a certain restoration effect on the occluded building compared with the traditional method. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 It is a logical flow diagram of the method of the present application.

[0041] Figure 2 It is an overall architecture diagram of the improved Unet++ network.

[0042] Figure 3 It is an instance schematic diagram of the channel attention mechanism of the CBAM attention module.

[0043] Figure 4 It is an instance schematic diagram of the spatial attention mechanism of the CBAM attention module.

[0044] Figure 5 It is one of the instance effect schematic diagrams of the building recognition module.

[0045] Figure 6 It is the second instance effect schematic diagram of the building recognition module.

[0046] Figure 7 It is an instance effect schematic diagram of the contour extraction module.

[0047] Figure 8 It is a frame schematic diagram of the graphical interactive interface GUI module. DETAILED DESCRIPTION

[0048] The present application will be further described in detail below in conjunction with the embodiments and drawings, but the embodiments of the present application are not limited thereto.

[0049] The present embodiment discloses a building contour extraction system, comprising:

[0050] The data preprocessing module is configured to load and preprocess images, divide the preprocessed images into a training set and a test set, and perform data enhancement on the training set.

[0051] The building recognition module is configured to input the enhanced training set into the improved Unet++ network for training, and input the test set into the trained improved Unet++ network for recognition after the training is completed. The improved Unet++ network outputs a binary mask image for each input image. The improved Unet++ network uses a SeResNeXt network as a backbone network and removes the classification layer at the end of the SeResNeXt network. The backbone network generates different scale feature maps s1, s2, s3, and s4 from shallow to deep, and inputs the feature maps into the improved Unet++ network in sequence. 1,1 The node is represented by X 4,1 in the node, X i,j represents the jth node of the ith layer in the improved Unet++ network. In addition, a CBAM attention module is added to the skip connection of the first node and the last node of each layer of the improved Unet++ network.

[0052] The contour extraction module is configured to process the binary mask image output by the building recognition module. First, morphological filtering is used to remove noise. Then, a contour detection algorithm is used to extract the building contour from the denoised image. Then, a regression algorithm is used to simplify the extracted building contour. Finally, the simplified contour is converted into a standard geographic coordinate format file, referred to as a building contour coordinate format file.

[0053] The graphical user interface (GUI) module is configured to provide user interaction functions, including: a first function for receiving user input images, calling the contour extraction module to process the images, and outputting corresponding building contour coordinate format files; and a second function for providing a user manual marking tool to enable the user to mark the building contour on the image and save the marking result.

[0054] As shown in Figure 1 The data flow process of the building contour extraction system is as follows: image data is first preprocessed and enhanced by the data preprocessing module, and a training set and a test set are divided for training and testing the improved Unet++ network in the building recognition module. The user uploads the image that needs to be extracted by the building contour extraction system through the graphical user interface (GUI) module. The building recognition module outputs the corresponding binary mask image to the contour extraction module after obtaining the image. The contour extraction module extracts the building contour and saves it in the format of a geographic coordinate file to the database. The user can upload the image automatically to extract the building contour through the graphical user interface, or manually mark the building contour in the image and save the marking result file to the database.

[0055] Specifically, the data preprocessing module performs the following steps:

[0056] First, a normalization operation is performed on the input images to reduce the pixel value differences between different images. Given that image pixel values are usually distributed in the [0, 255] interval, this operation divides each pixel value by 255. Normalization processing unifies the pixel values of all images to the [0, 1] range, which helps to improve the convergence speed and stability of model training. Taking satellite remote sensing images as an example, their original pixel values are in the [0, 255] interval, and after normalization, they are mapped to floating-point numbers in the [0, 1] interval. This data format is more suitable for the input requirements of deep learning models, which facilitates more effective learning of building contour features by the models.

[0057] Second, to meet the specific requirements of deep learning models for input image size and ensure processing efficiency, the images need to be resized. During the resizing process, bilinear interpolation is used for resampling. This method estimates the pixel value at a new position by calculating the weighted average of neighboring pixels, which can preserve image detail information to some extent and reduce image distortion caused by scaling. Specifically, down-sampling is performed on images with large original sizes, while up-sampling is performed on images with small original sizes, ultimately making all images conform to the input size specifications of the model.

[0058] In addition, a series of data augmentation operations are performed on the image data to expand the size of the data set, improve data diversity, and enhance the robustness of the model:

[0059] Random flipping: To simulate the symmetric distribution characteristics of buildings and improve the segmentation effect of the model on left-right symmetric buildings, a horizontal flipping operation is performed on the image and its corresponding mask with a preset probability during data loading. This strategy can generate new training samples, enhance the model's adaptability to different building spatial layouts, and thus improve the accuracy of building contour recognition.

[0060] Random rotation: aimed at enhancing the model's segmentation ability for buildings in different orientations. A random angle rotation transformation is applied to the image and mask to generate rotated samples. To maintain the accuracy of the mask's pixel values, nearest neighbor interpolation is used to process the mask after rotation. By training the model to learn such rotated samples, the model's ability to recognize building contours in different orientations can be effectively improved, enhancing its robustness and generalization performance.

[0061] Random cropping: by randomly cropping a fixed size region from the image and mask for training, the model is prompted to focus on different local regions of the building. This operation increases the model's ability to learn local features, enabling it to grasp the contour features of buildings in different spatial positions and scales, thereby improving the model's precision in extracting building detail contours.

[0062] The training set and test set obtained by the data preprocessing module will be used to improve the training and testing of the Unet++ network. Figure 2 The structure of the improved Unet++ network is shown, and the specific introduction is as follows:

[0063] ① Remove the global average pooling layer and the fully connected layer at the end of the SeResNeXt network, and keep the feature extraction part as the backbone network; for the input image data, the backbone network outputs different scale feature maps s1, s2, s3, s4 from shallow to deep;

[0064] ② The Unet++ network is a major improved version of the classic U-Net architecture, and the core innovation is the introduction of densely nested skip connections, a total of 4 layers, of which the i-th layer has 5-i nodes, and the j-th node of the i-th layer is denoted as X i,j The densely nested skip connections require features from different stages of the encoder, i.e., different depths and different resolution feature maps, so the output feature maps of certain stages in the backbone network, i.e., the different scale feature maps s1, s2, s3, s4, are intercepted as the input of the first node X i,1 of each layer of the improved Unet++ network, which receives the up-sampling results from the nodes of the i+1-th layer, receives the feature maps from the backbone network at the corresponding level, and accepts the feature maps with interpolated resolution from the previous nodes of the improved Unet++ network, and splices them together to obtain the initial feature map of the current node, and then performs convolution operation to fuse the information;

[0065] ③ In the improved Unet++ network, for the first node and the last node of each layer, the feature maps from the backbone network received by the skip connection need to be processed by the CBAM attention module, and the CBAM attention module is used to weight the spatial information and channel information of the features.

[0066] The CBAM attention module aims to overcome the limitations of traditional convolutional neural networks in processing different scales, shapes and direction information. To this end, the CBAM attention module introduces two attention mechanisms: channel attention mechanism and spatial attention mechanism, as shown in Figure 3 and Figure 4 The channel attention mechanism helps to enhance the feature representation of different channels, while the spatial attention mechanism helps to extract key information at different positions in the space.

[0067] The goal of the channel attention mechanism is to adaptively enhance the feature representation capability of each channel in the input feature map. The implementation of this mechanism includes the following steps: For the input feature map, the channel attention mechanism performs global max pooling and global average pooling operations respectively. Specifically, for each channel of the feature map, the maximum and average values ​​at all spatial locations are calculated. This step generates two one-dimensional vectors, where the dimension of each vector equals the number of channels in the input feature map, representing the global maximum response feature and the global average response feature of each channel, respectively. The obtained global max pooling feature vector and global average pooling feature vector are input into a fully connected layer with shared parameters. This fully connected layer learns and generates attention weights for each channel, enabling the channel attention mechanism to adaptively evaluate the importance of different channels to the current task. Then, the two feature vectors processed by this fully connected layer are added element-wise to fuse two types of global statistical information, and a sigmoid activation function is applied for non-linear transformation to normalize and generate a channel attention weight vector. Finally, the normalized channel attention weight vector is used to multiply the original input feature map element-wise by channel, that is, multiplying the original feature value of each channel by the corresponding attention weight. The effect of this operation is that, based on the learned weights, the feature representation of channels that contribute greatly to the current task is significantly enhanced, while the feature representation of channels that contribute less or are irrelevant is suppressed, thereby obtaining an output feature map weighted by the attention mechanism.

[0068] The goal of the spatial attention mechanism is to adaptively highlight the importance of key spatial locations in the input feature map. This mechanism is implemented through the following steps: For the input feature map, global max pooling and global average pooling operations are performed along the channel dimension. This process generates two two-dimensional feature maps, each representing spatial context information under different receptive fields. The two feature maps are concatenated along the channel dimension to form a multi-scale spatial feature descriptor. The concatenated feature map is then input into a 7*7 convolutional layer to generate the original spatial attention weight matrix, and a sigmoid activation function is applied. This convolutional layer adaptively captures the dependencies between spatial locations through learnable parameters. The resulting normalized spatial attention map is then multiplied element-wise with the original input feature map. This significantly enhances the response intensity of the target region and effectively suppresses interference from non-critical regions. After training, the network outputs a binary mask image for each input image data. Figure 5 , Figure 6 The results are presented using the Austin dataset. Austin features buildings that blend modern and traditional architectural styles, with varying density, shapes, and orientations, scattered throughout the images. The segmentation masks predicted by the network accurately delineate the building outlines, and the segmentation results for most areas are consistent with the ground truth masks, demonstrating the network's good adaptability to urban environments.

[0069] Specifically, the contour extraction module performs the following steps:

[0070] 1) Morphological filtering on the binary mask image output by the building recognition module: first perform a closing operation to fill small holes inside the building, and then perform an opening operation to remove small noise points; wherein the kernel size of the morphological operation is 3x3;

[0071] 2) For the binary mask image processed in step 1), use the Suzuki-Abe algorithm to obtain the pixel-based building contour line, and filter out the contour line with too small area, too large area, and contact with the image edge. For the case where multiple contour lines are detected for the same building, calculate the minimum bounding rectangle of each contour line; if the minimum bounding rectangles of two contour lines are similar in position, it is determined that they belong to the same building, and the contour lines are merged accordingly;

[0072] 3) Perform polygon simplification on the contour line processed in step 2) to obtain simplified contour line coordinates. A contour line may have several hundred points, which will occupy too much memory if not simplified, but excessive simplification will also lose important shape features of the building. Use the Douglas-Peucker algorithm for simplification, select the points Pmin and Pmax with the smallest and largest horizontal coordinates on the contour line, divide the contour line into two halves, and perform the Douglas-Peucker algorithm on each half. The tolerance parameter of the algorithm uses an adaptive strategy, which automatically selects according to the perimeter of the contour line;

[0073] The simplified contour also needs to be checked for quality to ensure that no important shape features are lost. We establish two evaluation indicators:

[0074] Area retention rate This indicator intuitively reflects the impact of simplification on the area of the building, where P o represents the initial contour polygon, P s represents the simplified contour polygon, Area(P s ) represents the area of the simplified contour, and Area(P o ) represents the area of the initial contour.

[0075] Hausdorff distance D H (P s , P o ): measures the maximum deviation of the contour before and after simplification.

[0076]

[0077] where p1 and p2 represent the initial contour P o and the simplified contour P sThe point on the image is denoted as p, and d(p1, p2) represents the Euclidean distance between points p1 and p2.

[0078] In practical applications, when the area retention rate is lower than 95% or the Hausdorff distance exceeds a preset threshold, the system automatically reduces the simplification strength to ensure that important architectural features are not lost.

[0079] 4) The simplified contour line coordinates obtained in step 3) are converted from the image pixel coordinate system to the geographic coordinate system through affine transformation and saved as a standard geographic coordinate format file, referred to as a building contour coordinate format file.

[0080] Figure 7 The effect of the polygon contour extraction by the contour extraction module is shown.

[0081] Specifically, the framework of the graphical interactive interface GUI module is as shown in Figure 8 It can be seen that the most important role is that the user realizes label management and file operation through the control panel, and the core includes:

[0082] The first function: load and display the image specified by the user; in response to user operation, call the contour extraction module to process the image, automatically extract the building contour and display the result, and output the corresponding building contour coordinate format file;

[0083] The second function: load and display the image specified by the user; provide a drawing tool to enable the user to manually draw and edit the building contour polygon on the image; save the contour polygon data drawn by the user.

[0084] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application are equivalent replacement methods and are included in the protection scope of the present application.

Claims

1. A building contour extraction system characterized by, The application relates to a building contour extraction method based on improved Unet++ network, which comprises the following steps: A data preprocessing module is used for loading and preprocessing images, dividing the preprocessed images into a training set and a test set, and performing data enhancement on the training set; The building recognition module is used for inputting the enhanced training set into the improved Unet++ network for training, and after the training is completed, inputting the test set into the trained improved Unet++ network for recognition. The improved Unet++ network outputs a binary mask image for each input image. The improved Unet++ network is a SeResNeXt network as a backbone network, and the last classification layer of the SeResNeXt network is removed. For the input image data, the backbone network generates different scale feature maps s1, s2, s3 and s4 from shallow to deep, and the feature maps are sequentially input into the X 1,1 The node is connected to X 4,1 In the node, X i,j represents the jth node of the ith layer in the improved Unet++ network, and a CBAM attention module is added to the skip connection of the first node and the last node of each layer of the improved Unet++ network. A contour extraction module is used for processing a binary mask image output by the building recognition module: firstly, noise is removed through morphological filtering; then, building contours are extracted from the denoised image by using a contour detection algorithm; the extracted building contours are simplified by using a regression algorithm; and finally, the simplified contours are converted into a standard geographic coordinate format file, which is called a building contour coordinate format file; A graphical user interface (GUI) module is used for providing user interaction functions, including: a first function of receiving a user input image, calling the contour extraction module to process the image, and outputting a corresponding building contour coordinate format file; and a second function of providing a user manual marking tool to enable the user to mark a building contour on the image and save the marking result.

2. The building contour extraction system of claim 1, wherein, The data preprocessing module performs the following steps: 1) Normalization is performed on the input image, each pixel value in the image is divided by 255, all the pixels are converted into floating-point numbers between 0 and 1, and the image is divided into a training set and a test set, wherein each image in the training set has a corresponding binary mask image; 2) Data enhancement is performed on the training set obtained in step 1), specifically, the same random horizontal flip, random rotation and random cropping are performed on each image and the corresponding binary mask image in the training set to obtain an enhanced training set; 3) The image size of the test set in step 1) and the enhanced training set in step 2) is adjusted, and all the images are adjusted to 512*512 pixels in size by using a bilinear interpolation method.

3. The building contour extraction system of claim 1, wherein, The improved Unet++ network is specifically introduced as follows: ① The global average pooling layer and the full connection layer at the end of the SeResNeXt network are removed, and the feature extraction part thereof is reserved as a backbone network; for input image data, the backbone network outputs different scale feature maps s1, s2, s3 and s4 from shallow to deep; ②Unet++ network is a major improved version of the classic U-Net architecture, the core innovation lies in the introduction of densely nested skip connections, a total of 4 layers, of which the i-th layer has 5-i nodes, and the j-th node of the i-th layer is denoted as X i,j The densely nested skip connections of the node need to come from different stages of the encoder, that is, different depth, different resolution feature maps, so the output feature maps of a certain few stages in the backbone network, that is, the above different scale feature maps s1, s2, s3, s4 are intercepted as the first node X of each layer of the improved Unet++ network i,1 The input of each node of the i-th layer of the improved Unet++ network receives the up-sampling result from the node of the i+1-th layer, receives the feature map from the backbone network at the corresponding level, and accepts the feature map of the interpolation resolution adjustment of the previous node of the improved Unet++ network, and splices them together to obtain the initial feature map of the current node, and then performs convolution operation to fuse the information; ③ In the improved Unet++ network, the feature maps from the backbone network received by the skip connection of each layer first node and last node need to be processed through a CBAM attention module, and the CBAM attention module is used for weighting the spatial information and channel information of the features.

4. The building contour extraction system of claim 1, wherein, The contour extraction module performs the following steps: 1) The binary mask image output by the building recognition module is subjected to morphological filtering: firstly, a closing operation is performed to fill small holes in the building interior, and then an opening operation is performed to remove small noise points; wherein the kernel size of the morphological operation is 3*3; 2) The binary mask image processed in step 1) is subjected to Suzuki-Abe algorithm to obtain a pixel-based building contour line, and the contour lines with an area not in a preset range and contacting the image edge are filtered out; for the case that multiple contour lines of the same building are detected, the minimum bounding rectangle of each contour line is calculated, and if the minimum bounding rectangles of two contour lines are similar, it is determined that the two contour lines belong to the same building, and the contour lines are combined according to the determination result. 3) Simplify the contour line processed in step 2) to get the simplified contour line coordinates, wherein the Douglas-Peucker algorithm is used for simplification, the contour line is divided into two halves by selecting the points Pmin and Pmax with the minimum and maximum horizontal coordinates on the contour line, and the Douglas-Peucker algorithm is executed on the two halves respectively, and the tolerance parameter of the algorithm is automatically selected according to the circumference of the contour line by using an adaptive strategy; 4) Convert the simplified contour line coordinates obtained in step 3) from the image pixel coordinate system to the geographic coordinate system by affine transformation, and save the result as a standard geographic coordinate format file, which is referred to as a building contour coordinate format file.

5. The building contour extraction system of claim 1, wherein, The functions of the graphical interactive interface GUI module include: A first function: loading and displaying a user-specified image; in response to a user operation, calling the contour extraction module to process the image, automatically extracting a building contour, and displaying the result, while outputting a corresponding building contour coordinate format file; A second function: loading and displaying a user-specified image; providing a drawing tool to enable the user to manually draw and edit a building contour polygon on the image; and saving the contour polygon data drawn by the user.