A Building Recognition and Extraction Method and System Based on an Improved Neural Network

By improving the DeepLabV3+ model, the limitations of building feature extraction in remote sensing images are solved by using lightweight and high-performance backbone networks and two ASPP modules, and efficient building profile detection and extraction are achieved.

CN116403109BActive Publication Date: 2025-06-13SHANGHAI UNIV OF MEDICINE & HEALTH SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310291665.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-23
Publication Date
2025-06-13
Estimated Expiration
2043-03-23

AI Technical Summary

Technical Problem

The prior art methods for extracting building elements in remote sensing images have limitations and are difficult to effectively overcome difficulties in practical applications. In addition, traditional methods rely on artificial design features, and there are uncertainty and subjectivity.

Method used

A building recognition and extraction method based on improved neural network is proposed. By improving the DeepLabV3+ model, the lightweight high-performance backbone network and two ASPP modules are used to improve the model's calculation speed and the extraction ability of high-level semantic features.

Benefits of technology

It realizes efficient detection and extraction of building profiles, reduces the model's parameter calculation amount and memory usage, and improves the calculation speed and edge feature extraction ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403109B_ABST
    Figure CN116403109B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for building recognition and extraction based on an improved neural network. The method includes the following steps: collecting original building images, preprocessing the original images to obtain preprocessed image data; constructing a DeepLabV3+ model and improving the DeepLabV3+ model to obtain an improved model; training the improved model based on the preprocessed image data to obtain a prediction model; collecting images of buildings to be recognized, predicting and generating building edge images based on the prediction model, and regularizing the building edge images to obtain building contour extraction results. The lightweight and high-performance backbone network proposed in the present application incorporates the design concept of ConvNeXt on the basis of the traditional DenseNet, greatly reducing the parameter calculation amount of the model, reducing memory occupancy, and improving the calculation speed of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly relates to a method and system for building recognition and extraction based on an improved neural network. Background Art

[0002] Most of the traditional methods for extracting building elements in remote sensing images are based on single-image information and use traditional methods such as segmentation, classification, and edge detection to achieve. These methods have great limitations and many difficulties that are difficult to overcome in practical applications.

[0003] In recent years, after many scholars used convolutional neural networks to achieve image recognition, deep learning methods have been increasingly applied in the field of feature extraction of remote sensing images. Deep learning technology has achieved remarkable achievements in various computer vision tasks such as image classification, segmentation, detection, and target recognition, and has gradually been used to solve the problem of extracting feature elements such as buildings, roads, and water foam lines in remote sensing images. The essential feature of deep learning is to use the learning algorithm of the computer to automatically learn high-level features from large sample data, so as to have the ability to predict the features of unknown data. Deep learning models do not require feature extraction, reducing the uncertainty and subjectivity of artificially designed features. Compared with traditional machine learning methods, its multi-level deep neural network shows stronger learning ability and representation ability for sample features, and can improve the information recognition efficiency and accuracy of massive image data. For the extraction of remote sensing image feature elements, it is essentially the segmentation of the target object, and the convolutional neural network in deep learning has a greater impact on the segmentation task because of its strong feature extraction and mining ability. End-to-end, graph-to-graph deep semantic segmentation algorithms have emerged one after another. Typical models of this type of algorithm include FCN, SegNet, U-Net, PSPNet, DeepLabV3+, etc. These deep learning models are increasingly applied to the segmentation task of remote sensing images, and continuous breakthroughs have been made in the accuracy of feature segmentation. Summary of the Invention

[0004] This application aims to solve the deficiencies of the prior art and proposes a method and system for building recognition and extraction based on an improved neural network. By improving the DeepLabV3+ model, the buildings in the target area are detected to obtain the building outlines.

[0005] To achieve the above object, this application provides the following solutions:

[0006] A method for building recognition and extraction based on an improved neural network, comprising the following steps:

[0007] Collect the original building images, and preprocess the original building images to obtain preprocessed image data;

[0008] Build the DeepLabV3+ model and improve the DeepLabV3+ model to obtain an improved model;

[0009] Train the improved model based on the preprocessed image data to obtain a prediction model;

[0010] Collect the building images to be recognized, predict and generate the building edge image based on the prediction model, and perform regularization processing on the building edge image to obtain the building contour extraction result.

[0011] Preferably, the method of the preprocessing includes:

[0012] Crop the original building image to obtain a cropped image dataset;

[0013] Perform data augmentation on the cropped image dataset, and screen the image dataset after data augmentation to obtain a screened image dataset;

[0014] Perform feature annotation on the screened image dataset based on ArcGIS Pro to obtain an annotated image dataset;

[0015] Divide the annotated dataset into a training set, a validation set and a test set to obtain the preprocessed image data.

[0016] Preferably, the method for obtaining the improved model includes:

[0017] Use a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model;

[0018] Add two ASPP modules to the DeepLabV3+ model to obtain the improved model.

[0019] Preferably, the training method of the prediction model includes:

[0020] Input the training set into the improved model to solve the network parameters in the case of minimizing the loss function;

[0021] Input the validation set into the improved model to minimize the overfitting of the improved model;

[0022] Input the test set into the improved model, compare the accuracy of the output result with the true classification result, and adjust the network parameters based on the comparison result to obtain the prediction model.

[0023] Preferably, the method of the regularization processing includes:

[0024] Extract the boundary image of the building edge image using the Marching Cubes model;

[0025] Polygonize the boundary image using the Douglas-Peucker model to obtain the building contour extraction result.

[0026] Preferably, the method for extracting the boundary image includes:

[0027] Eliminate errors in the building edge image based on a rough adjustment algorithm to obtain a rough-adjusted image;

[0028] Adjust the direction of the lines and the positions of the nodes in the rough-adjusted image based on a fine-tuning algorithm to obtain the boundary image.

[0029] This application also provides a building recognition and extraction system based on an improved neural network, including: an image preprocessing module, a model construction module, a model training module, and a recognition and extraction module;

[0030] The image preprocessing module is used to collect the original building image and preprocess the original building image to obtain preprocessed image data;

[0031] The model construction module is used to construct a DeepLabV3+ model and improve the DeepLabV3+ model to obtain an improved model;

[0032] The model training module is used to train the improved model based on the preprocessed image data to obtain a prediction model;

[0033] The recognition and extraction module is used to collect the image of the building to be recognized, predict and generate the building edge image based on the prediction model, and perform regularization processing on the building edge image to obtain the building contour extraction result.

[0034] Preferably, the method for obtaining the improved model includes:

[0035] Use a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model;

[0036] Add two ASPP modules to the DeepLabV3+ model to obtain the improved model.

[0037] Compared with the prior art, the beneficial effects of this application are:

[0038] (1) The lightweight and high-performance backbone network proposed in this application incorporates the design concept of ConvNeXt on the basis of the traditional DenseNet, significantly reducing the parameter calculation amount of the model, reducing memory occupancy, and improving the calculation speed of the model.

[0039] (2) This application uses two ASPP modules to fuse image features, thereby obtaining more advanced semantic information, enhancing the extraction of edge features, and further improving the ability to extract advanced semantics. Brief Description of the Drawings

[0040] To more clearly illustrate the technical solutions of this application, the following briefly introduces the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 It is a schematic flowchart of the method of the embodiment of this application;

[0042] Figure 2 It is a schematic diagram of the improved model structure of the embodiment of this application;

[0043] Figure 3 It is the bottleneck layer designed by DenseNeXt of the embodiment of this application;

[0044] Figure 4 It is a schematic diagram of the system structure of the embodiment of this application. Detailed Embodiments

[0045] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0046] To make the above objects, features, and advantages of this application more obvious and understandable, the following further details this application in conjunction with the drawings and specific embodiments.

[0047] Embodiment 1

[0048] In this embodiment, as Figure 1 shown, a method for building recognition and extraction based on an improved neural network includes the following steps:

[0049] S1. Collect the original building image, preprocess the original building image, and obtain the preprocessed image data.

[0050] The preprocessing method includes: cropping the original building image to obtain a cropped image dataset; performing data augmentation on the cropped image dataset and screening the augmented image dataset to obtain a screened image dataset; performing feature annotation on the screened image dataset based on ArcGIS Pro to obtain an annotated image dataset; dividing the annotated dataset into a training set, a validation set, and a test set to obtain the preprocessed image data.

[0051] In this embodiment, the original building image is from the remote sensing image captured by a drone in actual production. Since the size of the remote sensing image is usually large and deep learning cannot support large-size data training, the sample data is first cropped into blocks of 512×512 size to obtain a cropped image dataset. Aiming at the problem of overfitting caused by insufficient feature extraction during training due to the small number of training set samples, data augmentation is performed on the cropped image dataset. The image is augmented by means of horizontal flipping, vertical rotation, central cropping, random brightness contrast, elastic transformation, Gaussian noise, and channel transposition. At the same time, aiming at the phenomena of blank areas, blurred images, and incomplete annotations in the samples, data screening is carried out to obtain a screened image dataset.

[0052] The target area of the screened image dataset is annotated so that the deep learning model can learn the features of the building to distinguish it from other areas in the image. Therefore, the production of data labels is particularly important. The traditional production of data labels uses a pure manual annotation method, such as manually drawing using tools like Labelme. This manual annotation is cumbersome and time-consuming, while this embodiment uses a semi-automatic annotation method based on ArcGIS Pro to obtain an annotated image dataset.

[0053] The annotated image dataset is further divided into three parts: a training set, a validation set, and a test set. Among them, there are 18,481 images in the training set, 945 images in the validation set, and 475 images in the test set.

[0054] S2. Construct a DeepLabV3+ model and improve the DeepLabV3+ model to obtain an improved model.

[0055] The method for obtaining the improved model includes: using a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model; adding two ASPP modules in the DeepLabV3+ model to obtain the improved model.

[0056] The DeepLabV3+ network is one of the most excellent semantic segmentation models at present, and it has achieved excellent results on the VOC dataset. However, the DeepLabV3+ model also has some deficiencies. First, during the feature extraction process at the encoding end, the spatial dimension of the input data is gradually reduced, resulting in the loss of useful information, and it cannot well restore details during decoding. Second, although the introduction of the ASPP module can improve the model's ability to extract the boundaries of objects, it cannot fully simulate the connections between local features of the objects, resulting in a hole phenomenon in object segmentation and reducing the accuracy of object segmentation. Finally, in order to pursue segmentation accuracy, Xception with a larger number of network layers and parameters is selected as the feature extraction network, and the convolution method in the ASPP module is ordinary convolution, which further increases the number of parameters. The increase in model depth and the number of parameters leads to an increase in model complexity, higher requirements for hardware, an increase in network training difficulty, slower network training speed, and slower convergence.

[0057] In this embodiment, in order to improve the network segmentation performance and address the above deficiencies, as Figure 2 shown, the following improvements are made to the traditional DeepLabV3+ network structure: (1) Aiming at the problem of the large number of parameters in the Xception network for feature extraction in the traditional DeepLabV3+ model, a lightweight high-performance backbone network named DenseNeXt is proposed to replace the Xception network in the traditional DeepLabV3+. The proposed DenseNeXt network incorporates the design concept of ConvNeXt on the basis of the traditional DenseNet, greatly reducing the parameter calculation amount of the model, reducing memory occupancy, and improving the calculation speed of the model; (2) In order to further enhance the ability of the DeepLabV3+ model to extract high-level semantic features, after the DenseNeXt network extracts features from the input image, two ASPP modules are used to fuse the image features, thereby obtaining more high-level semantic information and enhancing the extraction of edge features; Through the above improvements, the improved model is obtained.

[0058] Among them, the stacking ratio of the 4-stage blocks of the proposed DenseNeXt network is set to 1:1:3:1. The specific number of layers in each stage is 8, 8, 24, and 8 respectively. The proposed DenseNeXt network designs two branches. One branch is a depthwise separable convolution with a 7×7 convolutional kernel, and the other is a depthwise separable convolution with a 3×3 convolutional kernel. Then the output feature maps of them are added together, and finally concatenated with the input feature map of the bottleneck layer as the output feature map of the bottleneck layer, so that the model obtains the ability of multi-scale feature extraction. Figure 3 is the bottleneck layer designed by DenseNeXt.

[0059] S3. Based on the preprocessed image data, train the improved model to obtain a prediction model.

[0060] The training method of the prediction model includes: inputting the training set into the improved model to solve the network parameters under the condition of minimizing the loss function; inputting the validation set into the improved model to minimize the overfitting of the improved model; inputting the test set into the improved model, comparing the accuracy of the output result with the true classification result, and adjusting the network parameters based on the comparison result to obtain the prediction model.

[0061] In this embodiment, the training data is used to solve the network parameters that minimize the loss function; the validation data is used to minimize overfitting; the test data is used to test the classification ability of the network after the network training is completed. Input the test data into the trained deep neural network structure, calculate the difference between the output result and its true classification result, estimate the classification accuracy of the network, optimize and adjust or appropriately increase the labeled samples according to the model verification situation. When the labeled sample library is optimized and adjusted, the parameters of the deep neural network can be fine-tuned, the network structure can be optimized, and the classification accuracy of the network can be further improved to obtain the prediction model.

[0062] S4. Collect the image of the building to be recognized, predict and generate the edge image of the building based on the prediction model, and regularize the edge image of the building to obtain the building contour extraction result.

[0063] The method of regularization processing includes: using the Marching Cubes model to extract the boundary image of the edge image of the building. The extraction method of the boundary image includes: eliminating the obvious errors of the edge image of the building based on the coarse adjustment algorithm to obtain the coarsely adjusted image; adjusting the direction of the lines and the positions of the nodes in the coarsely adjusted image based on the fine adjustment algorithm to obtain the boundary image; using the Douglas-Peucker model to polygonize the boundary image to obtain the building contour extraction result.

[0064] In this embodiment, since there are irregularities such as the edges of the house area generated by model prediction, this project regularizes the contour line of the building coverage area by extracting the key points of the building and the main direction of the building to eliminate the irregular boundaries and details in the geometry of the building range.

[0065] First, use the Marching Cubes algorithm to implement boundary extraction. The main steps are divided into two steps. The coarse adjustment algorithm eliminates the obvious errors in segmentation and polygonization, and the further fine adjustment algorithm adjusts the direction of the lines and the positions of the nodes.

[0066] Implementation process of the coarse-tuning algorithm: Remove polygons S with an area below the threshold; Delete edges Td with a length below the given side length; Remove overly acute angles α with a threshold; Remove overly smooth angles β with a threshold. Implementation process of the fine-tuning algorithm: Find long edges W with a threshold; Add the direction of the longest edge to the main direction list; Add the directions of other edges to the main direction list according to the angle threshold, where δ is between their direction and the direction in the list; Adjust the long edges according to the list and the angle, and adjust the short edges according to the list and the angle (judged by the threshold θ); If the distance between two lines is less than (or greater than) the threshold, merge (or connect) parallel lines d; Connect all adjusted lines to form the final polygon. The thresholds in this embodiment need to be set according to the actual situation.

[0067] Then use the Douglas-Peucker algorithm to implement polygonization.

[0068] Process the building image to be recognized, and input the processed image into the trained improved DeepLabV3+ network model to detect the buildings in the building image to be recognized, and finally generate a building extraction map.

[0069] Embodiment 2

[0070] In this embodiment, as Figure 4 shown, a building recognition and extraction system based on an improved neural network includes: an image preprocessing module, a model construction module, a model training module, and a recognition and extraction module.

[0071] The image preprocessing module is used to collect the original building image and preprocess the original building image to obtain preprocessed image data.

[0072] The preprocessing method includes: Cropping the original building image to obtain a cropped image dataset; Performing data augmentation on the cropped image dataset, and screening the image dataset after data augmentation to obtain a screened image dataset; Performing feature annotation on the screened image dataset based on ArcGIS Pro to obtain an annotated image dataset; Dividing the annotated dataset into a training set, a validation set, and a test set to obtain preprocessed image data.

[0073] In this embodiment, the original building images are from remotely sensed images taken by drones in actual production. Since the size of remotely sensed images is usually large and deep learning cannot support the training of large-sized data, the sample data is first cropped into blocks of 512×512 size to obtain the cropped image dataset. Aiming at the problem of overfitting caused by insufficient feature extraction during training due to the small sample size of the training set, data augmentation is performed on the cropped image dataset. The images are augmented by means such as horizontal flipping, vertical rotation, central cropping, random brightness contrast, elastic transformation, Gaussian noise, and channel transposition. At the same time, aiming at the phenomena of blank areas, blurred images, and incomplete annotations in the samples, data screening is carried out to obtain the screened image dataset.

[0074] The target areas of the screened image dataset are labeled so that the deep learning model can learn the features of the buildings to distinguish them from other areas in the image. Therefore, the production of data labels is particularly important. Traditional data label production uses a pure manual annotation method, such as using tools like Labelme for manual drawing. This manual annotation is cumbersome and time-consuming. In this embodiment, a semi-automatic annotation method based on ArcGIS Pro is adopted to obtain the labeled image dataset.

[0075] The labeled image dataset is further divided into three parts: a training set, a validation set, and a test set. Among them, there are 18,481 images in the training set, 945 images in the validation set, and 475 images in the test set.

[0076] The model construction module is used to construct a DeepLabV3+ model and improve the DeepLabV3+ model to obtain the improved model.

[0077] The method for obtaining the improved model includes: using a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model; adding two ASPP modules to the DeepLabV3+ model to obtain the improved model.

[0078] The DeepLabV3+ network is one of the most excellent semantic segmentation models at present, and it has achieved excellent results on the VOC dataset. However, the DeepLabV3+ model also has some deficiencies. First, during the feature extraction process at the encoding end, the spatial dimension of the input data is gradually reduced, resulting in the loss of useful information, and it cannot well restore details during decoding. Second, although the introduction of the ASPP module can improve the model's ability to extract the boundaries of targets, it cannot completely simulate the connections between local features of targets, resulting in hole phenomena in target segmentation and reducing the accuracy of target segmentation. Finally, in order to pursue segmentation accuracy, Xception with a larger number of network layers and parameters is selected as the feature extraction network, and the convolution method in the ASPP module is ordinary convolution, further increasing the number of parameters. The increase in model depth and the number of parameters leads to an increase in model complexity, higher requirements for hardware, increased network training difficulty, slower network training speed, and slower convergence.

[0079] In this embodiment, in order to improve the network segmentation performance and address the above deficiencies, the following improvements are made to the traditional DeeplabV3+ network structure: (1) Aiming at the problem of the large number of parameters in the Xception network for feature extraction in the traditional DeepLabV3+ model, a lightweight high-performance backbone network named DenseNeXt is proposed to replace the Xception network in the traditional DeepLabV3+. The proposed DenseNeXt network incorporates the design idea of ConvNeXt on the basis of the traditional DenseNet, significantly reducing the parameter calculation amount of the model, reducing memory occupancy, and improving the calculation speed of the model; (2) In order to further enhance the ability of the DeepLabV3+ model to extract high-level semantic features, after the DenseNeXt network extracts features from the input image, two ASPP modules are used to fuse the image features to obtain more high-level semantic information and enhance the extraction of edge features; Through the above improvements, the improved model is obtained.

[0080] Among them, the stacking ratio of the 4 stage blocks of the proposed DenseNeXt network is set to 1:1:3:1. The specific number of layers in each stage is 8, 8, 24, and 8 respectively. The proposed DenseNeXt network designs two branches. One branch is a depthwise separable convolution with a 7×7 convolution kernel, and the other is a depthwise separable convolution with a 3×3 convolution kernel. Then the output feature maps of them are added together, and finally concatenated with the input feature map of the bottleneck layer as the output feature map of the bottleneck layer, so that the model obtains the ability of multi-scale feature extraction.

[0081] The model training module is used to train the improved model based on the preprocessed image data to obtain the prediction model.

[0082] The training method of the prediction model includes: inputting the training set into the improved model to solve the network parameters under the condition of minimizing the loss function; inputting the validation set into the improved model to minimize the overfitting of the improved model; inputting the test set into the improved model, comparing the accuracy of the output result with the true classification result, and adjusting the network parameters based on the comparison result to obtain the prediction model.

[0083] In this embodiment, the training data is used to solve the network parameters that minimize the loss function; the validation data is used to minimize overfitting; the test data is used to test the classification ability of the network after the network training is completed. The test data is input into the trained deep neural network structure, the difference between the output result and its true classification result is calculated, the classification accuracy of the network is estimated, and according to the model validation situation, the labeled samples are optimized and adjusted or appropriately increased. When the labeled sample library is optimized and adjusted, the parameters of the deep neural network can be fine-tuned, the network structure can be optimized, and the network classification accuracy can be further improved to obtain the prediction model.

[0084] The recognition and extraction module is used to collect the building image to be recognized, predict and generate the building edge image based on the prediction model, and regularize the building edge image to obtain the building contour extraction result.

[0085] The method of regularization processing includes: using the Marching Cubes model to extract the boundary image of the building edge image. The extraction method of the boundary image includes: eliminating the errors of the building edge image based on the coarse adjustment algorithm to obtain the coarsely adjusted image; adjusting the direction of the lines and the positions of the nodes in the coarsely adjusted image based on the fine adjustment algorithm to obtain the boundary image; using the Douglas-Peucker model to polygonize the boundary image to obtain the building contour extraction result.

[0086] In this embodiment, since there are irregularities and other phenomena in the edges of the house areas predicted by the model, this project regularizes the contour lines of the building coverage area by extracting the key points of the building and the main direction of the building to eliminate the irregular boundaries and details in the geometry of the building range.

[0087] First, use the Marching Cubes algorithm to implement boundary extraction. The main steps are divided into two steps. The coarse adjustment algorithm is used to eliminate obvious errors in segmentation and polygonization, and the fine adjustment algorithm is further used to adjust the direction of the lines and the positions of the nodes.

[0088] Implementation process of the coarse adjustment algorithm: Remove the polygon S with an area lower than the threshold; Delete the edges Td with a length lower than the given side length; Remove the overly acute angle α with a threshold; Remove the overly smooth angle β with a threshold. Implementation process of the fine adjustment algorithm: Find the long side W with a threshold; Add the direction of the longest side to the main direction list; Add the directions of other sides to the main direction list according to the angle threshold, where δ is between their directions and the direction in the list; Adjust the long sides according to the list and the angle, and adjust the short sides according to the list and the angle (judging by the threshold θ); If the distance between two lines is less than (or greater than) the threshold, merge (or connect) the parallel lines d; Connect all the adjusted lines to form the final polygon. The thresholds in this embodiment need to be set according to the actual situation.

[0089] Then use the Douglas - Peucker algorithm to implement polygonization.

[0090] Process the building image to be recognized, and input the processed image into the trained improved DeepLabV3+ network model to detect the buildings in the building image to be recognized, and finally generate a building extraction map.

[0091] The embodiments described above are only descriptions of the preferred embodiments of the present application, and do not limit the scope of the present application. Without departing from the design spirit of the present application, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present application shall fall within the protection scope determined by the claims of the present application.

Claims

1. A building recognition and extraction method based on an improved neural network, characterized in that, it includes the following steps: Collect the original building image, preprocess the original building image to obtain preprocessed image data; Construct a DeepLabV3+ model and improve the DeepLabV3+ model to obtain an improved model; Based on the preprocessed image data, train the improved model to obtain a prediction model; Collect the building image to be recognized, based on the prediction model, predict and generate a building edge image, and regularize the building edge image to obtain a building contour extraction result; The method for obtaining the improved model includes: Use a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model; Add two ASPP modules to the DeepLabV3+ model to obtain the improved model. The method includes: (1) Use the DenseNeXt network to replace the Xception network in the traditional DeepLabV3+ model. The proposed DenseNeXt network incorporates ConvNeXt on the basis of the traditional DenseNet; (2) Use two ASPP modules to fuse image features, thereby obtaining more high-level semantic information and enhancing the extraction of edge features; Through the above improvements, an improved model is obtained; The stacking ratio of the 4-stage blocks of the DenseNeXt network is set to 1:1:3:1, and the specific number of layers in each stage is 8, 8, 24, and 8 respectively; The DenseNeXt network designs two branches, one branch is a depthwise separable convolution with a 7×7 kernel size, and the other is a depthwise separable convolution with a 3×3 kernel size. Add their output feature maps, and finally concatenate them with the input feature map of the bottleneck layer as the output feature map of the bottleneck layer.

2. The building recognition and extraction method based on an improved neural network according to claim 1, characterized in that, the preprocessing method includes: Crop the original building image to obtain a cropped image dataset; Perform data augmentation on the cropped image dataset, and screen the image dataset after data augmentation to obtain a screened image dataset; Based on ArcGIS Pro, perform feature annotation on the screened image dataset to obtain an annotated image dataset; Divide the annotated image dataset into a training set, a validation set, and a test set to obtain the preprocessed image data.

3. The building recognition and extraction method based on an improved neural network according to claim 2, characterized in that, the training method of the prediction model includes: Input the training set into the improved model to solve the network parameters under the condition of minimizing the loss function; Input the validation set into the improved model to minimize the overfitting of the improved model; Input the test set into the improved model, compare the accuracy of the output result with the true classification result, and adjust the network parameters based on the comparison result to obtain the prediction model.

4. A method for building recognition and extraction based on an improved neural network according to claim 1, wherein, the method of regularization processing includes: using the Marching Cubes model to extract the boundary image of the building edge image; using the Douglas-Peucker model to polygonize the boundary image to obtain the building contour extraction result.

5. A method for building recognition and extraction based on an improved neural network according to claim 4, wherein, the method for extracting the boundary image includes: eliminating errors in the building edge image based on a coarse adjustment algorithm to obtain a coarsely adjusted image; adjusting the direction of the lines and the node positions in the coarsely adjusted image based on a fine adjustment algorithm to obtain the boundary image.

6. A building recognition and extraction system based on an improved neural network, wherein, it includes: an image preprocessing module, a model construction module, a model training module, and a recognition and extraction module; the image preprocessing module is used to collect the original building image and preprocess the original building image to obtain preprocessed image data; the model construction module is used to construct a DeepLabV3+ model and improve the DeepLabV3+ model to obtain an improved model; the model training module is used to train the improved model based on the preprocessed image data to obtain a prediction model; the recognition and extraction module is used to collect the building image to be recognized, predict and generate a building edge image based on the prediction model, and perform regularization processing on the building edge image to obtain a building contour extraction result; the method for obtaining the improved model includes: using a lightweight and high-performance backbone network to replace the Xception network in the DeepLabV3+ model; adding two ASPP modules to the DeepLabV3+ model to obtain the improved model; adding two ASPP modules to the DeepLabV3+ model to obtain the improved model, and the method includes: (1) using the DenseNeXt network to replace the Xception network in the traditional DeepLabV3+ model, and the proposed DenseNeXt network incorporates ConvNeXt on the basis of the traditional DenseNet; (2) using two ASPP modules to fuse the image features to obtain more high-level semantic information and enhance the extraction of edge features; through the above improvements, the improved model is obtained. The stacking ratio of the four stage blocks of the DenseNeXt network is set to 1:1:3:1, and the specific number of layers in each stage is 8, 8, 24, and 8 respectively; the DenseNeXt network designs two branches, one of which is a depthwise separable convolution with a 7×7 convolutional kernel, and the other is a depthwise separable convolution with a 3×3 convolutional kernel, and their output feature maps are added together, and finally concatenated with the input feature map of the bottleneck layer as the output feature map of the bottleneck layer.