A method for extracting building outlines from aerial images based on machine learning

Through the machine learning-based aerial image building profile extraction method, the self-net model is used for training, which solves the problem of inefficiency in building profile extraction by traditional surveying and mapping methods, and realizes efficient and accurate data production and model training.

CN114821290BActive Publication Date: 2025-07-01NINGBO SHANGHANG SURVEYING & MAPPING
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210104501.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-07-01
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

Traditional surveying and mapping methods have high accuracy but low efficiency in building profile extraction, which cannot meet the demand for high-precision and rapid data production in smart cities.

Method used

Using a machine learning-based aerial image building profile extraction method, a custom-resolution data set is produced and trained using the Self-net model. The model includes four sets of downsampling blocks and four sets of upsampling blocks, feature extraction and image augmentation through convolution, activation and normalization processing layers.

Benefits of technology

It realizes efficient and accurate building profile extraction, optimizes the data set production and model training process, reduces the amount of parameters, and improves the integrity and computing efficiency of data features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821290B_ABST
    Figure CN114821290B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting building outlines from aerial images based on machine learning, including: Step 1: Produce a dataset of building outlines in aerial images; Step 2: Input the dataset into the Self-net model for training, and the Self-net model includes four groups of downsampling blocks and four groups of upsampling blocks; Step 3: Verify the obtained results, calculate the loss value of the model, and finally compare and calculate with the true value, propagate the error back to the model, continuously update the model, and perform iterative learning; Step 4: Finally, the model performs convolutional classification on the data and outputs the data. The present invention realizes that the dataset can be output with a customizable resolution, the outlines are more regular and neat, and the dataset production can be fully automated compared with other means; the improved model has an overall better effect than the original model in this dataset. Under the circumstances of this dataset and the current initialization parameters, the initial accuracy is relatively high, and a relatively high training accuracy can be achieved within a relatively small number of cycles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision three-dimensional image reconstruction, and more specifically, to a method for extracting building outlines from aerial images based on machine learning. Background Art

[0002] With the in-depth construction and development of digital cities and smart cities, people's demand for geographic information data is becoming more and more refined, and the update speed is accelerating day by day. Whether in terms of the expression form or the data production speed, the traditional geographic information data production process can no longer meet the current social needs. The hardware equipment and technical methods based on new surveying and mapping are constantly being updated, posing higher requirements for data production in terms of time and accuracy. Currently, the oblique photogrammetry software developed based on graphics can only meet visual viewing, but cannot meet the needs of data applications. In the production of single-body data applied in smart cities, many software rely on building contour lines. The extraction and separation of buildings based on contour lines are involved in the construction process of smart cities, which can greatly improve the refinement of smart cities. It helps the full-automatic modeling of smart cities from Mesh to single-body, helps smart cities better express each component and structure, and provides a carrier for the intelligent decision-making of the city.

[0003] How to obtain more accurate building outlines and more information about buildings, and combine new technologies of artificial intelligence has become a common topic in the surveying and mapping industry. The application potential of artificial intelligence in the surveying and mapping industry is huge. How to achieve wider coverage and more efficient applications has become the key to the future development of artificial intelligence.

[0004] In the development of national basic geographic information, the realization of urban building informatization is crucial. The level of urban building informatization will affect the efficiency of work such as digital city construction and land use survey. Traditional surveying and mapping methods have the advantage of high accuracy in the extraction of urban buildings, but most traditional surveying and mapping rely on manual measurement, so there are disadvantages such as low work efficiency and high work costs, and can no longer meet the current work requirements of urban building informatization. With the continuous development of new technologies, it is accelerating the transformation from traditional manual operations to machine learning, further optimizing the operation process, improving work efficiency, and enabling the surveying and mapping industry to better adapt to the development of society. As an emerging technology, machine learning is being vigorously applied in all walks of life. Especially for the development of the surveying and mapping industry, it requires the support of new technology means. Although the building extraction results based on the fusion of high-resolution remote sensing images and deep neural networks perform well, there is still great room for development in the research of using improved deep neural network structures to improve building extraction accuracy.

[0005] The research directions at home and abroad mainly use two methods, namely image features and deep learning, to extract the outlines of buildings in remote sensing images. The research on image features mainly focuses on methods such as the gradient of graphics, the features of building outlines, based on the watershed principle, based on morphological transformation, and using corner detection to extract building outlines. Deep learning mainly uses convolutional neural networks and their extended models to conduct experiments in different frameworks. By modifying and optimizing the models, ideal data results can be obtained. The methods used include extraction based on pure images and extraction through multi-source data fusion. The multi-source data includes different types such as images, point clouds, and DSM. Different methods are used according to the characteristics of the data to obtain different data results. Whether it is machine learning or fully automatic image segmentation, the outline extraction of buildings in the data can be achieved. Summary of the Invention

[0006] The purpose of the present invention is to make up for the above deficiencies and disclose to the society a method for extracting building outlines from aerial images based on machine learning, realizing an efficient and accurate method for making data sets, and at the same time optimizing a machine learning method for extracting building outlines, so as to combine the existing surveying and mapping results and new technologies well and efficiently and accurately make data samples.

[0007] The technical solution of the present invention is realized as follows:

[0008] A method for extracting building outlines from aerial images based on machine learning, comprising the following steps:

[0009] Step 1: Make a custom-resolution aerial image building outline data set;

[0010] Step 2: Input the data set into the Self-net model for training. The Self-net model includes four groups of downsampling blocks and four groups of upsampling blocks. The training process includes:

[0011] Input the data set into the downsampling blocks to achieve data downsampling. Each group of downsampling blocks is connected through a pooling layer to extract features from the data. The data obtained by pooling is used as the input for the next downsampling block to perform multiple consecutive downsamplings on the data; perform an upsampling operation on the data finally output by the downsampling blocks. Each group of upsampling blocks is connected through one convolution, and the output of each layer is used as the input for the next layer to be associated. A closed loop is formed for the model from input to output, and the code blocks are combined to make the model form an end-to-end processing flow;

[0012] Step 3: Verify the result obtained in Step 2, calculate the loss value of the model, and finally compare it with the true value to calculate. Propagate the error back to the model and repeat Step 2, so as to continuously update the model and perform repeated learning;

[0013] Step 4: The final model performs convolutional classification on the data and outputs and displays the data.

[0014] In step 2, the downsampling block consists of three layers: a convolutional layer, an activation layer, and a normalization processing layer. The three layers are combined twice, and the two combinations are arranged in the order of convolutional layer, activation layer, normalization processing layer, convolutional layer, activation layer, and normalization processing layer, and dimensionality reduction processing is performed on the input and output of the model in different code blocks.

[0015] Measures to further optimize this technical solution are:

[0016] As an improvement, when the convolutional layer performs convolutional sampling, it adopts a structure with padding=(1, 1) and stride=(1, 1), and uses a convolutional kernel with kernel_size=(3, 3) to perform pixel-by-pixel convolution.

[0017] As an improvement, in step 2, the pooling process uses a pooling function, and the pooling function used is the max pooling function. The max pooling layer is used to reduce the sensitivity of the convolutional layer to position, so that the weight of position is reduced for data with similar features.

[0018] As an improvement, in step 2, the upsampling block consists of an upsampling layer, a skip connection layer, and a downsampling block. The model realizes image expansion by performing deconvolution on the data, and finally transforms the graph from the feature map block into a graph with the same size as the original input data. Then, the output graph of the upsampling block is skip-connected with the corresponding downsampling block, and after another downsampling block operation, the corresponding multi-level feature convolution result is obtained as the original data for the next operation.

[0019] As an improvement, in step 2, during the upsampling process, the connection layer used in the Self-net model adopts the Relu function as the activation layer, which well preserves the features of the non-linear expression of the data, enabling the model to largely transfer the image features.

[0020] As an improvement, in step 3, in the output link, a method of connecting data from different upsampling layers is adopted to form transitional data with the same size as the input data, so as to participate in the subsequent loss calculation.

[0021] As an improvement, in step 4, the sigmoid cross-entropy function is adopted as the classifier in the output link.

[0022] As an improvement, step 1 includes the following operations:

[0023] S1 Read the original file;

[0024] S1.1 Read the shapefile to obtain the information corresponding to the shp, including the coordinate system information and the coordinate extrema: Xmin, Xmax, Ymin, and Ymax;

[0025] S1.2 Read the original image file to obtain the coordinate system information of the image and the ground resolution gsd of the image;

[0026] S2 Calculate the image size;

[0027] S2.1 Calculate the coordinate system transformation parameters based on the two coordinate systems obtained in S1, and project the shp file onto the coordinate system of the original image to unify the coordinate systems;

[0028] S2.2 Calculate the image size based on the coordinate extrema calculated by the following formula, and construct a blank image:

[0029]

[0030]

[0031] Where:

[0032] W represents the width of the image, that is, the number of horizontal pixels;

[0033] H represents the height of the image, that is, the number of vertical pixels;

[0034] gsd is the ground resolution of the image;

[0035] S3 Obtain the feature objects of the shapefile one by one;

[0036] S3.1 Obtain the shapefile feature object set;

[0037] S3.2 Obtain the feature objects one by one and get the point set corresponding to the feature objects;

[0038] S3.3 Calculate the positions of each point in the point set in the blank image one by one:

[0039]

[0040] Where: the point set is represented by points, Pi ∈ points, Pix represents the x coordinate value of the Pi point, and Piy represents the y coordinate value of the Pi point;

[0041] S3.4 Until all data is drawn;

[0042] S4 Save the corresponding image;

[0043] S4.1 Save the just-drawn sample label image;

[0044] S4.2 Crop the corresponding area of the original image according to the calculated coordinate maximum and minimum values;

[0045] S4.3 Output data according to the defined map size of the dataset.

[0046] As an improvement, the original file includes road, bridge, river, vegetation, and building data.

[0047] The advantages of the present invention compared with the prior art are:

[0048] (1) It realizes the output of the dataset with customizable resolution, and the contour is more regular and neat. Compared with other means, the dataset can be made fully automatically.

[0049] (2) Improve the relevant layer structures and functions of the model, and redesign the forward and backward propagation processes. Compared with the existing model, the number of model parameters is reduced by about 20%.

[0050] (3) The Self-net model mainly adopts the architecture of the MC-FCN model, and makes adjustments on this basis. Reduce one convolution, activation, and normalization process for downsampling in the model respectively, and then use its skip connection layer to enhance the existence of data features. At the same time, perform cross-computation on the multi-layer structure to reduce the loss of data features, so as to ensure the integrity of data features as much as possible. At the same time, use cross-entropy calculation when calculating the loss, output the results of multiple levels as features for cross-validation, and update the parameters by reverse calculation. Description of the Drawings

[0051] Figure 1 It is a flowchart for dataset production;

[0052] Figure 2 It is a flowchart of the Self-net model;

[0053] Figure 3 It is a flowchart of the downsampling block;

[0054] Figure 4 It is a flowchart of the upsampling block;

[0055] Figure 5 It is a flowchart of forward and backward propagation;

[0056] Figure 6 It is an effect diagram of different iteration times;

[0057] Figure 7 It is an accuracy diagram of the Self-net model;

[0058] Figure 8 It is the Self-net structure;

[0059] Figure 9It is a comparison chart of different model parameters;

[0060] Figure 10 It is a line chart of the verification accuracy of different models;

[0061] Figure 11 It is a bar chart of the verification accuracy of different models. Detailed implementation manners

[0062] The present invention will be further described in detail below with reference to the accompanying drawings:

[0063] A method for extracting building outlines from aerial images based on machine learning includes the following steps:

[0064] Step 1: Make a custom-resolution aerial image building outline dataset;

[0065] Step 2: Input the dataset into the Self-net model for training. The Self-net model includes four groups of downsampling blocks and four groups of upsampling blocks. The training process includes:

[0066] Input the dataset into the downsampling blocks to achieve data downsampling. Each group of downsampling blocks is connected through a pooling layer to extract features from the data. The data obtained by pooling is used as the input for the next downsampling block to perform multiple consecutive downsamplings on the data; perform an upsampling operation on the data finally output by the downsampling blocks. Each group of upsampling blocks is connected through one convolution, and the output of each layer is used as the input for the next layer to be associated. A closed loop is formed for the model from input to output, and the basic structure of the model is used to combine the code blocks to make the model form an end-to-end processing flow;

[0067] The basic structure of the model refers to the upsampling blocks and downsampling blocks, which are composed of a series of sub-processes such as convolutional layers, activation layers, and normalization. In the whole model, it is combined according to the process of input, downsampling, upsampling, and output. Among them, the upsampling and downsampling are composed of multiple basic structures. Among them, the upsampling includes three upsampling blocks, and the downsampling includes three downsampling blocks. Each code block (upsampling block, downsampling block) is combined in an arrangement relationship where the output of the previous one is used as the input of the next one.

[0068] Step 3: Verify the result obtained in Step 2, calculate the loss value of the model, and finally compare it with the true value for calculation, and propagate the error back to the model, and repeat Step 2, so as to continuously update the model and perform iterative learning;

[0069] Step 4: Finally, the model performs convolutional classification on the data and outputs and displays the data.

[0070] As Figure 1 shown, the overall process of making the dataset from shapefile data in Step 1 is as follows:

[0071] S1. Read the original file

[0072] The selected data file contains roads, bridges, rivers, vegetation, buildings, and other features, ensuring data diversity from the data aspect. It provides various supports for subsequent data processing and calculations, is conducive to ensuring data reliability, and thus enhances the credibility of the experiment. The building distribution covers multiple types of data, reducing the model application risk caused by data feature limitations.

[0073] S1.1 Read the shapefile and obtain the corresponding information of the shp, including coordinate system information and coordinate extrema: Xmin, Xmax, Ymin, and Ymax;

[0074] S1.2 Read the original image file and obtain the coordinate system information of the image and the ground sampling distance (gsd) of the image;

[0075] S2. Calculate the image size

[0076] S2.1 Calculate the coordinate system transformation parameters based on the two coordinate systems obtained in S1, and project the shp file onto the coordinate system of the original image to unify the coordinate systems;

[0077] S2.2 Calculate the image size based on the coordinate extrema calculated by the following formula and construct a blank image:

[0078]

[0079] Where:

[0080] W represents the width of the image, that is, the number of horizontal pixels;

[0081] H represents the height of the image, that is, the number of vertical pixels;

[0082] gsd is the ground sampling distance of the image;

[0083] S3. Obtain the feature objects of the shapefile one by one

[0084] S3.1 Obtain the feature object set of the shapefile, and the feature object set is represented by Features;

[0085] S3.2 Obtain the feature objects (features) one by one and get the corresponding point sets of the feature objects (features);

[0086] S3.3 Calculate the positions of each point (point) in the point set in the blank image one by one:

[0087]

[0088] Among them: the point set is represented by points, Pi ∈ points, Pix represents the x - coordinate value of point Pi, and Piy represents the y - coordinate value of point Pi;

[0089] S3.4 until all data is drawn;

[0090] S4. Save the corresponding image

[0091] S4.1 Save the just - drawn sample label image;

[0092] S4.2 Crop the corresponding area of the original image according to the calculated coordinate maximum and minimum values;

[0093] S4.3 Output data according to the defined map size of the data set.

[0094] Improved Self - net model:

[0095] By changing the relevant structure of the deep - learning model, the structural organization of the Self - net model is proposed (as Figure 2 shown). The Self - net model mainly adopts the architecture of the MC - FCN model. On this basis, it is adjusted by reducing one convolution, activation, and normalization process for downsampling in the model respectively, and then using its skip connection layer to enhance the existence of data features. At the same time, cross - calculation is performed on the multi - layer structure to reduce the loss of data features, so as to ensure the integrity of data features as much as possible. At the same time, cross - entropy calculation is used when calculating the loss, and the multi - level results are output as features for cross - validation, and the parameters are updated by back - calculation.

[0096] The structure of the Self - net model is mainly composed of four groups of downsampling blocks and four groups of upsampling blocks. In downsampling, each group is connected through max - pooling, and each group of upsampling blocks is connected through one convolution. The output of each layer is used as the input of the next layer for correlation, forming a closed loop from the input to the output of the model. The basic structure is used to combine the model to form an end - to - end processing flow.

[0097] The entire self-net model is mainly composed of two block structures: an upsampling block and a downsampling block. In other parts, simple pooling or activation functions are used to connect various parts of the model and output some of the required data. In the downsampling process, the block structure mainly uses max pooling to extract features from the data, reducing the data from w*h to w*h / 4, thereby achieving the feature transfer of data downsampling. The data obtained by max pooling is used as the input for the next sampling block, and the data is continuously downsampled. In the upsampling block, through the application of transposed convolution to the data, the data is expanded from the original w*h to w*h*4, enabling the model to achieve image expansion. Finally, the graph is transformed from the feature map block into a graph with the same size as the original input data, so as to calculate the loss function and output the results.

[0098] In this process, the pooling function used is the max pooling function, and the max pooling layer is used to reduce the sensitivity of the convolutional layer to position, so that the weight of the position is reduced for data with similar features. For the activation functions used in the Self-net model of the present invention, they are all ReLu functions, which perform non-linear processing on the data, making it maintain the characteristics of non-linear functions, while reducing the possibility of discreteness, so that the model can extract features well.

[0099] (1) Self-net downsampling block

[0100] The downsampling block used in the Self-net model is mainly composed of three layers: convolution, activation, and normalization processing. These three layers are combined twice, and the two combinations are arranged in the order of convolutional layer, activation layer, normalization processing layer, convolutional layer, activation layer, and normalization processing layer, and dimensionality reduction processing is performed on the input and output of the model in different code blocks; the structure is as shown in Figure 3 the figure. The data is input into the upsampling block in the model to achieve data downsampling and obtain data features at different levels. In convolutional sampling, to ensure the consistency of data size, a structure with padding=(1, 1) and stride=(1, 1) is adopted to ensure the consistency of data during the convolution process. The entire block structure does not change the size of the input data, but only extracts the features of the data, and a convolutional kernel with kernel_size=(3, 3) is used to perform pixel-by-pixel convolution to ensure the acquisition of eight-directional features of the data.

[0101] In the downsampling block of the Self-net model, the parameter changes and related data of different layers are shown in Table 1 below. According to the different parameter changes, the theoretical data volume changes can be clearly seen, and at the same time, it is ensured whether the data input and output before model training conform to the design and whether the results can be calculated normally.

[0102] Table 1 Relationships between layers of the downsampling block

[0103]

[0104] Among them, h and w respectively represent the height and width of the input layer. *The normalization layer _2 is connected to the upsampling layer of the upsampling block.

[0105] (2) Self-net upsampling block

[0106] In the Self-net model, the second main part is the upsampling block, which is composed of an upsampling layer, a skip connection layer, and a downsampling block. Through an upsampling operation on the data, the input of the graph is expanded to 4 times the size of the input graph. Then, the output graph of the upsampling block is skip-connected to the corresponding downsampling block. After that, through two operations of convolution, activation, and normalization (i.e., a downsampling block), the corresponding multi-level feature convolution results are obtained, which are used as the original data for the next operation. The upsampling operation performed in this upsampling block adopts an extended convolution mode to ensure that the input data is consistent with the downsampled data of the corresponding layer, so as to perform a skip connection calculation on the two data. The structure is as Figure 4 shown. In addition, other convolutional layers all adopt the convolution mode with the same input and output size to ensure the consistency of data before and after, reduce the loss of data during convolution, and facilitate data operation at the same time.

[0107] The parameter changes of different layers in the upsampling block are shown in Table 2 below. In different layers of the upsampling block, the parameter changes are different. It can be seen from the upsampling that the data volume of the model expands, which has an intuitive display for the cross-operation of different layers of the model. The number of data can be used to check whether the two-layer data can perform skip connection calculation to ensure the consistency of data and the normal operation of the program.

[0108] Table 2 Relationships of each layer in the upsampling block

[0109]

[0110] Among them: h', w', and d respectively represent the height, width, and depth of the input layer. *The normalization layer _2' generates predictions.

[0111] (3) Self-net forward and backward propagation

[0112] After the main parts of the Self-net model are sorted out, the subsequent focus is to input data into the model and associate each part, mainly to enable the data to be operated normally, so that the model can perform automatic batch input of data, automatic gradient calculation, and continuous update of model parameters, as Figure 5The model is organized and calculated as shown. Each part mainly combines the input and output parameter requirements of different layers through a process-based operation to combine the model, combines the forward and backward propagation of data, mainly combines the above-mentioned downsampling, upsampling, and backpropagation. In downsampling, the pooling layer is used to connect to ensure the reduction of data. In upsampling, the model is expanded with data through convolutional upsampling, and then the data is connected to calculate the loss value of the model. Finally, it is compared with the true value, and the error is propagated back to the model, so as to continuously update the model, perform repeated learning, and finally the model performs sigmoid convolution classification on the data and outputs and displays the data.

[0113] From Figure 5 It can be seen from the flowchart in that the model is basically in an end-to-end black box state in data processing. Only when the data is backpropagated, that is, the gradient is updated or the loss value is calculated, the parameters of the model are updated and replaced. Different parameters will affect the convolution operation result of the data, thus causing changes in the output results of different layers of the model. In the model of the present invention, cross-entropy is mainly used as a loss calculation, the difference between the calculated true value and the predicted value is calculated, and then it is propagated back to each parameter according to the gradient, so as to realize the replacement and iteration of the parameters, so as to ensure the continuous change of the data and the continuous optimization of the model.

[0114] After the data verification of the model, the data set is input into the Self-net model, and combined with the initialization parameters of the above model training, the processes of data input, model training, data verification, etc. are executed in sequence, and the relevant results of the operation are output and the optimal parameters during the model operation are saved. Figure 6 They are the original image, label image, and predicted image with the best prediction accuracy in different iteration times when the Self-net model trains a large range of data.

[0115] From Figure 6 It can be seen that when the model is trained, there is a certain error between the predicted result and the sample label. However, as the number of iterations increases, the model corrects some errors in the prediction, and the prediction ability of the model is improved to a certain extent. However, some details are still insufficient, and the positive training data of the model needs to be continuously improved.

[0116] The three models of FCN32s, Unet, and MC-FCN are used as reference models and compared with the Self-net model of the present invention.

[0117] Figure 7 It shows the change of the prediction accuracy and loss accuracy of the Self-net model in large and small range data, indicating the change of the Self-net model in different cycles, and overall conforms to the overall expectation of the previously designed model. FromFigure 7 As can be seen from the calculation results, during the training process, the prediction accuracy of the Self-net model is continuously improving. As the number of iterations increases, the accuracy improves significantly at first and then gradually shows a slow growth trend, with the growth rate becoming smaller and the growth gradient gradually tending to 0, indicating that the accuracy of the model tends to be stable after training to a certain extent, which conforms to the general law of machine learning. By observing Figure 7 the loss function in, it can be found that the gradient descent in the model training slows down from fast to slow, indicating that the initial parameters of the model are updated relatively quickly, and then the update rate slows down. However, there is no high-low fluctuation like that in the test set, indicating that the gradient update during the training process does not show a diffusion state, and the change of model parameters is a normal phenomenon.

[0118] Compared with the basic model MC-FCN, the Self-net model changes the number of layers in downsampling. In the downsampling block, the original model uses three convolutional, activation, and normalization operations, while the model of the present invention uses two of the above operations. The goal is to reduce the loss of data features during the convolution process and ensure that the data can be well utilized during feature extraction. Secondly, during the upsampling process, the connection layer used by the Self-net model uses the Relu function as the activation layer, which well preserves the features of the non-linear expression of the data, enabling the model to transfer the image features to a large extent. Finally, the model adopts a method of connecting the data of different upsampling layers at the output link to form transitional data with the same size as the input data, so as to participate in the subsequent loss calculation. The purpose is to ensure that multi-level features participate in the calculation, reduce the loss of data features, and use the data layer by layer for good expression. After the improvement, all sub-parts participate in backpropagation and have an impact on the output results, thus avoiding the serious loss of data features. At the output link, the sigmoid cross-entropy function is used as the classifier, as Figure 8 shown, a holistic improvement is made to the model. Compared with the FCN32S model, the difference lies in adding skip connections and multi-level features of the result output participating in the loss calculation. At the same time, in terms of the number of model layers, fewer layers are used to process data; compared with the Unet model, the main improvement is in the multi-level cross-entropy calculation of the model upsampling output and the input data. The specific process of the Self-net model is shown in Figure 8 as shown.

[0119] When the number of layers of the Self-net model decreases, the number of weights and biases involved in the model also decreases as the number of layers decreases, which will change both the forward propagation and backward propagation of the model. Under the standard of input data size of 224*224*3, the number of parameters of Fcn32s is 98,253,497, Unet is 3,408,553, MC-FCN is 3,414,028, and Self-net is 2,749,361, as shown in Table 3. It can be found that when the input data is the same, the number of parameters of the entire Self-net model has decreased significantly. Compared with the previous models, it can be found that the number of parameters of the Self-net model is significantly less. As the number of parameters decreases, the resources consumed in the entire calculation process will also decrease, thus reducing the hardware configuration requirements of the model, that is, the corresponding memory occupancy also decreases with the reduction of the data volume.

[0120] Table 3 Comparison of the number of parameters of different models

[0121]

[0122] In Figure 9 it can be clearly seen the parameter situations of different models. In the comparison of the number of model parameters, the change in the number of parameters of the Self-net model is relatively obvious. The number of parameters of the Self-net model is basically less than one-thirtieth of FCN32S and about two-thirds of the Unet and MC-FCN models. It can be seen that the number of parameters of the Self-net model has decreased significantly. The number of model parameters is conducive to the fast calculation of the forward and backward propagation of data and less resource sacrifice.

[0123] By observing Table 4, it can be found that the time required for the model to complete the processing of input data to output prediction (i.e., a single epoch) is inconsistent. The original model MC-FCN processes one epoch in 13 minutes for small-scale data, and Self-net processes one epoch in 14 minutes for small-scale data. MC-FCN processes one epoch in 3 hours and 6 minutes for large-scale data, and Self-net processes one epoch in 3 hours and 14 minutes for large-scale data. From the perspective of time, the efficiency of each epoch is inconsistent. As the data volume increases, it can be found that there is a process of time increase for a single epoch of the Self-net model. As the data volume increases, the computing power of the Self-net model will be inferior to the original model. Theoretically analyzed, mainly because the calculation amount of its subsequent loss function increases. When calculating the loss of the modified loss function, it changes from the original cross-entropy calculation of two images to the calculation of multi-level images and the ground truth, which increases a part of the calculation amount, so the efficiency after improvement is reduced.

[0124] Table 4 Completion time of a single epoch

[0125]

[0126] During the entire experiment, the results of the overall data processing efficiency of the model are shown in Table 5. It can be seen from the inference speed (FPS) of the model that during the training process, the Self-net model is relatively slow in terms of efficiency, while during the verification process, the number of processed data is basically the same as that of the existing models. The reason for this inconsistency lies in the different reverse calculation parameters of the model, which has an impact on the model's learning ability. The speed of the same training data is mainly affected by the complexity of the data's backpropagation.

[0127] Table 5 Inference speed of the model

[0128]

[0129] Figure 10 is the accuracy graph after different models complete data verification in different data ranges. From Figure 10 it can be seen that in a small range of data, the verification accuracy of the Self-net model is superior to other models in the initial stage, and the accuracy gradually decreases in the later stage, indicating that the model is moving in a negative direction during the gradual learning process and has not been corrected. In the results of a large range of data, the verification accuracy of the Self-net model is higher than that of other models throughout the process. Even if there is a decrease once, it is still better than other models. It can be explained that the Self-net model has characteristic small-range model learning due to insufficient data, which has a certain directional impact on the model. While a large range of data has more data features, even if the model shows a negative shift, it can be corrected to repair the model, enabling the model to maintain a high accuracy, thereby enhancing the stability of the model. Generally speaking, the Self-net model demonstrates strong stability in the verification process when the feature learning is rich. The accuracy after training with different data volumes depends on the degree of feature learning.

[0130] The quality of the verification accuracy determines the model's ability to process new data, that is, whether the model can distinguish new data well and the model's prediction ability during the application process. It can be found from Table 6 that the verification accuracies of different models are inconsistent in a small range of data compared with those in a large range of data. The differences in the verification accuracies of different models in a small range are not significant, while in the comparison of the accuracies of models in a large range, the abilities of the models are different, mainly showing the high or low abilities of the models in learning training data and predicting new data.

[0131] Table 6 Verification accuracies of different models

[0132]

[0133] From the following barFigure 11 It can be seen that the application ability of the Self-net model in small-scale data is the same as that of the existing model, but its application ability in large-scale data is better than that of the existing model, indicating that the performance of the Self-net model supported by big data will be better than the application of the original model.

[0134] The overall effect of the Self-net model in the dataset is better than that of the existing model (MC-FCN model). After the model training and verification are completed, the loss function, training accuracy, and verification accuracy of each model are compared and analyzed. The results show that the improved model also belongs to the convergent state. Under the dataset and initial parameters, the initial accuracy is relatively high, and a relatively high training accuracy can be achieved within a small number of epochs; in terms of verification accuracy, the two are basically the same in the small range of 3,888 data pairs, but when dealing with the large-scale data of 59,414 data pairs, the improved model is 4.2% better than the existing model, indicating that the Self-net model has better effects on large-scale data.

[0135] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention accordingly. For those skilled in the art, it should be able to realize that all the equivalent replacements and obvious changes made by using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for extracting building outlines from aerial images based on machine learning, characterized in that: It includes the following steps: Step 1: Create a custom resolution aerial image building contour dataset, including the following steps: S1 Read the original file; S1.1 Read the shapefile to obtain the information corresponding to the shp, including coordinate system information and coordinate extrema: Xmin, Xmax, Ymin, and Ymax; S1.2 Read the original image file to obtain the coordinate system information of the image and the ground resolution gsd of the image; S2 Calculate the image size; S2.1 Calculate the coordinate system conversion parameters based on the two coordinate systems obtained in S1, project the shp file onto the coordinate system of the original image to unify the coordinate systems; S2.2 Calculate the image size based on the calculated coordinate extrema and construct a blank image: Where: W represents the width of the image, that is, the number of horizontal pixels; H represents the height of the image, that is, the number of vertical pixels; gsd is the ground resolution of the image; S3 Obtain the feature objects of the shapefile one by one; S3.1 Obtain the shapefile feature object set; S3.2 Obtain the feature objects one by one and get the corresponding point set of the feature objects; S3.3 Calculate the positions of each point in the point set in the blank image one by one: Where: the point set is represented by points, Pi ∈ points, Pix represents the x coordinate value of point Pi, and Piy represents the y coordinate value of point Pi; S3.4 Until all data is drawn; S4 Save the corresponding image; S4.1 Save the just-drawn sample label image; S4.2 Crop the corresponding area of the original image according to the calculated coordinate extrema; S4.3 Output data according to the defined dataset map size; Step 2: Input the dataset into the Self-net model for training. The Self-net model includes four groups of downsampling blocks and four groups of upsampling blocks. The downsampling block consists of three layers: a convolutional layer, an activation layer, and a normalization processing layer. The upsampling block consists of an upsampling layer, a skip connection layer, and a downsampling block. The training process includes: Input the dataset into the downsampling block to achieve data downsampling. Each group of downsampling blocks is connected through a pooling layer to extract features from the data. Use the data obtained by pooling as the input for the next downsampling block to perform multiple consecutive downsamplings on the data; perform an upsampling operation on the data finally output by the downsampling block. Each group of upsampling blocks is connected through a single convolution. Associate the output of each layer as the input for the next layer to form a closed loop from the input to the output of the model, and combine the code blocks to make the model form an end-to-end processing flow; Step 3: Verify the result obtained in Step 2, calculate the loss value of the model, finally compare it with the true value, propagate the error back to the model, and repeat Step 2 to continuously update the model and perform iterative learning; Step 4: Finally, the model performs convolutional classification on the data and outputs and displays the data.

2. The method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the second step, the downsampling block consists of three layers: a convolutional layer, an activation layer, and a normalization layer. The three layers are combined twice, and the two combinations are arranged in the order of convolutional layer, activation layer, normalization layer, convolutional layer, activation layer, and normalization layer, and dimensionality reduction processing is performed on the input and output of the model in different code blocks.

3. A method for extracting building outlines from aerial images based on machine learning according to claim 2, characterized in that: When performing convolutional sampling, the convolutional layer adopts a structure with padding=(1, 1) and stride=(1, 1), and a convolutional kernel with kernel_size=(3, 3) is used for pixel-by-pixel convolution.

4. A method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the second step, the pooling process uses a pooling function, and the pooling function used is the max pooling function.

5. A method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the second step, the upsampling block consists of an upsampling layer, a skip connection layer, and a downsampling block. The model realizes image expansion by performing deconvolution on the data, and finally transforms the graph from the feature map block into a graph with the same size as the original input data. Then, the output graph of the upsampling block is skip-connected to the corresponding downsampling block, and after another downsampling block operation, the corresponding multi-level feature convolution result is obtained as the original data for the next operation.

6. The method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the second step, during the upsampling process, the connection layer used in the Self-net model adopts the Relu function as the activation layer.

7. A method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the third step, in the output link, a method of connecting data from different upsampling layers is adopted to form intermediate data with the same size as the input data to participate in the subsequent loss calculation.

8. A method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: In the fourth step, in the output link, the sigmoid cross-entropy function is adopted as the classifier.

9. A method for extracting building outlines from aerial images based on machine learning according to claim 1, characterized in that: The original file contains data on roads, bridges, rivers, vegetation, and building structures.

Citation Information

Patent Citations

  • High-resolution image building extraction method based on multi-scale residual network model

    CN113205018A