A method for extracting impervious surfaces from high-resolution remote sensing images based on weakly supervised learning

By combining the labeling samples of high-score and medium-score remote sensing images, the training data set is expanded, and weakly supervised learning method is adopted, the problem of insufficient labeling of impermeable surface data of high-score remote sensing images is solved, and high-precision impermeable surface extraction is achieved, reducing the labeling cost and training cycle.

CN114581729BActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210170422.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-05-13
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The existing technology has fewer data labels for impermeable surfaces of high-score remote sensing images and too long manual labeling of image labels, resulting in low extraction accuracy and cannot meet the needs of massive remote sensing data processing in the big data era.

Method used

By combining a small number of high-score remote sensing images with fine labeling samples and a large number of coarse-grained medium-score remote sensing images, the training data set is expanded, and weakly supervised learning method is adopted to train using convolutional neural networks to gradually improve the model accuracy.

Benefits of technology

It achieves a significant saving of impermeable surface labeling costs while ensuring accuracy, shortens the cycle of training data generation, and improves the automation level and efficiency of impermeable surface extraction of high-score remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581729B_ABST
    Figure CN114581729B_ABST
Patent Text Reader

Abstract

A high-resolution impervious surface extraction method based on weakly supervised deep learning, firstly determines the remote sensing ground object target according to the extraction task and produces a small number of accurate high-resolution remote sensing image sample labels and a large number of coarse-grained medium-resolution remote sensing image sample labels. Then, a small sample set impervious surface model is trained using a small number of accurate sample labels. Then, a large number of high-resolution unlabeled images are predicted using the small sample set impervious surface model to obtain a large number of high-resolution remote sensing image impervious surface pseudo labels, from which a large number of more accurate high-resolution remote sensing image sample labels combined with a small number of accurate sample labels are selected as new training sets, and a new impervious surface model is generated through multiple cycles of training until the model accuracy is no longer improved, and the final sample labels are predicted by the final model, and then a large sample set impervious surface model is generated by combining a large number of medium-resolution remote sensing impervious surface weak label data for training, and finally, the large sample set impervious surface model is used to predict the high-resolution remote sensing image to be produced to obtain the final impervious surface extraction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for extracting impervious surfaces from high-resolution remote sensing images based on weakly supervised learning, and relates to the fields of remote sensing, neural networks, and weakly supervised learning. Background Art

[0002] Impervious surface is an important indicator of urban ecological environment and key data for understanding urban environment. It plays an important guiding role in urbanization process, environmental quality assessment and the development of the relationship between man and nature.

[0003] At present, the methods for extracting impervious surfaces mainly include index method, supervised classification method, regression-based method and threshold-based method, etc. However, due to the high complexity of impervious surface objects and the limitations of classification accuracy and processing speed on the operator's technical level and experience knowledge, the overall feature extraction automation level is poor; and the extraction efficiency cannot meet the needs of massive remote sensing data processing in the big data era. With the development of deep learning computer vision, it has achieved a series of breakthrough results in image classification, semantic segmentation, target detection and other fields. Therefore, deep learning algorithms have gradually been applied to remote sensing image information extraction. The use of deep learning methods can solve the problems of complex objects and large data volume in high-resolution remote sensing images, which is of great significance for the study of urban impervious surfaces. However, the premise for the remote sensing image impervious surface extraction method based on deep learning to achieve ideal accuracy is to have high-quality and large-scale data sets, but the number of high-resolution impervious surface remote sensing image samples is limited and the annotation cost is high. It is difficult to obtain remote sensing image label data sets that meet quality requirements. The development of weakly supervised deep learning provides a new solution for extracting impervious surfaces from remote sensing images. Compared with fully supervised deep learning, the supervised information it uses is not complete, that is, not all samples need to have labels or the labels are not necessarily accurate. Therefore, weakly supervised deep learning technology can be applied to extracting impervious surfaces. Under the premise of ensuring accuracy, it can greatly save the cost of labeling impervious surfaces and shorten the cycle of training data generation. Summary of the invention

[0004] The present invention aims to solve the shortcomings of the prior art in that the impervious surface data labels of high-resolution remote sensing images are relatively few, the time for manual image labeling is too long, and the impervious surface extraction accuracy of high-resolution remote sensing images is low, and a high-resolution remote sensing impervious surface extraction method based on weak supervised learning is provided.

[0005] The present invention expands the training data set by combining a small number of high-resolution remote sensing image fine-labeled samples and a large number of coarse-grained medium-resolution remote sensing image labelled samples, thereby improving the model accuracy.

[0006] A high-resolution remote sensing impervious surface extraction method based on weakly supervised learning of the present invention comprises the following steps:

[0007] Step 1: Train a convolutional neural network with a small amount of accurate high-resolution image sample label data; select a suitable semantic segmentation network model for training based on the characteristics of high-resolution remote sensing images. The steps are as follows:

[0008] Step 1.1: Select a semantic segmentation network model suitable for high-resolution remote sensing image object extraction. Here, the D-LinkNet semantic segmentation network model is selected.

[0009] Step 1.2: Create high-resolution remote sensing image samples. Select 10% to 20% of the high-resolution remote sensing image data set as a small sample data set. Manually perform fine annotation of the impervious surface areas in the small sample data set to generate a small number of accurate impervious surface label images.

[0010] Step 1.3: Perform data enhancement on the small number of accurately labeled impervious surface images obtained in step 1.2. Here, random horizontal flipping, vertical flipping, and noise addition are used to reduce the overfitting problem of the neural network. The same operation is performed on the masks of the training set.

[0011] Step 1.4: Select a suitable loss function and network optimizer during training. Select Dice Loss and Bce Loss to form the loss function. The definition of Dice Loss is:

[0012]

[0013] Where |X∩Y| is the intersection of X and Y, |X| and |Y| represent the number of elements of X and Y respectively, and the coefficient of the numerator is 2 because the denominator repeats the calculation of the common elements between X and Y.

[0014] Bce Loss is defined as:

[0015] l n =-w n [y n ·logx n +(1-y n )·log(1-x n )] (2)

[0016] In the formula, x n Indicates the actual predicted value of the nth sample, y n represents the actual label, w n is a hyperparameter, l n Indicates the loss corresponding to the nth sample.

[0017] The neural network optimizer uses the adaptive moment estimation optimizer (Adam) to train the network parameters.

[0018] After that, the network weights are updated for each epoch, and the neural network training is accelerated by setting early stopping and adjusting the learning rate to converge faster.

[0019] Step 1.5: Perform training to obtain a neural network training model.

[0020] Step 2: Obtain a large number of high-resolution mask images;

[0021] Use the test set of high-resolution image data and other high-resolution uncalibrated high-resolution image data as input to the trained network, select high-confidence samples from the prediction results, use the labeled data and pseudo-labels in step 1 to train a new model, replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, use the final model to predict the remaining unknown label data, and get the final label. The specific steps are as follows:

[0022] Step 2.1: Load the neural network model trained in step 1.

[0023] Step 2.2: Read the test data set and split the large image into rows and columns.

[0024] Step 2.3: Obtain the results through forward propagation, and output and save the results step by step.

[0025] Step 2.4: Filter high-confidence samples from the prediction results and train a new model using the labeled data in step 1 and the pseudo-labels.

[0026] Step 2.5: Replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, and use the final model to predict the remaining unknown label data to obtain the final label.

[0027] Step 3: Use the large number of high-resolution mask images obtained in step 2 combined with a large number of medium-resolution impervious surface remote sensing images with weakly supervised label data to train a more accurate convolutional neural network model;

[0028] According to the geographic coordinates of the high-resolution image used in step 1 and step 2 and the correspondence with the medium-resolution image, the pixel coordinates of the high-resolution image are obtained by row and column, the medium-resolution image at the corresponding position is cropped, and then the result is used as a new input to retrain a neural network of high-resolution image data to predict the impervious surface of the high-resolution image. The specific steps are as follows:

[0029] Step 3.1: Read in the high-resolution mask image obtained in step 1 and step 2.

[0030] Step 3.2: Use the tool class method of the gdal library function to obtain the corresponding coordinates according to rows and columns.

[0031] Step 3.3: Crop the medium-resolution remote sensing image data at the corresponding position according to the coordinate correspondence between the high-resolution image and the medium-resolution image.

[0032] Step 3.4: Obtain the coordinates of the center point of the impervious surface of the high-resolution remote sensing image, mark it on the medium-resolution remote sensing image data according to the coordinate correspondence, and then use the bounding box to mark the impervious surface area of ​​the medium-resolution remote sensing image to generate a large number of weakly supervised labels for the impervious surface of the medium-resolution remote sensing image.

[0033] The obtained medium-resolution remote sensing image data is combined with a large number of high-resolution mask images obtained in step 1 and step 2 to train and generate a high-resolution impervious surface extraction model, thereby realizing the automatic extraction of impervious surfaces in high-resolution remote sensing images.

[0034] Preferably, the high-resolution image described in step 3 has one pixel per meter, while the medium-resolution image has one pixel per 30 meters, so one pixel on the medium-resolution image corresponds to 900 pixels on the high-resolution image.

[0035] The present invention mainly proposes a new impervious surface extraction algorithm for medium-resolution and high-resolution collaborative remote sensing images based on weakly supervised deep learning. Different from the relatively simple use of high-resolution remote sensing image data for impervious surface extraction, the main principle of this algorithm is to use weakly supervised learning, which only requires a small amount of finely labeled high-resolution remote sensing impervious surface labels, combined with a large number of coarse-grained medium-resolution remote sensing impervious surface labels, and the medium-resolution image data is cropped according to the correspondence between geographic coordinates, thereby expanding the training set.

[0036] The present invention has the following advantages:

[0037] (1) A deep learning convolutional neural network model is used. Compared with other traditional algorithms, convolutional neural networks have greatly improved the accuracy of extraction and reduced the intervention of human factors. By adjusting the parameters through network self-learning, the subjectivity of manually adjusting a series of index thresholds is avoided. Neural networks can also replace us to extract unknown feature values ​​in images, especially in the extraction of center-resolution remote sensing images, which have seven-band data including infrared band and ultraviolet band instead of ordinary three-band RGB images. This algorithm can be used to extract feature values ​​in these bands for prediction.

[0038] (2) With the continuous development of computers and the emergence of deep learning frameworks, relying on the powerful computing GPU, the training and optimization of deep learning neural network models can be controlled in a relatively short time, which is more efficient than manual adjustment.

[0039] (3) The present invention adopts a weakly supervised learning method and effectively utilizes the coarse-grained annotation method, which can achieve good high-resolution remote sensing impervious surface extraction results with only a small amount of manual annotation time and cost. In addition, the present invention makes extensive use of coarse-grained manually annotated samples with a wide distribution range, expands the research area, and improves the robustness of the deep learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0041] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0042] The deep learning-based collaborative high-resolution remote sensing image impervious surface extraction algorithm of the present invention comprises the following steps:

[0043] Step 1: Train a convolutional neural network with a small amount of accurate high-resolution image sample label data; select a suitable semantic segmentation network model for training based on the characteristics of high-resolution remote sensing images. The steps are as follows:

[0044] Step 1.1: Select a semantic segmentation network model suitable for high-resolution remote sensing image object extraction. Here, the D-LinkNet semantic segmentation network model is selected.

[0045] Step 1.2: Create high-resolution remote sensing image samples. Select 10% to 20% of the high-resolution remote sensing image data set as a small sample data set. Manually perform fine annotation of the impervious surface areas in the small sample data set to generate a small number of accurate impervious surface label images.

[0046] Step 1.3: Perform data enhancement on the small number of accurate impervious surface label images obtained in step 1.2. Here, we mainly use methods such as random horizontal flipping, vertical flipping, and adding noise to reduce the overfitting problem of the neural network. It should be noted that the same operation needs to be performed on the masks of the training set at the same time.

[0047] Step 1.4: Select a suitable loss function and network optimizer during the training process. In the present invention, Dice Loss and Bce Loss are selected to form the loss function. The definition of Dice Loss is:

[0048]

[0049] Where |X∩Y| is the intersection of X and Y, |X| and |Y| represent the number of elements of X and Y respectively, and the coefficient of the numerator is 2 because the denominator repeats the calculation of the common elements between X and Y.

[0050] Bce Loss is defined as:

[0051] l n =-w n [y n ·logx n +(1-y n )·log(1-x n )] (2)

[0052] In the formula, x n Indicates the actual predicted value of the nth sample, y n represents the actual label, w n is a hyperparameter, l n Indicates the loss corresponding to the nth sample.

[0053] The neural network optimizer uses the adaptive moment estimation optimizer (Adam) to train the network parameters.

[0054] After that, the network weights are updated for each epoch, and the neural network training is accelerated by setting early stopping and adjusting the learning rate to converge faster.

[0055] Step 1.5: Perform training to obtain a neural network training model.

[0056] Step 2: Obtain a large number of high-resolution mask images;

[0057] Use the test set of high-resolution image data and other high-resolution uncalibrated high-resolution image data as input to the trained network, select high-confidence samples from the prediction results, use the labeled data and pseudo-labels in step 1 to train a new model, replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, and use the final model to predict the remaining unknown label data to obtain the final label. The specific steps are as follows:

[0058] Step 2.1: Load the neural network model trained in step 1.

[0059] Step 2.2: Read the test data set and split the large image into rows and columns.

[0060] Step 2.3: Obtain the results through forward propagation, and output and save the results step by step.

[0061] Step 2.4: Filter high-confidence samples from the prediction results and train a new model using the labeled data in step 1 and the pseudo-labels.

[0062] Step 2.5: Replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, and use the final model to predict the remaining unknown label data to obtain the final label.

[0063] Step 3: Use the large number of high-resolution mask images obtained in step 2 combined with a large number of medium-resolution impervious surface remote sensing images with weakly supervised label data to train a more accurate convolutional neural network model;

[0064] According to the geographic coordinates of the high-resolution image used in step 1 and step 2 and the correspondence with the medium-resolution image, the pixel coordinates of the high-resolution image are obtained by row and column, the medium-resolution image at the corresponding position is cropped, and then the result is used as a new input to retrain a neural network of high-resolution image data to predict the impervious surface of the high-resolution image. The specific steps are as follows:

[0065] Step 3.1: Read in the high-resolution mask image obtained in step 1 and step 2.

[0066] Step 3.1: Use the tool class method of the gdal library function to obtain the corresponding coordinates according to rows and columns.

[0067] Step 3.3: The high-resolution image used here has one pixel per meter, while the medium-resolution image has one pixel per 30 meters. Therefore, one pixel on the medium-resolution image corresponds to 900 pixels on the high-resolution image. The medium-resolution remote sensing image data at the corresponding position is cropped according to the coordinate correspondence.

[0068] Step 3.4: Obtain the coordinates of the center point of the impervious surface of the high-resolution remote sensing image, mark it on the medium-resolution remote sensing image data according to the coordinate correspondence, and then use the bounding box to mark the impervious surface area of ​​the medium-resolution remote sensing image to generate a large number of weakly supervised labels for the impervious surface of the medium-resolution remote sensing image.

[0069] The obtained medium-resolution remote sensing image data is combined with a large number of high-resolution mask images obtained in step 1 and step 2 to train and generate a high-resolution impervious surface extraction model, thereby realizing the automatic extraction of impervious surfaces in high-resolution remote sensing images.

[0070] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for extracting impervious surfaces from high-resolution remote sensing images based on weakly supervised learning, comprising the following steps: Step 1: Train a convolutional neural network with a small amount of accurate high-resolution image sample label data; select a suitable semantic segmentation network model for training based on the characteristics of high-resolution remote sensing images. The steps are as follows: Step 1.1: Select a semantic segmentation network model suitable for high-resolution remote sensing image object extraction. Here, the D-LinkNet semantic segmentation network model is selected; Step 1.2: Prepare high-resolution remote sensing image samples. Select 10% to 20% of the high-resolution remote sensing image images from the high-resolution remote sensing image dataset as a small sample dataset. Manually perform fine annotation on the impervious surface areas in the small sample dataset to generate a small number of accurate impervious surface label images. Step 1.3: Perform data enhancement on the small number of accurate impervious surface label images obtained in step 1.

2. Here, random horizontal flipping, vertical flipping, and noise addition are used to reduce the overfitting problem of the neural network. The same operation is performed on the masks of the training set. Step 1.4: Select a suitable loss function and network optimizer during training. Select Dice Loss and BceLoss to form the loss function. The definition of Dice Loss is: Where |X∩Y| is the intersection of X and Y, |X| and |Y| represent the number of elements of X and Y respectively, where, The coefficient of the numerator is 2 because the denominator contains repeated counting of common elements between X and Y; Bce Loss is defined as: l n =-w n [y n ·logx n +(1-y n )·log(1-x n )] (2) In the formula, x n Indicates the actual predicted value of the nth sample, y n represents the actual label, w n is a hyperparameter, l n Indicates the loss corresponding to the nth sample; The neural network optimizer uses the adaptive moment estimation optimizer (Adam) to train the network parameters; After that, the network weights are updated for each epoch, and the neural network training is accelerated by setting early stopping and adjusting the learning rate to converge faster. Step 1.5: Perform training to obtain a neural network training model; Step 2: Obtain a large number of high-resolution mask images; Use the test set of high-resolution image data and other high-resolution uncalibrated high-resolution image data as input, select high-confidence samples from the prediction results, use the labeled data and pseudo-labels in step 1 to train a new model, replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, use the final model to predict the remaining unknown label data, and obtain the final label. The specific steps are as follows: Step 2.1: Load the neural network model trained in step 1; Step 2.2: Read the test data set and split the large image into rows and columns; Step 2.3: Get the result through forward propagation, and output and save the result step by step; Step 2.4: Filter high-confidence samples from the prediction results and train a new model using the labeled data and pseudo labels from step 1; Step 2.5: Replace the model generated in step 1 with the new model, repeat the above steps until the model effect is no longer improved, and use the final model to predict the remaining unknown label data to obtain the final label; Step 3: Use the large number of high-resolution mask images obtained in step 2 combined with a large number of medium-resolution impervious surface remote sensing images with weakly supervised label data to train a more accurate convolutional neural network model; According to the geographic coordinates of the high-resolution image used in step 1 and step 2 and the correspondence with the medium-resolution image, the pixel coordinates of the high-resolution image are obtained by row and column, the medium-resolution image at the corresponding position is cropped, and then the result is used as a new input to retrain a neural network of high-resolution image data to predict the impervious surface of the high-resolution image. The specific steps are as follows: Step 3.1: Read the high-resolution mask image used in step 1 and obtained in step 2; Step 3.2: Use the tool class method of the gdal library function to obtain the corresponding coordinates according to rows and columns; Step 3.3: Crop the corresponding position of the medium-resolution remote sensing image data according to the coordinate correspondence between the high-resolution image and the medium-resolution image; Step 3.4: Obtain the coordinates of the center point of the impervious surface of the high-resolution remote sensing image, mark it on the medium-resolution remote sensing image data according to the coordinate correspondence, and then use the boundingbox to mark the impervious surface area of ​​the medium-resolution remote sensing image to generate a large number of weakly supervised labels for the impervious surface of the medium-resolution remote sensing image; The obtained medium-resolution remote sensing image data is combined with a large number of high-resolution mask images obtained in step 1 and step 2 to train and generate a high-resolution impervious surface extraction model, thereby realizing the automatic extraction of impervious surfaces in high-resolution remote sensing images.

2. The method for extracting impervious surfaces from high-resolution remote sensing images based on weakly supervised learning as claimed in claim 1, characterized in that: The high-resolution image described in step 3 has one pixel per meter, while the medium-resolution image has one pixel per 30 meters. Therefore, one pixel on the medium-resolution image corresponds to 900 pixels on the high-resolution image.

Citation Information

Patent Citations

  • Mobile phone light guide plate defect visual detection method based on multi-task learning network

    CN111951249A

  • High-resolution remote sensing target boundary extraction method based on weak supervised learning

    CN112084871A