Remote Sensing Image Urban Building Extraction Method Based on Shadow Compensation and U-net

By performing shadow detection and compensation in remote sensing images, combined with U-net network model, the problem of building edge extraction under shadow interference is solved, and the accuracy and accuracy of building extraction is improved.

CN114005042BActive Publication Date: 2025-06-27QINGDAO HAIDEMAN PHOTOELECTRIC TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111221175.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2025-06-27
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

In remote sensing images, it is difficult to extract the edge of the building intact due to shadow interference, and traditional methods are difficult to avoid the problem of missed divisions, which affects the accuracy of building extraction.

Method used

Shadow compensation and U-net methods are used to detect and compensate shadows through the spectral characteristics and brightness characteristics of the image to reduce the impact of shadows on building edge extraction, and the building profile is extracted in combination with the U-net network model.

Benefits of technology

It effectively reduces the impact of shadows in the image, improves the accuracy of the U-net network model to extract the building edges, and improves the accuracy of building extraction through auxiliary water and vegetation information re-judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005042B_ABST
    Figure CN114005042B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for extracting urban buildings from remote sensing images based on shadow compensation and U-net, comprising the following steps: acquisition and correction of high-resolution remote sensing images; extraction of spectral features and brightness features of multi-spectral images, shadow detection and shadow compensation for multi-spectral images, and normalization processing; production of training images and prediction images using the normalized remote sensing images, and production of a sample data set using the training images; training of a U-net network model using the sample data set, and prediction of the prediction images using the trained U-net network model to obtain a preliminary prediction image for extracting building outlines; identification of water body and vegetation areas using spectral features, and optimization of the preliminary prediction image. The method disclosed by the present invention optimizes shadows in remote sensing data, can effectively reduce the influence brought by shadows, and uses auxiliary information for re-judgment of the prediction results of deep learning, thereby improving the accuracy of building extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image processing technology, and in particular to a remote sensing image urban building extraction method based on shadow compensation and U-net. Background Art

[0002] Accurate and timely monitoring of urban buildings is of great significance to the study of urban construction and the management of urban land resources. Remote sensing technology is a technology that can obtain ground information over a large range at a long distance. The use of high-resolution remote sensing images to monitor urban buildings has the advantages of wide coverage, strong data reliability, and low labor costs. It is widely used in urban planning, illegal construction monitoring, land rights confirmation, etc. Due to the relatively low resolution of satellite remote sensing images, the many interference factors in the imaging process, and the complex urban landscape, high-precision building extraction has always been a difficult point. Traditional supervised classification methods often cannot avoid the problem of misclassification and omission. Many practical applications that require high monitoring accuracy can only be completed by manual visual interpretation. In recent years, new progress in deep learning technology has gradually been applied in the field of remote sensing monitoring, and building extraction based on deep learning has become a cutting-edge topic.

[0003] Building extraction based on deep learning often uses pixel-level semantic segmentation methods. However, since building extraction requires accurate determination of building edges, more attention should be paid to shallow features of the image. U-net is a U-shaped convolutional neural network that downsamples the image through convolution to obtain deep image features, and then upsamples to increase the resolution of deep features, combining deep and shallow features to improve the accuracy of target detection. U-net was first applied to medical imaging. Due to its superior performance under small samples and good segmentation effect for small targets, it is also suitable for the remote sensing field. Due to the problem of satellite observation altitude, images are easily affected by clouds and mountain shadows. In high-resolution remote sensing images, building shadows are also more obvious, and building edges are difficult to extract completely due to shadow interference. Therefore, there is an urgent need for a method to improve the accuracy of U-net edge extraction of buildings. Summary of the invention

[0004] To solve the above technical problems, the present invention provides a remote sensing image urban building extraction method based on shadow compensation and U-net, which uses the spectral characteristics and brightness characteristics of the image to extract shadows, and compensates the shadow area with reference to the reflectivity of the bright pixels in the neighborhood window to reduce the influence of various shadows in the image on the extraction of building edges.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net, comprising the following steps:

[0007] Step 1: Obtain and correct high-resolution remote sensing images to obtain four-band multispectral images;

[0008] Step 2: Extract the spectral features and brightness features of the multispectral images, perform shadow detection and shadow compensation on the multispectral images, and perform normalization processing:

[0009] Step 3: Use the normalized remote sensing images to produce training images and prediction images, and use the training images to produce a sample data set;

[0010] Step 4: Use the sample data set to train the U-net network model, and use the trained U-net network model to predict the prediction images to obtain preliminary prediction images for extracting building outlines;

[0011] Step 5: Use spectral features to identify water and vegetation areas, and optimize the preliminary prediction images for extracting building outlines.

[0012] In the above solution, in Step 1, the high-resolution remote sensing images used are GF-2 L1A PMS images, including one panchromatic band and four multispectral bands.

[0013] In a further technical solution, in Step 1, the orthorectification of the PMS images is performed using the RPC model with the precisely corrected Sentinel-2 images as the reference. The radiometric calibration of each band image is carried out using the absolute radiometric calibration parameters provided by the China Resources Satellite Application Center. Then, the nearest neighbor diffusion method is used to fuse the multispectral bands with the panchromatic band to obtain four-band multispectral images resampled to a resolution of 1 meter.

[0014] In the above solution, in Step 2, the implementation method of shadow detection is to enhance and quantify the features of shadows according to the spectral features and brightness features of the multispectral images, and use the threshold decision method to preliminarily extract the shadow areas from the multispectral images. The calculation logic is as follows:

[0015]

[0016]

[0017]

[0018] Wherein, NDSI is the Normalized Difference Shadow Index, Y is the brightness feature, and NDWI is the Normalized Difference Water Index; Blue represents the reflectance layer of the blue band, NIR represents the reflectance layer of the near-infrared band, Green represents the reflectance layer of the green band, Red represents the reflectance layer of the red band. The addition and subtraction in the formula represent the operations between bands, Max() represents finding the maximum pixel value of the reflectance layer obtained by accumulating the three bands within the parentheses, and C1, C2, and C3 are the respective decision thresholds;

[0019] The area that simultaneously meets the above three conditions is the shadow area; after detecting the shadow area, a binary Mask map containing the shadow area and the non-shadow area is generated for the next step of shadow compensation.

[0020] In the above solution, in step two, the method of shadow compensation is as follows:

[0021] First, traverse the generated binary Mask map. When a shadow pixel is detected, a 3×3 window is generated with this shadow pixel as the center, and then the proportion of shadow pixels in this window is calculated. If this proportion is greater than 50%, the window size is expanded to 5×5, 7×7, …, until the proportion of shadow pixels in the window is less than 50%; if the window boundary touches the image boundary and still cannot meet the condition that the proportion of shadow pixels is less than 50%, stop expanding the window, and set the compensation gain k b of each band of this shadow pixel to 1;

[0022] Then, obtain the pixel values within the current window at the same position in the multi-spectral image, and calculate the compensation gain k b band by band:

[0023]

[0024] where m is the number of non-shadow pixels in the current window, n is the number of shadow pixels in the current window, the initial value of i is 0, and B_ns i represents the pixel value of the i-th non-shadow pixel in the current window in the b band, and B_s i represents the pixel value of the i-th shadow pixel in the current window in the b band;

[0025] Finally, use the obtained compensation gain to perform shadow compensation on each band of the shadow area. The calculation formula is as follows:

[0026] B′ b =k b B b , b ∈ [1, 4]

[0027] where B′ b represents the pixel value after shadow compensation in the b band, k b is the compensation gain in the b band, and Bb is the original pixel value in the b band.

[0028] In the above solution, in step two, the calculation formula for normalization processing is:

[0029]

[0030] where B_n b represents the pixel value after normalization processing in the b band, and B b represents the original pixel value in the b band. Max(B b ) represents finding the maximum pixel value in the b band, and Min(B b ) represents finding the minimum pixel value in the b band.

[0031] In the above solution, the specific method for step three is as follows:

[0032] First, select a representative area with dense buildings in the remotely sensed image after normalization processing, and make training images and prediction images through cropping;

[0033] After that, use the method of manual annotation to completely check the building outlines in the training images to make binary labels;

[0034] Then, divide the training images and binary labels respectively in the way of sliding windows, and crop them into small images with the same size of 512*512 pixel dimensions to obtain an image set and a corresponding label set respectively;

[0035] Secondly, batch flip the small images in the image set and label set segmented in the previous step in three directions: horizontal, vertical, and diagonal, so that both the image set and the corresponding label set become four times the original;

[0036] Finally, randomly select 70% of the images in the image set and the corresponding label set as the training data set, and the other 30% as the validation data set.

[0037] In the above solution, in step four, the training method of the U-net network model is as follows:

[0038] Construct a U-net network model, and input the training data set and the validation data set into the U-net network model. Among them, the training data is used to fit the U-net network model, and the validation data set is used to evaluate the effect of the U-net network model and adjust the hyperparameters to optimize the prediction performance of the U-net network model;

[0039] The U-net network model consists of a total of 19 convolutional processes to extract image features, performs 4 downsamplings and 4 upsamplings, and connects the feature layers of the same size before and after upsampling and downsampling to simultaneously retain the deep and shallow features of the image and prevent information loss caused by downsampling. The cross-entropy loss function is used to calculate the difference between the label value and the predicted value, and the adaptive moment estimation optimizer is used to correct the hyperparameters of the U-net network model. After 100 rounds of iterative optimization, the trained U-net network model is saved.

[0040] In a further technical solution, in step four, after the U-net network model is trained, the cropped predicted image is used as the target area image. First, the target area image is sequentially translated and window-cropped into small images of 512*512 size, then sequentially input into the trained U-net network model, and finally all the predicted result small images are stitched in the order of segmentation to obtain a preliminary predicted image for extracting the building contour.

[0041] In the above solution, the specific implementation method of step five is as follows:

[0042] According to the spectral characteristics of water bodies and vegetation respectively, traverse the predicted image to extract these two types of targets. The water body is extracted using the normalized difference water index, and the vegetation is extracted using the normalized difference vegetation index. The calculation logic is as follows:

[0043] NDWI < C4

[0044]

[0045] Among them, NDWI is the normalized difference water index, NDVI is the normalized difference vegetation index, NIR represents the near-infrared band reflectance layer, Red represents the red band reflectance layer, and C4 and C5 are their respective decision thresholds;

[0046] If the pixel at this position is identified as a water body or vegetation, then the same position of the preliminary predicted image for extracting the building contour is marked as non-building.

[0047] Through the above technical solution, the method provided by the present invention has the following beneficial effects:

[0048] 1. The present invention performs shadow detection and shadow compensation on remote sensing data, which can effectively reduce the influence of shadows in the image;

[0049] 2. The present invention performs normalization processing on the multi-spectral image after shadow compensation, which can reduce the difference in reflectance brightness of remote sensing images in different periods;

[0050] 3. The present invention combines shadow compensation and the U-net network model, which improves the edge accuracy of the U-net network model for extracting buildings;

[0051] 4. The prediction results of deep learning in the present invention are rejudged using auxiliary water body and vegetation information, which helps to improve the accuracy of building extraction. Description of the Drawings

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0053] Figure 1 It is a schematic diagram of the overall process of a method for extracting urban buildings from remote sensing images based on shadow compensation and U-net disclosed in the embodiments of the present invention;

[0054] Figure 2 It is a flowchart of shadow detection and shadow compensation of the present invention;

[0055] Figure 3 It is a flowchart of making the sample data set of the present invention. Detailed Embodiments

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention.

[0057] The present invention provides a method for extracting urban buildings from remote sensing images based on shadow compensation and U-net, as Figure 1 shown, including the following steps:

[0058] Step 1, obtaining and correcting a high-resolution remote sensing image to obtain a four-band multispectral image;

[0059] The high-resolution remote sensing image used is a GF-2 L1A PMS image, including a panchromatic band and four multispectral bands.

[0060] Using the precisely corrected Sentinel-2 image as a reference, the PMS image is orthorectified using the RPC model, the radiometric calibration of each band image is performed using the absolute radiometric calibration parameters provided by the China Resources Satellite Application Center, and then the multispectral band and the panchromatic band are fused using the nearest neighbor diffusion method to obtain a four-band multispectral image with a resampled resolution of 1 meter.

[0061] Step 2, extracting the spectral features and brightness features of the multispectral image, performing shadow detection and shadow compensation on the multispectral image, and performing normalization processing, as Figure 2 shown, including the following steps:

[0062] 1. Shadow detection

[0063] The implementation method of shadow detection is to enhance and quantify the features of shadows based on the spectral features (NDSI, NDWI) and brightness feature (Y) of multispectral images.

[0064] Since it is difficult for light to irradiate the areas under clouds and buildings, the reflectance of each band is very low, and the information received by the satellite is mainly scattered light. Moreover, the shorter the wavelength, the stronger the scattering. Therefore, by using the ratio difference feature between the near-infrared band and the blue band, the normalized shadow index (NDSI) can be constructed to highlight shadow information. Another method for enhancing shadow features is to utilize its brightness feature (Y). The brightness value of the shadow area is often low. Calculate the image brightness based on the red, green, and blue bands, and then extract the shadow area. In addition, since the reflectance and brightness of water bodies are similar to those of shadows, it is still not easy to distinguish them after the above two shadow feature enhancements. At this time, the normalized difference water index (NDWI) is used to remove water bodies.

[0065] Based on the above three points, the threshold decision method is used to preliminarily extract the shadow area from the multispectral image, and its calculation logic is as follows:

[0066]

[0067]

[0068]

[0069] Among them, NDSI is the normalized shadow index, Y is the brightness feature, and NDWI is the normalized difference water index; Blue represents the reflectance layer of the blue band, NIR represents the reflectance layer of the near-infrared band, Green represents the reflectance layer of the green band, Red represents the reflectance layer of the red band. The addition and subtraction in the formula represent operations between bands, Max() represents finding the maximum pixel value of the reflectance layer obtained by accumulating the three bands in the parentheses. C1, C2, and C3 are the respective decision thresholds, which can be obtained from prior analysis combined with specific experimental data. In the present invention, C1 = -0.53, C2 = 0.13, and C3 = 0.52 are adopted.

[0070] The area that simultaneously satisfies the above three conditions is the shadow area; after detecting the shadow area, a binary Mask map including the shadow area and the non-shadow area is generated for the next step of shadow compensation.

[0071] 2. Shadow compensation

[0072] The implementation method of shadow compensation is to calculate the spectral pixel value difference between the shadow area and the surrounding normally illuminated area, so as to adjust each band of the shadow area proportionally. For each pixel in the shadow area, due to differences in the thickness of the clouds and building forms that cause shadows, the shadow compensation gain k b is not fixed.

[0073] To accurately compensate each shadow pixel, first, traverse the binary Mask image generated in the previous step. When a shadow pixel is detected, a 3×3 window is generated with this shadow pixel as the center. Then, calculate the proportion of shadow pixels in this window. If this proportion is greater than 50%, the window size is expanded to 5×5, 7×7, …, until the proportion of shadow pixels in the window is less than 50%. If the condition that the proportion of shadow pixels is less than 50% still cannot be met when the window boundary touches the image boundary, stop expanding the window and set the compensation gain k of each band of this shadow pixel b all to 1;

[0074] Then, obtain the pixel values within the current window at the same position in the multispectral image, and calculate the compensation gain k band by band b , since the image used is a 4-band image, the value range of b is [1, 4].

[0075]

[0076] Among them, m is the number of non-shadow pixels in the current window, n is the number of shadow pixels in the current window, the initial value of i is 0, and B_ns i represents the pixel value of the i-th non-shadow pixel within the current window in the b band, and B_s i represents the pixel value of the i-th shadow pixel within the current window in the b band;

[0077] Finally, perform shadow compensation on each band of the shadow area using the obtained compensation gain. The calculation formula is as follows:

[0078] B′ b = k b B b , b ∈ [1, 4]

[0079] Among them, B′ b represents the pixel value after shadow compensation in the b band, k b is the compensation gain in the b band, and B b is the original pixel value in the b band.

[0080] 3. Image normalization

[0081] After the multispectral image undergoes shadow compensation, it needs to be normalized to reduce the difference in the reflection brightness of remote sensing images in different periods. The calculation formula for normalization is:

[0082]

[0083] Among them, B_n b represents the pixel value after normalization processing in the b band, and B brepresents the original pixel value in the b band, Max(B b ) represents finding the maximum pixel value in the b band, Min(B b ) represents finding the minimum pixel value in the b band.

[0084] Step 3: Use the normalized remote sensing image to produce training images and prediction images, and use the training images to produce a sample data set;

[0085] As Figure 3 shown, the method for producing the sample data set is as follows:

[0086] First, select a representative area with dense buildings in the normalized remote sensing image, and produce training images and prediction images through cropping;

[0087] After that, use the method of manual annotation to completely check the building outlines in the training images to produce binary labels;

[0088] Then, divide the training images and binary labels respectively in the way of sliding windows, crop them into small images with the same size of pixel dimension 512*512, and obtain an image set and the corresponding label set respectively;

[0089] Secondly, batch flip the small images in the image set and label set obtained in the previous step in the horizontal, vertical and diagonal directions, so that both the image set and the corresponding label set become four times the original;

[0090] Finally, randomly select 70% of the images in the image set and the corresponding label set as the training data set, and the other 30% as the validation data set.

[0091] Step 4: Use the sample data set to train the U-net network model, and use the trained U-net network model to predict the prediction image to obtain a preliminary prediction image for extracting the building outline;

[0092] Construct the U-net network model, input the training data set and the validation data set into the U-net network model. Among them, the training data is used to fit the U-net network model, and the validation data set is used to evaluate the effect of the U-net network model and adjust the hyperparameters to optimize the prediction performance of the U-net network model;

[0093] The U-net network model contains a total of 19 convolutional processes to extract image features, performs 4 downsamplings and 4 upsamplings, and connects the feature layers of the same size before and after upsampling and downsampling to simultaneously retain the deep and shallow features of the image and prevent information loss caused by downsampling. The cross-entropy loss function is used to calculate the difference between the label value and the predicted value, and the adaptive moment estimation optimizer is used to correct the hyperparameters of the U-net network model. After 100 rounds of iterative optimization, the trained U-net network model is saved.

[0094] After the U-net network model is trained, the cropped predicted image is used as the target area image. First, the target area image is sequentially translated and window-cropped into small images of size 512*512, then input into the trained U-net network model in turn, and finally all the predicted result small images are stitched in the segmentation order to obtain the preliminary predicted image for extracting the building contour.

[0095] Step five, use spectral features to identify water bodies and vegetation areas and optimize the preliminary predicted image for extracting the building contour.

[0096] According to the spectral features of water bodies and vegetation respectively, traverse the predicted image to extract these two types of targets. The water body is extracted using the normalized difference water index, and the vegetation is extracted using the normalized difference vegetation index. The calculation logic is as follows:

[0097] NDWI < C4

[0098]

[0099] Among them, NDWI is the normalized difference water index, NDVI is the normalized difference vegetation index, NIR represents the near-infrared band reflectance layer, Red represents the red band reflectance layer, and C4 and C5 are the respective decision thresholds, which are determined according to prior knowledge and the actual situation of the target area image. In the experiments of the present invention, C4 = 0 and C5 = 0.3 are adopted.

[0100] If the pixel at this position is identified as a water body or vegetation, the same position in the preliminary predicted image for extracting the building contour is marked as non-building.

[0101] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net, characterized in that, The method includes the following steps: Step 1: Obtain and correct the high-resolution remote sensing image to obtain a four-band multispectral image; Step 2: Extract the spectral features and brightness features of the multispectral image, perform shadow detection and shadow compensation on the multispectral image, and perform normalization processing: Step 3: Use the normalized remote sensing image to produce a training image and a prediction image, and use the training image to produce a sample data set; Step 4: Use the sample data set to train the U-net network model, and use the trained U-net network model to predict the prediction image to obtain a preliminary prediction image for extracting building outlines; Step 5: Use the spectral features to identify water bodies and vegetation areas, and optimize the preliminary prediction image for extracting building outlines; In Step 2, the implementation method of shadow detection is to enhance and quantify the features of the shadow according to the spectral features and brightness features of the multispectral image, and use the threshold decision method to preliminarily extract the shadow area from the multispectral image. The calculation logic is as follows: ; ; ; Among them, NDSI is the normalized shadow index, Y is the brightness feature, and NDWI is the normalized water index; Blue represents the blue-band reflectance layer, NIR represents the near-infrared band reflectance layer, Green represents the green-band reflectance layer, Red represents the red-band reflectance layer. The addition and subtraction in the formula represent inter-band operations, Max( ) represents finding the maximum pixel value of the reflectance layer obtained by accumulating the three bands in the parentheses, and C1, C2, and C3 are the respective decision thresholds; The area that satisfies the conditions shown in the above three inequalities at the same time is the shadow area; after detecting the shadow area, generate a binary Mask map including the shadow area and the non-shadow area for the next step of shadow compensation; In Step 2, the method of shadow compensation is as follows: First, traverse the generated binary Mask image. When a shaded pixel is detected, a 3×3 window is generated with this shaded pixel as the center. Then, calculate the proportion of shaded pixels in this window. If this proportion is greater than 50%, the window size is expanded to 5×5, 7×7, …, until the proportion of shaded pixels in the window is less than 50%. If the condition that the proportion of shaded pixels is less than 50% still cannot be met when the window boundary touches the image boundary, stop expanding the window and set the compensation gains of all bands of this shaded pixel all to 1; Then, obtain the pixel values within the current window at the same position of the multispectral image, and calculate the compensation gain band by band : ; Where m is the number of non-shadow pixels in the current window, n is the number of shadow pixels in the current window, and the initial value of i is 0. represents the pixel value of the i-th non-shadow pixel in the current window in the b band. represents the pixel value of the i-th shadow pixel in the current window in the b band. Finally, use the obtained compensation gain to perform shadow compensation on each band of the shadow area. The calculation formula is as follows: ; Among them, represents the pixel value after shadow compensation in the b band, is the compensation gain in the b band, is the original pixel value in the b band.

2. The method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 1, characterized in that, In Step 1, the high-resolution remote sensing image used is the GF-2 L1A PMS image, which includes a panchromatic band and four multispectral bands.

3. A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 1, characterized in that, In Step 2, the calculation formula for normalization processing is: ; Among them, represents the pixel value after b-band normalization processing, represents the original pixel value of the b-band, represents finding the maximum pixel value in the b-band, represents finding the minimum pixel value in the b-band.

4. A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 1, characterized in that, The specific method of Step 3 is as follows: First, select a representative area with dense buildings in the normalized remote sensing image, and make a training image and a prediction image after cropping; After that, use the method of manual annotation to completely check the building outlines in the training image to make a binary label; Then, divide the training image and the binary label in the way of a sliding window respectively, and crop them into small images with a consistent pixel size of 512*512 to obtain an image set and a corresponding label set respectively; Secondly, flip the small images in the image set and the label set obtained in the previous step in three directions: horizontal, vertical, and diagonal, so that both the image set and the corresponding label set become four times the original; Finally, randomly select 70% of the images in the image set and the corresponding label set as the training data set, and the other 30% as the validation data set.

5. A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 4, characterized in that, In Step 4, the training method of the U-net network model is as follows: Build a U-net network model, and input the training dataset and the validation dataset into the U-net network model. Among them, the training dataset is used to fit the U-net network model, and the validation dataset is used to evaluate the performance of the U-net network model and adjust the hyperparameters to optimize the prediction performance of the U-net network model; The U-net network model contains a total of 19 convolutional processes to extract image features, performs 4 downsamplings and 4 upsamplings, and connects the feature layers of the same size before and after upsampling and downsampling to simultaneously retain the deep and shallow features of the image and prevent information loss caused by downsampling; Use the cross-entropy loss function to calculate the difference between the label value and the predicted value, and use the adaptive moment estimation optimizer to correct the hyperparameters of the U-net network model. After 100 rounds of iterative optimization, save the trained U-net network model.

6. The method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 4, characterized in that In step four, after the U-net network model is trained, use the cropped predicted image as the target area image. First, translate the window of the target area image in order to crop it into small images of 512*512 size, then input them into the trained U-net network model in turn, and finally splice all the predicted result small images in the order of segmentation to obtain a preliminary predicted image of the extracted building outline.

7. A method for extracting urban buildings from remote sensing images based on shadow compensation and U-net according to claim 1, characterized in that, The specific implementation method of step five is as follows: According to the spectral characteristics of water bodies and vegetation respectively, traverse the predicted image to extract these two types of targets; Use the Normalized Difference Water Index (NDWI) for water body extraction and the Normalized Difference Vegetation Index (NDVI) for vegetation extraction; The calculation logic is as follows: ; ; Among them, NDWI is the Normalized Difference Water Index, NDVI is the Normalized Difference Vegetation Index, NIR represents the near-infrared band reflectance layer, Red represents the red band reflectance layer, and C4 and C5 are their respective decision thresholds; If a pixel at a certain position is identified as a water body or vegetation, mark the same position in the preliminary predicted image of the extracted building outline as non-building.

Citation Information

Patent Citations

  • Multi-temporal high-resolution remote sensing image building extraction method based on multi-feature LSTM network

    CN111582194A