High-resolution remote sensing image urban impervious surface extraction method

By applying data augmentation, mixed sampling and deep convolutional neural network semantic segmentation models in high-resolution remote sensing images, the problem of insufficient accuracy of urban impermeable surface information extraction in the prior art is solved, and higher-precision impermeable surface extraction and urban ecological environment monitoring are achieved.

CN120107784AActive Publication Date: 2025-06-06CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

Patent Information

Application Number
CN202510162910.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing technology is difficult to quickly and accurately extract high-precision urban impermeable information, which makes it difficult to effectively monitor and solve problems such as urban heat island effect and heavy rain and waterlogging disasters.

Method used

High-resolution remote sensing images are used to combine data augmentation, mixed sampling and deep convolutional neural network semantic segmentation models to build a feature fusion deep convolutional neural network semantic segmentation model to extract urban impermeable surfaces.

Benefits of technology

The extraction accuracy of impermeable surface types and spatial distribution is improved, and the dependence on prior knowledge and artificial intervention is reduced, and the building shape can be more accurately portrayed and road information under the canopy cover is extracted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107784A_ABST
    Figure CN120107784A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution remote sensing image urban impervious surface extraction method, and relates to the technical field of impervious surface remote sensing extraction. Comprising the following steps: acquiring high-spatial-resolution remote sensing data of a to-be-analyzed region; making a sample data set by combining the characteristics of the impervious surface; unbalanced sample data distribution quality in the sample data set is improved, and a high-quality sample data set is obtained and divided into a training set and a test set; constructing a feature fusion deep convolutional neural network semantic segmentation model; inputting the training set into a semantic segmentation model for model training to obtain a trained semantic segmentation model; inputting the test set into a trained semantic segmentation model to carry out ground feature classification, and generating a regional dense semantic city ground feature category classification result; and finally integrating impervious surface extraction results of a non-shadow region and a shadow region by adopting a shadow region impervious surface extraction scheme of multi-scale segmentation and a knowledge layering model to obtain urban impervious surface distribution. The method can improve the extraction precision of the impervious surface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing extraction of impervious surfaces, and in particular to a method for extracting urban impervious surfaces from high-resolution remote sensing images. Background Art

[0002] The continuous improvement of urbanization has caused a rapid expansion of impervious surfaces. Impervious surfaces refer to natural or artificial surfaces that prevent surface water from penetrating into the soil, such as roofs of buildings covered with waterproof materials, urban roads, squares, parking lots, etc., which are impervious. The material of impervious surfaces absorbs heat quickly and has a small specific heat capacity. It is easy to form a high heat storage body in the city, which leads to an increase in urban surface temperature and the formation of an urban heat island effect. The expansion of urban impervious surfaces prevents the infiltration of surface water, resulting in an increase in urban surface runoff, exacerbating urban rainstorm waterlogging disasters and causing ecological and environmental problems such as water pollution. Therefore, impervious surface information not only reflects the changes in land use and land cover caused by urbanization, but also has a profound impact on the urban ecological environment. Impervious surfaces are an important indicator for evaluating the quality of urban ecological environment. How to automatically extract high-precision impervious surface information in a timely and accurate manner is of great significance for monitoring the degree of urbanization expansion, sponge city planning and ecological construction, and achieving sustainable development in urban and rural areas.

[0003] The traditional way to obtain data on impervious surfaces mainly relies on manual ground survey and mapping, but this method is costly and slow to update. With the rapid development of satellite earth observation technology, the data obtained by remote sensing technology has the advantages of low cost, high efficiency and timeliness, wide coverage and repeated observation, and has been widely used in related research on impervious surface extraction. At present, the methods for extracting impervious surfaces based on different remote sensing data can be divided into spectral hybrid decomposition method based on optical remote sensing data and radar remote sensing data, regression model, decision tree model, spectral index method and classification method.

[0004] Therefore, it is an urgent problem for those skilled in the art to propose a method for extracting urban impervious surfaces from high-resolution remote sensing images to solve the difficulties existing in the prior art. Summary of the invention

[0005] In view of this, the present invention provides a method for extracting urban impervious surfaces from high-resolution remote sensing images, which can improve the extraction accuracy of impervious surface types and spatial distribution.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] A method for extracting urban impervious surfaces from high-resolution remote sensing images, comprising:

[0008] S1. Obtain high spatial resolution remote sensing data, high precision DSM data and open source POI geospatial data of the area to be analyzed;

[0009] S2. Create a sample data set based on the characteristics of urban impervious surfaces;

[0010] S3. Use data enhancement and mixed sampling techniques to improve the distribution quality of unbalanced sample data in the sample data set, obtain a high-quality sample data set, and divide the high-quality sample data set into a training set and a test set;

[0011] S4. Construct a deep convolutional neural network semantic segmentation model with feature fusion;

[0012] S5, inputting the training set into the feature-fused deep convolutional neural network semantic segmentation model for model training, and obtaining a trained feature-fused deep convolutional neural network semantic segmentation model;

[0013] S6. Input the test set into the trained feature-fused deep convolutional neural network semantic segmentation model to classify objects and generate the classification results of urban objects with regional dense semantics;

[0014] S7. Adopt the impervious surface extraction scheme of shadow area based on multi-scale segmentation and knowledge hierarchical model, and integrate the impervious surface extraction results of non-shadow area and shadow area to obtain the final urban impervious surface distribution.

[0015] In the above method, optionally, the high spatial resolution remote sensing data of the area to be analyzed in S1 include: high spatial resolution optical satellite GF-2 remote sensing data of the area to be analyzed, high-precision DSM data generated by stereo image pairs of Ziyuan-3 satellite of the area to be analyzed, and open source POI geospatial data of the area to be analyzed crawled based on an open interface.

[0016] In the above method, optionally, the sample data set in S2 includes: remote sensing data, spectral index data and annotated data;

[0017] Each sample data is in the form of: image-annotation-spectral index data;

[0018] The preparation of sample data sets includes: data preprocessing, sample annotation, and spectral index extraction;

[0019] Sample labeling: By directly labeling the original pixels with categories and vectorization, the urban features are classified into categories including bare land, grass, water, trees, roads, buildings, squares and shadows. Roads, buildings and squares are all impervious surface extraction results in non-shadow areas.

[0020] Spectral index extraction The spectral indexes of impervious surfaces include: NDWI index and NDVI index;

[0021] The NDVI index is calculated using the near-infrared band and red band in the acquired high-spatial-resolution remote sensing data images, and the NDWI index is calculated using the near-infrared band and green band.

[0022] In the above method, optionally, data enhancement in S3 includes: translation, rotation, mirroring, scaling, color space switching and adding interference;

[0023] Hybrid sampling: Combining random undersampling and random oversampling, it enhances the distribution of the minority class balanced data, reduces the classifier's bias towards the majority class, and obtains a high-quality sample data set.

[0024] In the above method, optionally, in S4, a feature-fused deep convolutional neural network semantic segmentation model is constructed based on the ResNet50 residual network and the VGG-16 convolutional neural network, and the feature-fused deep convolutional neural network semantic segmentation model includes: an encoding module, a decoding module, and a Softmax layer;

[0025] The encoding module adopts a dual-channel structure. The first channel uses the ResNet50 residual network as the skeleton network, combined with the hole convolution to extract features and obtain the RGB image features. The second channel is a branch network for NDVI data and NDWI data input. The five convolutional layers of the VGG-16 convolutional neural network are used as the skeleton trunk. After four times of 2-fold sampling, the feature output of 16-fold downsampling is finally generated to obtain the spectral index features. The stacking and bonding fusion method of inputting multiple vectors on the specified axis is used to explicitly establish the correlation between the RGB image features and the spectral index features NDVI and NDWI multi-feature mapping.

[0026] The decoding module adopts a multi-step decoder. In the decoding stage, the input of each layer includes the output of the previous deconvolution layer, and also includes the convolution output of the corresponding layer in the encoding stage, gradually restoring the feature map to the original resolution; the global-local Transformer block is introduced. The Transformer block consists of an efficient global-local attention mechanism, a multi-layer perceptron, two batch normalization operations and two addition operations.

[0027] The above method is optional, and the specific content of S5 is: using the training set data, iteratively training the constructed feature-fused deep convolutional neural network semantic segmentation model, updating the full convolutional network parameters in the feature-fused deep convolutional neural network semantic segmentation model, and completing the model training.

[0028] In the above method, optionally, in S7, a shadow-based segmentation recognition object is obtained by multi-scale segmentation;

[0029] The vegetation, bare soil, water bodies and impervious surfaces in the shadow area are reclassified to assist in distinguishing the impervious surfaces under the tree canopy cover. The impervious surface extraction results of the non-shadow area and the shadow area are integrated to obtain the final urban impervious surface distribution.

[0030] Through the above technical solutions, it can be seen that compared with the prior art, the present invention provides a method for extracting urban impervious surfaces from high-resolution remote sensing images, which has the following beneficial effects: 1) The present invention uses a deep learning method with multi-feature fusion to extract impervious surface information, which can obtain the optimal model parameters, reduce the dependence on prior knowledge and human intervention, and obtain higher impervious surface extraction accuracy; 2) A pixel-level spatial scale urban feature classification model is constructed, which can realize large-scale and refined impervious surface extraction and improve the accuracy of model extraction of impervious surfaces; 3) Shadows are extracted as a separate category, and secondary classification is performed based on the classified shadow results, while spectral index is used to extract the impervious surface information. The input of numerical data includes vegetation index and water body index, which are input into the convolutional neural network classification model to help distinguish water bodies, vegetation and shadow categories; 4) The classification accuracy of ground objects and non-ground objects is improved, and the shape of buildings can be more accurately portrayed, which is conducive to solving the problem of distinguishing between the same object with different spectra / different objects with the same spectrum and impervious surface buildings and trees in shadow areas in high-resolution images; 5) Optimize the extraction of roads and extract the road information under the canopy cover, merge the building categories and road / square categories obtained after hierarchical classification into the final impervious surface information, and obtain more accurate impervious surface coverage extraction by combining multi-source and multi-modal data fusion with deep learning methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0032] Figure 1 A flowchart of a method for extracting urban impervious surfaces from high-resolution remote sensing images provided by the present invention;

[0033] Figure 2 A module composition diagram of the deep convolutional neural network semantic segmentation model with feature fusion provided by the present invention;

[0034] Figure 3 This is a structural block diagram of the feature-fused deep convolutional neural network semantic segmentation model provided by the present invention. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0037] Reference Figure 1 As shown, the present invention discloses a method for extracting urban impervious surfaces from high-resolution remote sensing images, comprising:

[0038] S1. Obtain high spatial resolution remote sensing data, high precision DSM data and open source POI geospatial data of the area to be analyzed;

[0039] S2. Create a sample data set based on the characteristics of urban impervious surfaces;

[0040] S3. Use data enhancement and mixed sampling techniques to improve the distribution quality of unbalanced sample data in the sample data set, obtain a high-quality sample data set, and divide the high-quality sample data set into a training set and a test set;

[0041] S4. Construct a deep convolutional neural network semantic segmentation model with feature fusion;

[0042] S5, inputting the training set into the feature-fused deep convolutional neural network semantic segmentation model for model training, and obtaining a trained feature-fused deep convolutional neural network semantic segmentation model;

[0043] S6. Input the test set into the trained feature-fused deep convolutional neural network semantic segmentation model to classify objects and generate the classification results of urban objects with regional dense semantics;

[0044] S7. Adopt the impervious surface extraction scheme of shadow area based on multi-scale segmentation and knowledge hierarchical model, and integrate the impervious surface extraction results of non-shadow area and shadow area to obtain the final urban impervious surface distribution.

[0045] Furthermore, the high spatial resolution remote sensing data of the area to be analyzed in S1 include: high spatial resolution optical satellite GF-2 remote sensing data of the area to be analyzed, high-precision DSM data generated by the stereo image pair of the Ziyuan-3 satellite in the area to be analyzed, and open source POI geospatial data of the area to be analyzed crawled based on an open interface.

[0046] Furthermore, in S2, a sample database is created based on the characteristics of urban impervious surfaces, and rich samples of multi-band and multi-object types are added, thereby providing sample data set support for multi-feature fusion high-resolution remote sensing image object classification;

[0047] The sample data set includes: remote sensing data, spectral index data and annotation data;

[0048] Each sample data is in the form of "image-annotation-spectral index data", that is, for each original image, its corresponding annotation data and spectral index data can be obtained;

[0049] The preparation of sample data sets includes: data preprocessing, sample annotation, and spectral index extraction;

[0050] Sample labeling: By directly labeling the original pixels with categories and vectorization, the urban features are classified into categories including bare land, grassland, water, trees, roads, buildings, squares and shadows. Roads, buildings and squares are all impervious surface extraction results in non-shadow areas. The original pixels are directly labeled with categories, including: polyline, rectangle, straight line, ellipse, circle, arbitrary rectangle (three-point method), parallelogram and arbitrary polygon.

[0051] Spectral index extraction selects spectral indicators of impervious surfaces, including NDWI and NDVI, which can further improve the classification accuracy of water, trees and shadows. As the input factor of the feature fusion network, it corresponds one-to-one with the original image and the annotated data;

[0052] The NDVI index is calculated using the near-infrared band and red band in the acquired high-spatial-resolution remote sensing data images, and the NDWI index is calculated using the near-infrared band and green band.

[0053] Furthermore, data enhancement in S3 includes: translation, rotation, mirroring, scaling, color space switching and adding interference, and new samples are generated through data augmentation technology;

[0054] Hybrid sampling improves the unbalanced distribution of sample sets from both global and local perspectives. It combines remote sensing images with other different spectral indices and adds other multi-source data that are beneficial to extracting minority classes such as water bodies and roads as much as possible. It can help distinguish minority classes and other types of objects in complex areas.

[0055] Hybrid sampling combines random undersampling and random oversampling to balance the distribution of training data by enhancing minority classes such as permeable surface classes, and attempts to reduce the classifier's bias towards majority classes such as grassland and bare land. Ultimately, a sufficiently high-quality training sample set is obtained, which helps to effectively use samples in the training phase, so that the trained deep learning model can also achieve better classification performance.

[0056] Furthermore, in S4, based on the ResNet50 residual network and the VGG-16 convolutional neural network, combined with the Transformer module for extracting global information and the encoding-decoding structure, the multi-scale features of the full convolutional network and the contextual information are effectively combined to enhance the feature extraction of remote sensing image objects and construct a feature-fused deep convolutional neural network semantic segmentation model.

[0057] Reference Figure 2 As shown, the feature fusion deep convolutional neural network semantic segmentation model includes: encoding module, decoding module, and Softmax layer;

[0058] The network model consists of two parallel networks, and the feature fusion module enables the entire network to learn branch fusion features. This parallel structure simultaneously processes RGB image data and impervious surface-specific spectral index data (NDVI data and NDWI data), and realizes the extraction of regional dense semantic impervious surface categories and other ground object categories through the network structure. It can not only identify different ground object targets, but also restore clearer ground object boundaries, significantly improving the recognition accuracy of impervious surface categories, forming an end-to-end complete high-resolution remote sensing image classification model;

[0059] Further, refer to Figure 3As shown in the figure, the encoding module adopts a dual-channel structure. The first channel uses the ResNet50 residual network as the skeleton network, combined with the hole convolution to extract features and obtain RGB image features; the ResNet50 residual network consists of a convolution layer and four blocks, each block has several bottleneck units, and inside the bottleneck unit, there is a shortcut connection between the input and the output. ResNet connects the bypass information and uses the bottleneck block to solve the gradient disappearance problem, and obtains higher accuracy and a smaller model under a deep network. The standard ResNet50 is a 32-fold downsampling structure; the present invention changes the convolution step size of the first bottleneck block 3*3 in block3 to 1, and at the same time, in order to maintain the receptive field of the remaining convolution kernels in block3, the standard convolution is replaced by a hole convolution with a hole rate of 2 to obtain the original The image is downsampled 16 times; in order to better capture global information, the last feature map of the module uses global average pooling plus 1*1 convolution to change the number of channels. This method can provide richer global information than maximum pooling; the second channel is a branch network for NDVI data and NDWI data input. The 5 convolutional layers of the VGG-16 convolutional neural network are used as the backbone. After four times of 2-fold sampling, the 16-fold downsampled feature output is finally generated to obtain the spectral index feature. The stacking and bonding fusion method of inputting multiple vectors on the specified axis is used to explicitly establish the correlation between the RGB image features and the spectral index features NDVI and NDWI multi-feature mapping. This method not only enhances the model's recognition ability of impervious surface features, but also avoids the introduction of redundant features and noisy images;

[0060] The decoding module uses a multi-step decoder. In the decoding stage, the input of each layer includes the output of the previous deconvolution layer, and also includes the convolution output of the corresponding layer in the encoding stage, gradually restoring the feature map to the original resolution. The global-local Transformer block is introduced. The Transformer block consists of an efficient global-local attention mechanism, a multi-layer perceptron, two batch normalization operations, and two addition operations.

[0061] The global-local attention mechanism combines the global attention branch and the convolutional local branch to capture global and local context information. In the global branch, the window-based multi-head self-attention and cross-window context interaction modules are introduced to capture the global context with low complexity. In the local branch, the convolutional layer is applied to extract the local context.

[0062] The calculation formula of multi-head self-attention is:

[0063] MultiHead(Q,K,V)=Concat(head 1 ,...,head h )W O

[0064]

[0065] Among them, Q represents the query vector (Query), K represents the key vector (Key), and V represents the value vector (Value); head i (i=1, 2, ..., h, h is the number of heads) represents the attention result calculated by the i-th “head”; W O Represents a learnable weight matrix that linearly transforms the results of the concatenated multiple "heads" and maps them to the appropriate output space dimension; represents the weight matrix corresponding to the query vector for the i-th “head”, represents the weight matrix corresponding to the key vector of the i-th "head", represents the weight matrix corresponding to the value vector of the i-th "head", represents the linear transformation matrix from the model hidden dimension to the key vector dimension, represents the linear transformation matrix from the model hidden dimension to the value vector dimension, Represents the linear transformation matrix from the concatenated dimension of the multi-head attention to the hidden dimension of the model. The attention function calculated by each head is usually the scaled dot product attention. The matrices Q, K, and V can be defined as: The i-th row is the vector q i , k i and v i A matrix that linearly transforms a vector:

[0066]

[0067] in, represents the dot product of each query vector in Q and each key vector in K; d k represents the dimension of the key vector K, which is used to scale the attention score; Represents the value vector (Value) matrix, which is obtained by linear transformation of the input sequence

[0068] The weighted sum operation (WeightedSum, WS) can selectively weight these two features to learn more effective fusion features based on their contribution to segmentation accuracy; the batch normalization layer of the Transformer block makes the features extracted by the network more stable. The input of a batch is set to x = x 1,…,m , where m is the batch size and the algorithm training parameters are γ and β,

[0069] Calculate the average value of the batch data as:

[0070]

[0071] The variance is:

[0072]

[0073] The normalization method is expressed as:

[0074]

[0075] where ε is a small positive number (usually 10 -8 The purpose is to prevent the variance A division by zero error occurs when the time is very small;

[0076] The output of the algorithm is:

[0077]

[0078] Among them, γ is used to scale the normalized data, and β is used to translate the normalized data. By adjusting γ and β, the model can learn the most appropriate data distribution; BN γ,β represents a batch normalization operation with learnable parameters γ and β, which are used to adjust the data distribution and optimize model training.

[0079] Finally, the Softmax layer is used for classification, and the final feature map is converted into a category probability distribution, so that the network can output the probability that each pixel belongs to each category, providing accurate classification results for the extraction of impervious surfaces and surrounding landforms.

[0080] Furthermore, the specific content of S5 is: using the training set data, iteratively training the constructed feature fused deep convolutional neural network semantic segmentation model, updating the full convolutional network parameters in the feature fused deep convolutional neural network semantic segmentation model, and completing the model training;

[0081] Furthermore, in S7, the shadow area segmentation recognition object is first obtained through multi-resolution segmentation, and the selection of scale parameters, determination of thresholds, and rule-based hierarchical knowledge strategies are formulated;

[0082] Due to the limitations of optical remote sensing sensors, multi-source heterogeneous data such as DSM and POI are used as a supplementary data source for decision-level fusion. For example, impervious surfaces in a large geographic space have low proportion and high aggregation characteristics. Through a graphics-based target domain reduction method, thresholds for multiple features in multi-source heterogeneous data are established based on the knowledge base. As shown in Table 1, the prior knowledge of the knowledge base is constructed by candidate samples provided by the impervious surface extraction results of the fully convolutional network. The vegetation (including grass and trees), bare soil, water bodies and impervious surfaces in the shadow area are reclassified, and the impervious surfaces under the canopy cover are assisted to be distinguished. Finally, the impervious surface extraction results of the non-shadow area and the shadow area are integrated to obtain the final urban impervious surface distribution.

[0083] Table 1. Knowledge base of impervious surface identification model

[0084]

[0085]

[0086] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only schematic, wherein the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.

[0087] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for extracting urban impervious surfaces from high-resolution remote sensing images, characterized in that: include: S1. Obtain high spatial resolution remote sensing data, high precision DSM data and open source POI geospatial data of the area to be analyzed; S2. Create a sample data set based on the characteristics of urban impervious surfaces; S3. Use data enhancement and mixed sampling techniques to improve the distribution quality of unbalanced sample data in the sample data set, obtain a high-quality sample data set, and divide the high-quality sample data set into a training set and a test set; S4. Construct a deep convolutional neural network semantic segmentation model with feature fusion; S5, inputting the training set into the feature-fused deep convolutional neural network semantic segmentation model for model training, and obtaining a trained feature-fused deep convolutional neural network semantic segmentation model; S6. Input the test set into the trained feature-fused deep convolutional neural network semantic segmentation model to classify objects and generate the classification results of urban objects with regional dense semantics; S7. Adopt the impervious surface extraction scheme of shadow area based on multi-scale segmentation and knowledge hierarchical model, and integrate the impervious surface extraction results of non-shadow area and shadow area to obtain the final urban impervious surface distribution.

2. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: The high spatial resolution remote sensing data of the area to be analyzed in S1 include: high spatial resolution optical satellite GF-2 remote sensing data of the area to be analyzed, high-precision DSM data generated by the stereo image pairs of the Ziyuan-3 satellite in the area to be analyzed, and open source POI geospatial data of the area to be analyzed crawled based on open interfaces.

3. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: The sample data set in S2 includes: remote sensing data, spectral index data and annotation data; Each sample data is in the form of: image-annotation-spectral index data; The preparation of sample data sets includes: data preprocessing, sample annotation, and spectral index extraction; Sample labeling: By directly labeling the original pixels with categories and vectorization, the urban features are classified into categories including bare land, grass, water, trees, roads, buildings, squares and shadows. Roads, buildings and squares are all impervious surface extraction results in non-shadow areas. Spectral index extraction The spectral indexes of impervious surfaces include: NDWI index and NDVI index; The NDVI index is calculated using the near-infrared band and red band in the acquired high-spatial-resolution remote sensing data images, and the NDWI index is calculated using the near-infrared band and green band.

4. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: Data enhancement in S3 includes: translation, rotation, mirroring, scaling, color space switching, and adding interference; Hybrid sampling: Combining random undersampling and random oversampling, it enhances the distribution of the minority class balanced data, reduces the classifier's bias towards the majority class, and obtains a high-quality sample data set.

5. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: In S4, based on the ResNet50 residual network and the VGG-16 convolutional neural network, a feature-fused deep convolutional neural network semantic segmentation model is constructed. The feature-fused deep convolutional neural network semantic segmentation model includes: an encoding module, a decoding module, and a Softmax layer; The encoding module adopts a dual-channel structure. The first channel uses the ResNet50 residual network as the skeleton network, combined with the hole convolution to extract features and obtain the RGB image features. The second channel is a branch network for NDVI data and NDWI data input. The five convolutional layers of the VGG-16 convolutional neural network are used as the skeleton trunk. After four times of 2-fold sampling, the feature output of 16-fold downsampling is finally generated to obtain the spectral index features. The stacking and bonding fusion method of inputting multiple vectors on the specified axis is used to explicitly establish the correlation between the RGB image features and the spectral index features NDVI and NDWI multi-feature mapping. The decoding module adopts a multi-step decoder. In the decoding stage, the input of each layer includes the output of the previous deconvolution layer, and also includes the convolution output of the corresponding layer in the encoding stage, gradually restoring the feature map to the original resolution; the global-local Transformer block is introduced. The Transformer block consists of an efficient global-local attention mechanism, a multi-layer perceptron, two batch normalization operations and two addition operations.

6. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: The specific content of S5 is: use the training set data to iteratively train the constructed feature-fused deep convolutional neural network semantic segmentation model, update the full convolutional network parameters in the feature-fused deep convolutional neural network semantic segmentation model, and complete the model training.

7. The method for extracting urban impervious surfaces from high-resolution remote sensing images according to claim 1, characterized in that: In S7, shadow-based segmentation and object recognition are obtained through multi-scale segmentation; The vegetation, bare soil, water bodies and impervious surfaces in the shadow area are reclassified to assist in distinguishing the impervious surfaces under the tree canopy cover. The impervious surface extraction results of the non-shadow area and the shadow area are integrated to obtain the final urban impervious surface distribution.

Citation Information

Patent Citations

  • Urban main built-up area remote sensing extraction method based on impervious surface aggregation density

    CN105095888A

  • Land cover remote sensing monitoring method based on multi-source feature fusion

    CN115527123A

  • Remote sensing multispectral image water body identification method based on two-channel segmentation network

    CN117726936A

  • Accurate inversion method and system for aboveground biomass of urban vegetations considering vegetation type

    US20240312206A1

Cited By

  • Remote sensing image water body extraction method and system based on RGB + X data view angle

    CN121982515A