GAN-based high-resolution remote sensing image ground object extraction result optimization method

Through the GAN-based method, the feature extraction network, generator network and discriminator network are trained to optimize the geometry extraction results for high-resolution remote sensing images, solving the problem of poor extraction effect of complex geometry boundaries and special morphological buildings in the prior art, and achieving higher extraction accuracy and regularity.

CN120219979APending Publication Date: 2025-06-27CHANGGUANG SATELLITE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510303210.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has poor extraction effect on complex land boundaries and special morphological buildings in high-resolution remote sensing images, and there are problems of overfitting, category imbalance and insufficient generalization capabilities.

Method used

Using the GAN-based high-resolution remote sensing image object extraction result optimization method, the optimized data set is constructed, the feature extraction network, the generator network and the discriminator network are trained, and the optimization results are generated and post-processed to improve the accuracy and regularity of the extraction results.

Benefits of technology

The accuracy and reliability of the boundary extraction results of special morphological buildings and complex land objects are improved, and the missed extraction and missed extraction are reduced. The morphological characteristics of the generated optimization results are more consistent with the truth value label, which is suitable for optimization models of different remote sensing land objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219979A_ABST
    Figure CN120219979A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing technology application, and provides a GAN-based high-resolution remote sensing image ground feature extraction result optimization method in order to improve the accuracy and regularity of special form building and complex ground feature boundary extraction in a high-resolution remote sensing image, which comprises the following steps: constructing an optimization data set of ground feature extraction results; taking the constructed optimization data set as input data for training to obtain an optimization model, wherein the optimization model comprises a feature extraction network, a generator network and a discriminator network; optimizing a ground feature extraction result; the feature extraction network extracts multi-scale features, the generator network extracts and learns the features of an input image and generates an optimization result, and the discriminator network and the generator network perform adversarial learning, so that the optimization result generated by the generator network effectively reduces the phenomena of wrong extraction and missing extraction; and the accuracy and the reliability of the boundary extraction result of the special-form building and the complex ground feature are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing technology applications, and specifically relates to a method for optimizing the extraction results of ground objects in high-resolution remote sensing images based on GAN. Background Art

[0002] With the development of remote sensing technology, high-resolution satellite and aerial images have become important means for obtaining geographical information and are of great value in fields such as ground object extraction, urban planning, and environmental monitoring. Although the image resolution has been improved, the automated ground object extraction method based on artificial intelligence cannot accurately extract in the face of complex ground object boundaries and scene distributions; although deep learning technology has improved the accuracy of ground object boundary extraction, it still faces technical problems such as overfitting, class imbalance, and insufficient generalization ability. The generative adversarial network GAN provides a new idea for optimizing the extraction of ground objects in remote sensing images due to its excellent ability in image generation and enhancement.

[0003] Existing methods for optimizing the extraction results of ground objects in remote sensing images, such as the patent "A Method and System for Extracting Building Images in Remote Sensing Images" with the publication number CN110334719B, provide a method including: obtaining a convolutional neural network model; the convolutional neural network model is a trained neural network model with remote sensing images as input and building images as output; obtaining remote sensing images of the area to be collected; inputting the remote sensing images into the convolutional neural network model to extract the building images of the area to be collected, obtaining a preliminary extraction result; optimizing the preliminary extraction result using morphological closing operation to obtain the final extraction result of the building images in the area to be collected; the above patent has technical problems of high dependence on the result accuracy and still poor extraction effect for buildings with special shapes. Summary of the Invention

[0004] In order to improve the accuracy and regularity of the extraction of buildings with special shapes and complex ground object boundaries in high-resolution remote sensing images, the present invention proposes a method for optimizing the extraction results of ground objects based on GAN and post-processing, which can further optimize the extraction results in vector format.

[0005] As Figure 1 shown, the method for optimizing the extraction results of ground objects in high-resolution remote sensing images based on GAN specifically includes:

[0006] Step 1: Construct an optimization dataset for the extraction results of ground objects, so that each group of samples has three corresponding pictures: the original image, the extraction result, and the true value label.

[0007] Step 2: Use the optimization dataset constructed in Step 1 as input data to train and obtain an optimization model.

[0008] The optimization model includes a feature extraction network, a generator network, and a discriminator network. The feature extraction network can fuse the spatial details of low-scale features and the semantic information of high-scale features in the input data to extract features. The generator network converts the extracted intermediate feature representation back to the image space to generate an optimized result image. The discriminator network works in cooperation with the generator network and uses adversarial training. It can receive the optimized result image generated by the generator network, the ground truth label reconstruction result, and the ground truth label, and use them as inputs for classification and judgment.

[0009] Step 3: Optimization of the ground object extraction result:

[0010] Input the ground object extraction result to be optimized into the optimization model trained in Step 2 to obtain a raster ground object extraction result.

[0011] Obtain a vector ground object patch from the raster ground object extraction result through the raster-to-vector function.

[0012] After reducing the point data volume and smoothing the outer contour burrs of the vector ground object patch, an optimized result in vector format is obtained, thus completing the optimization of the ground object extraction result.

[0013] Technical effects:

[0014] The present invention can automatically optimize the results extracted from satellite remote sensing images: handle the phenomena of misclassification, missed classification, noise, and irregular contour edges caused by the complex and variable features of ground objects and the limitations of algorithms, and improve the accuracy and reliability of the extraction results of special-shaped buildings and complex ground object boundaries. The feature extraction network extracts multi-scale features, the generator network extracts and learns the features of the input image to generate an optimized result, and the discriminator network performs adversarial learning with the generator network and uses adversarial training to make the optimized result generated by the generator network not only improve the consistency of the contour of the extraction result with the ground truth label in morphological features but also effectively reduce the phenomena of mis-extraction and missed extraction, obtaining an optimization model applicable to different remote sensing ground objects.

[0015] In order to better demonstrate the performance of the optimization model in the present invention, FT-UNetFormer and RS3Mamba are selected as the basic feature extraction methods, and the extraction results are obtained by extracting two different features, buildings and roads. P (precision), R (recall), F1-Score and IoU (interSEction over union) are used as evaluation indicators to compare the present invention with the boundary regularization method based on the loss function (hereinafter referred to as optimization 1) and the regularization-based optimization method (hereinafter referred to as optimization 2). Among them, P refers to the proportion of pixels correctly classified as positive in all pixels classified as positive; R refers to the proportion of pixels correctly classified as positive in all pixels of positive; F1-Score is the harmonic mean result of P and R; IoU is the intersection of all pixels predicted as positive and true positive pixels over their union.

[0016] The comparison results are as follows:

[0017] 1. Optimization and comparison of building extraction results

[0018] Buildings usually have clear geometric shapes and obvious regional properties, but the results obtained through automated extraction cannot fit the actual boundaries and need to be regularized. The specific evaluation indicators are shown in the following table. The comparison of the extraction optimization results is shown in the figure below. Figures 3 - 4 shown.

[0019] Method P R F1 IoU FT-UNetFormer 67.62% 87.99% 76.47% 61.91% FT-UNetFormer - Optimization 1 67.74% 92.27% 78.12% 64.10% FT-UNetFormer - Optimization 2 71.23% 91.94% 80.27% 67.05% FT-UNetFormer - The present invention 71.23% 92.35% 80.43% 67.26% RS3Mamba 63.17% 90.41% 74.37% 59.20% RS3Mamba - Optimization 1 66.99% 92.01% 77.53% 63.30% RS3Mamba - Optimization 2 70.32% 88.96% 78.55% 64.67% RS3Mamba - The present invention 69.98% 92.63% 79.73% 66.29%

[0020] The present invention is applicable to different basic extraction and building extraction methods. The average evaluation index can be improved by about 5% compared with the original method, which has a great optimization effect. And compared with the two comparative optimization methods, such as Figure 3 As shown, the present invention can achieve better optimization results in the case of buildings that are closer to each other; and in the process of optimization, it can better form a discreteness of the buildings, which can make the results more accurate for subsequent statistical tasks related to single-building data. Figure 4 It can be intuitively observed that the building after optimization by the present invention presents regularity, which is more similar to the shape of the label. Compared with the comparison method, the proposed method can obtain a more regular building shape, and the attributes such as size are more in line with the actual building, which proves the advancement of the proposed method in the task of optimizing building results.

[0021] 2. Optimization and comparison of road extraction results

[0022] The road is more complex than other ground object attributes. In addition to obtaining accurate extraction results, the results also need to meet the characteristics of regular edges and accurate connectivity. Therefore, it is necessary to optimize the road extraction results. The specific evaluation indicators are shown in the following table. Due to the dual requirements of accuracy and connectivity logic for the road, it requires a higher optimization effect. The present invention focuses on supplementing the connectivity on the road, and the recall rate has been greatly improved. Due to the great difficulty of optimization, the other comparison methods have a small improvement in the F1 and IoU values. The results optimized by the method in this paper are still the best indicators.

[0023] Method P R F1 IoU FT-UNetFormer 75.99% 66.03% 70.66% 54.63% FT-UNetFormer - Optimization 1 75.64% 79.38% 77.46% 63.22% FT-UNetFormer - Optimization 2 83.32% 77.61% 80.37% 67.18% FT-UNetFormer - The present invention 83.94% 78.63% 81.20% 68.35% RS3Mamba 77.17% 64.62% 70.34% 54.25% RS3Mamba - Optimization 1 77.80% 65.04% 70.85% 54.86% RS3Mamba - Optimization 2 77.48% 65.22% 70.82% 54.83% RS3Mamba - The present invention 70.59% 72.22% 71.40% 55.52%

[0024] The comparison diagram of the extraction optimization results is as Figures 5 - 6 , such as Figure 5 shown. Each method has connected the truncated parts in the FT-UNetFormer road results, but Optimization 2 generates noise, which affects the actual use effect. For the 3 breakpoints generated in the RS3Mamba road results, although the 2 comparison methods have optimized them, they are not completely joined due to excessive truncation. However, the present invention directly generates a road optimization result with better connectivity, completely connecting all 3 breakpoints. As Figure 6 shown, the present invention also obtains better results in optimizing the FT-UNetFormer road results. For the discrete roads obtained by RS3Mamba, the present invention can also optimize them into a road that is more similar to the actual one, and the effect is more prominent. The road optimized by this paper has neater edges and is more in line with the actual road shape.

[0025] The extraction results processed by the present invention can ensure the consistency in form, further meet the high-level requirements for distribution characteristics in practical applications. Designing a general processing method can reduce the R & D investment in the post-processing field, and the same set of technologies can be used for the optimization processes of different ground objects. At the same time, it can assist the entire information extraction process to achieve fully automated processing, reduce the labor cost of the entire information extraction process, and conform to the future development direction. Brief Description of the Drawings

[0026] Figure 1 is the overall flow block diagram of the present invention.

[0027] Figure 2 is the schematic diagram of the feature extraction network structure.

[0028] Figure 3 is the comparison diagram of the extraction optimization results for the building areas with relatively close distances.

[0029] Figure 4 is the comparison diagram of the extraction optimization results for the regularized building areas.

[0030] Figure 5Comparison chart of optimization results for multi-breakpoint road area extraction.

[0031] Figure 6 Comparison chart of optimization results for discrete road area extraction.

[0032] in, Figures 3 - 6 In the figure, a is the original image; b is the GT; c is the extraction result of the FT-UNetFormer method; d is the optimization result of the FT-UNetFormer extraction result using optimization 1; e is the optimization result of the FT-UNetFormer extraction result using optimization 2; f is the optimization result of the FT-UNetFormer extraction result using the present invention; h is the extraction result of the RS3Mamba method; i is the optimization result of the RS3Mamba extraction result using optimization 1; j is the optimization result of the RS3Mamba extraction result using optimization 2; k is the optimization result of the RS3Mamba extraction result using the present invention. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by ordinary technicians in the field without creative work by adopting the embodiments of the present invention are within the scope of protection of the present invention.

[0034] The present invention provides a method for optimizing the extraction results of high-resolution remote sensing images based on GAN, which can improve the accuracy and regularity of the boundaries of the extraction results of high-resolution remote sensing images. Figure 1 As shown, the specific steps include:

[0035] Step 1: construct an optimized dataset of ground object extraction results based on the existing ground object extraction result dataset, so that each group of samples has a one-to-one correspondence of three pictures: original image, true value label and extraction result;

[0036] In this embodiment, an optimized data set of 256×256 size is used, and samples with a large number of mis-extractions and missed extractions are eliminated to obtain a final optimized data set.

[0037] Step 2: Use the optimized data set constructed in step 1 as input data to train the optimized model;

[0038] The optimization model includes a feature extraction network, a generator network, and a discriminator network. The feature extraction network can fuse the spatial details of low-scale features and the semantic information of high-scale features in the input data to extract features. The generator network converts the extracted intermediate feature representation back to the image space to generate an optimized result image. The discriminator network works in cooperation with the generator network and uses adversarial training. It can receive the optimized result image generated by the generator network, the ground truth label reconstruction result, and the ground truth label as inputs for classification and judgment.

[0039] The feature extraction network is as Figure 2 shown, which is a multi-scale information extraction network that comprehensively utilizes global structure and local detail features to improve feature expression ability. Taking the original image x, the ground truth label y, and the extraction result z as input data, the feature extraction network uses SE_ResNeXt101 as the backbone network for feature extraction, giving full play to the advantages of multiple features in the backbone network, and fusing the spatial details of low-scale features and the semantic information of high-scale features to provide more comprehensive information. It also includes a fusion module that can connect multiple features by combining 3×3Conv, BN, and ReLU activation functions, as well as a 2×2 max pooling layer.

[0040] The generator network includes multiple residual connection layers, 2×2 upsampling layers, BN, and ReLU activation functions. The last convolutional layer uses a 1×1 convolutional kernel and a Sigmoid activation function to generate an optimized result image. Adopting the decoding part of a typical convolutional autoencoder architecture, it gradually converts the intermediate feature representation of the encoder back to the image space to generate an optimized result image. At the same time, combined with loss functions, including regularization loss functions, by applying terms such as Potts loss and normalized cut loss, the ground feature learned by the generator network is more in line with the geometric shape of the actual ground features, improving the quality of the segmentation results. The Potts loss and normalized cut loss are defined as:

[0041] L Potts (G) = E x,z ∑ k S k W(1 - S k )

[0042]

[0043] where S k is the vectorized representation of the k-th channel of the softmax mask S, representing one of the optimized results output by the generator network. In this embodiment, k = 2; W and are pairwise discontinuity cost matrices used to quantify the discontinuity between pixels or regions.

[0044] It also includes a reconstruction loss function:

[0045] L recG (G) = -E x,z [x·logG(x,z)], ensuring that the data generated by the generator network is geometrically consistent with the ground truth data and helping the generator network learn the distribution characteristics of the data.

[0046] The discriminator network uses multiple 3×3 Conv, 2×2 max pooling layers, and fully connected layers to extract features and make classification judgments. It receives the optimized result images, ground truth label reconstruction results, and ground truth labels generated by the generator network as inputs. By learning to distinguish the patterns between the generated data and the real data, it outputs a probability value representing the likelihood that the input data is real data. The discriminator network works in collaboration with the generator network to guide the generation results of the generator network to better fit the actual contours and distribution characteristics of the ground objects.

[0047] Furthermore, the discriminator network and the generator network perform adversarial learning. Using adversarial training, the optimized result images generated by the generator network can not only improve the morphological features of the result contours to be consistent with the ground truth labels but also effectively reduce the phenomena of mis-extraction and missed extraction, obtaining an optimized model applicable to the ground object extraction results of different remote sensing images. The adversarial loss function L GAN of the generator network and the adversarial loss function L D of the discriminator network are as follows:

[0048] L GAN (G,D) = E x,z [log(1 - D(G(x,z)))],

[0049] L D (G,D) = E y [log(y)] + E y [logD(y)] + E x,z [log(1 - D(G(x,z)))],

[0050] where G(x,z) is the result generated by the generator network based on the input original image x and the extraction result z, y is the ground truth label, and D(y) and D(G(x,z)) are the outputs of the discriminator network for the ground truth label y and the result of the generator network respectively. The adversarial loss function can achieve a dynamic balance between data generation and discrimination, improve the effect of the generator network on optimizing the ground object extraction results, and enhance the ability of the discriminator network to distinguish the quality of the optimization results.

[0051] That is, the total loss function of the generator network is the weighted sum of each loss function:

[0052] L(G,D) = αL GAN (G,D) + βL recG(G) + γL Potts (G) + δL ncut (G),

[0053] Among them, α, β, γ, and δ are weight parameters used to balance the contributions of different loss functions. By minimizing the total loss function, the generator network learns to generate optimized results that are more in line with actual ground object features, while the discriminator network learns to evaluate the authenticity of these optimized results, improving the generalization of the method.

[0054] Step 3: Optimization of the ground object extraction result:

[0055] Input the ground object extraction result to be optimized into the optimization model trained in Step 2 to obtain the raster ground object extraction result;

[0056] Convert the raster ground object extraction result through the raster-to-vector function, set parameters to retain the target ground object part, and delete the background value part to complete the conversion from raster to vector, obtaining the vector ground object patch;

[0057] After reducing the point data volume and smoothing the outer contour burrs of the vector ground object patch, an optimized result in vector format is obtained, thus completing the optimization of the ground object extraction result. Specifically, apply the Douglas-Peucker algorithm, set the threshold parameter of the algorithm, and perform iterative processing on the boundary of each ground object patch to remove the detail parts smaller than the threshold and retain the main contour features. At the same time, smooth the boundary to eliminate the burr phenomenon. After processing all ground object patches, the finally optimized result in vector format is obtained, further improving the accuracy and practicality of the ground object extraction result.

[0058] The above are only embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A GAN-based method for optimizing the extraction results of high-resolution remote sensing images, characterized in that: Specifically include: Step 1: Construct an optimized dataset of the ground feature extraction results so that each group of samples has a one-to-one correspondence between the original image, the extraction result, and the true value label. Step 2: Use the optimized data set constructed in step 1 as input data to train the optimized model; The optimization model includes a feature extraction network, a generator network and a discriminator network. The feature extraction network can fuse the spatial details of low-scale features in the input data with the semantic information of high-scale features to extract features; the generator network converts the extracted intermediate feature representation back to the image space to generate an optimized result image; the discriminator network works in conjunction with the generator network and uses adversarial training, and can receive the optimized result image, the true value label reconstruction result and the true value label generated by the generator network, and use them as input for classification judgment; Step 3: Optimize the ground feature extraction results: Input the object extraction result to be optimized into the optimization model trained in step 2 to obtain the raster object extraction result; The raster feature extraction result is converted into a vector feature map by a raster-to-vector function; The vector feature map is processed by reducing the amount of point data and smoothing the outer contour burrs to obtain an optimized result in vector format, thereby completing the optimization of the feature extraction result.

2. The method for optimizing high-resolution remote sensing image object extraction results based on GAN according to claim 1, characterized in that: The feature extraction network is a multi-scale information extraction network, which uses SE_ResNeXt101 as the backbone network for feature extraction, and also includes a fusion module that can connect multiple features by combining 3×3Conv, BN and ReLU activation functions, and a 2×2 maximum pooling layer.

3. The method for optimizing high-resolution remote sensing image object extraction results based on GAN according to claim 1, characterized in that: The generator network includes multiple residual connection layers, 2×2 upsampling layers, BN, ReLU activation functions, and the last convolution layer uses a 1×1 convolution kernel and a Sigmoid activation function to generate an optimized result image; at the same time, combined with loss functions, including regularization loss functions, Potts loss and normalized cut loss are defined as: L Potts (G)=E x,z ∑ k S k W(1-S k ), Among them, S k is the vectorized representation of the kth channel of the softmax mask S, representing one of the optimized results output by the generator network, k = 2; W and is a pairwise discontinuity cost matrix that quantifies discontinuities between pixels or regions; Also includes the reconstruction loss function: L recG (G) = -E x,z [x·logG(x,z)] ensures that the data generated by the generator network is consistent with the true data in geometry, helping the generator network learn the distribution characteristics of the data.

4. The method for optimizing high-resolution remote sensing image object extraction results based on GAN according to claim 1, characterized in that: The discriminator network uses multiple 3×3Conv, 2×2 maximum pooling layers and fully connected layers to extract features and perform classification judgment.

5. The method for optimizing high-resolution remote sensing image object extraction results based on GAN according to claim 1, characterized in that: The discriminator network is trained adversarially with the generator network, and the generator network is trained adversarially against the loss function L GAN And the discriminator network adversarial loss function L D As shown below: L GAN (G,D)=E x,z [log(1-D(G(x,z)))], L D (G,D)=E y [log(y)]+E y [logD(y)]+E x,z [log(1-D(G(x,z)))], Among them, G(x,z) is the result generated by the generator network based on the input original image x and the extraction result z, y is the true value label, and D(y) and D(G(x,z)) are the outputs of the discriminator network for the true value label y and the generator network result respectively. The adversarial loss function can achieve a dynamic balance between data generation and discrimination, improve the effect of the generator network on optimizing the object extraction results, and enhance the ability of the discriminator network to distinguish the quality of the optimization results.

6. The method for optimizing high-resolution remote sensing image object extraction results based on GAN according to claim 1, characterized in that: The optimization of the ground feature extraction result in step 3 is specifically as follows: the ground feature extraction result to be optimized is input into the optimization model trained in step 2 to obtain the raster ground feature extraction result; The raster object extraction result is converted to vector by a raster-to-vector function, and parameters are set to retain the target object part and delete the background value part, so as to complete the raster-to-vector conversion and obtain a vector object patch; The vector feature map applies the Douglas-Peucker algorithm, sets the threshold parameters of the algorithm, iteratively processes the boundaries of each feature map, removes the details smaller than the threshold, and retains the main contour features; at the same time, the boundaries are smoothed to eliminate burrs, and after processing all feature maps, the optimization result of the final vector format is obtained.

Citation Information

Patent Citations

  • A method and system for extracting building images from remote sensing images

    CN110334719B