Landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology
By combining the generative adversarial network and the improved U-Net model, the problems of low-resolution image details loss and low recognition accuracy in landslide remote sensing image recognition are solved, and efficient detail recovery and accurate recognition of landslide remote sensing images are achieved.
Patent Information
- Application Number
- CN202510019405.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art faces the problems of low-resolution image details loss and low recognition accuracy in landslide remote sensing image recognition, especially in landslide area recognition in complex backgrounds, which are difficult to extract and restore subtle features.
By collecting multi-source remote sensing images and labeling landslide areas, unifying image size and resolution, data enhancement and standardization processing are performed to generate low-resolution images. Generative adversarial networks are designed to reconstruct high-resolution images through generators, introducing visual perception loss and pixel-level loss optimization detail recovery. Landslide segmentation was performed using an improved U-Net model, and landslide area details were determined in combination with encoder feature extraction, multi-scale cavity convolution and attention mechanism.
It effectively improves the detail recovery and recognition accuracy of landslide remote sensing images, solves the problem of low resolution image recognition accuracy, and provides a more accurate and reliable landslide monitoring solution.
Smart Images

Figure CN120107072A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of remote sensing intelligent identification, and more specifically to a method for remote sensing intelligent identification of landslide targets based on super-resolution reconstruction technology. Background Art
[0002] Landslide disasters, as one of the common geological disasters, have brought serious impacts on the natural environment and human society, especially in mountainous areas and high-risk areas. Landslide events often cause mountain collapse and traffic interruptions. With climate change, accelerated urbanization and the continuous deepening of human activities, the frequency and destructiveness of landslide disasters have gradually increased. Therefore, timely and accurate monitoring and identification of landslide areas has become an important task in geological disaster prevention and control and risk assessment.
[0003] Deficiencies of existing technologies: In landslide remote sensing image recognition, low-resolution image detail loss and low recognition accuracy are usually faced. Traditional remote sensing image processing methods often rely on low-resolution images for direct analysis, resulting in blurred boundaries and unclear details of landslide areas, which affects the recognition and segmentation accuracy of the model. Although some methods have improved data quality through data enhancement and annotation, they are still unable to effectively restore high-frequency details in low-resolution images, especially for landslide area recognition under complex backgrounds. Although existing landslide segmentation models (such as U-Net) perform well in image segmentation tasks, they have poor adaptability to low-resolution images and are difficult to extract and restore subtle features of landslide areas. In addition, the significant imbalance between landslide areas and background areas causes the model to make too many predictions about background areas during training, while being insufficiently sensitive to landslide areas. Summary of the invention
[0004] In order to overcome the above defects of the prior art, there is a solution as follows to solve the problem of low recognition accuracy of remote sensing images at low resolution in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] Collect multi-source remote sensing images and mark the landslide area, unify the image size and resolution, enhance and standardize the data, generate low-resolution images and save them;
[0007] Design a generative adversarial network and reconstruct high-resolution images through the generator, introduce visual perception loss and pixel-level loss to optimize detail restoration, and enhance the landslide area features according to the total loss function of the generator;
[0008] The improved U-Net model is used for landslide segmentation, and encoder feature extraction, multi-scale hole convolution and attention mechanism are combined to determine the details of the landslide area, and multiple loss functions are used to optimize the segmentation accuracy.
[0009] The difference in recognition accuracy between super-resolution reconstructed images and low-resolution images was compared, the segmentation accuracy was compared through evaluation indicators, and the improvement in the accuracy of landslide identification by super-resolution technology was verified.
[0010] In a preferred embodiment, multi-source remote sensing images are collected and landslide areas are marked, image size and resolution are unified, data enhancement and standardization are performed, low-resolution images are generated and saved, and the specific steps are as follows:
[0011] Multi-source remote sensing images include high-resolution satellite images, drone images, and public remote sensing datasets. The acquired multi-source remote sensing images are used to construct a landslide segmentation dataset, and the landslide segmentation dataset is annotated to record the image as a mask corresponding to the landslide area and the image;
[0012] Perform data enhancement on images in the landslide segmentation dataset, including rotation, flipping, cropping, and contrast adjustment;
[0013] Perform image cropping and standardization, cropping each image and the corresponding mask into a fixed-size window;
[0014] The cropped images are normalized to normalize the pixel values of each image to the interval [0, 1];
[0015] Downsampling the high-resolution image using the bicubic interpolation algorithm to generate a low-resolution image;
[0016] All images generated by data augmentation and the corresponding masks are used as augmented data sets;
[0017] The generated low-resolution images and corresponding masks are used as low-resolution datasets;
[0018] The high-resolution images and the corresponding landslide masks are stored as a high-resolution dataset.
[0019] In a preferred embodiment, a generative adversarial network is designed, and a high-resolution image is reconstructed through a generator, and visual perception loss and pixel-level loss are introduced to optimize detail restoration. The specific steps include:
[0020] Use a deep convolutional generative adversarial network to convert low-resolution images into high-resolution images. The deep convolutional generative adversarial network includes a generator network and a discriminator network.
[0021] The generator network uses a multi-layer convolutional neural network structure, adds a residual module and a skip connection generator, and introduces a sub-pixel convolution layer;
[0022] The loss function of the discriminator network is based on a combination of adversarial loss and pixel-level loss to train the generator and discriminator;
[0023] The Charbonnier loss is used to measure the pixel-level difference. The Charbonnier loss is defined as: Where N is the total number of pixels in the image, and I HR,i Generate images for and the real image I HR The i-th pixel value of ∈ is 0.001;
[0024] Perceptual loss calculates the high-level feature differences of images through convolutional neural networks to determine the visual similarity of images. Perceptual loss is defined as: Among them, φ k represents the feature map of the k-th convolutional network, K is the total number of layers in the network, is the L2 norm;
[0025] The adversarial loss is used to train the generator, and the adversarial loss is expressed as: Among them, P data represents data distribution, E is the expected value symbol, G SR is the image generated by the generator, D SR is the judgment of the discriminator on the generated image.
[0026] In a preferred embodiment, the landslide area characteristics are enhanced according to the total loss function of the generator, and the specific steps are as follows:
[0027] The total loss function of the generator is the weighted sum of pixel-level loss, perceptual loss, and adversarial loss, expressed as: L total =λ 1 L MSE +λ 2 L LPIPS +λ 3 L adv , where λ 1 ,λ 2 , and λ 3 It is the weight coefficient of each loss item, which is used to balance the impact of each loss on the training process.
[0028] In a preferred embodiment, an improved U-Net model is used for landslide segmentation, and encoder feature extraction, multi-scale hole convolution and attention mechanism are combined to determine the details of the landslide area, and multiple loss functions are used to optimize the segmentation accuracy, including the following steps:
[0029] Add multi-scale hole convolution modules and attention mechanisms to the U-Net architecture;
[0030] The dilated convolution module is applied in the encoder layer to extract multi-scale features at different levels: for the input feature map, the feature map obtained after the dilated convolution operation;
[0031] The attention mechanism focuses on the key parts of the landslide area in the image and suppresses the interference of the background area. For the feature map obtained after the input hole convolution operation, the attention mechanism calculates the weighted feature map;
[0032] Use cross entropy loss to measure the difference between the segmentation results predicted by the model and the actual labels;
[0033] Dice loss is used to measure the overlap between the landslide area predicted by the model and the actual landslide area;
[0034] The cross entropy loss and Dice loss are combined as the total loss function to balance the training objectives.
[0035] In a preferred embodiment, the difference in recognition accuracy between the super-resolution reconstructed image and the low-resolution image is compared, the segmentation accuracy is compared and evaluated by an indicator, and the improvement in the accuracy of landslide recognition by the super-resolution technology is verified, including the following steps:
[0036] Reconstruct low-resolution images using super-resolution models;
[0037] Perform image analysis using peak signal-to-noise ratio and structural similarity on the reconstructed super-resolution images and the real high-resolution images to determine their effectiveness;
[0038] When the validity is confirmed, the reconstructed high-resolution image is input into the landslide recognition model to obtain the landslide segmentation result;
[0039] The low-resolution image and the high-resolution image are input into the landslide identification model respectively to obtain the low-resolution image results and the high-resolution image results. The two sets of results are compared and analyzed to determine the degree of restoration of the details of the landslide area.
[0040] The technical effects and advantages of the landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology of the present invention are as follows:
[0041] The present invention constructs a data set of uniform size and resolution by collecting multi-source remote sensing images and marking landslide areas, and ensures the diversity and consistency of training data through data enhancement and standardization. At the same time, low-resolution images are generated and saved, providing high-quality input for super-resolution reconstruction and landslide identification. In the super-resolution reconstruction process, a generative adversarial network is designed, and a generator is used to reconstruct low-resolution images into high-resolution images. Visual perception loss and pixel-level loss are introduced to optimize detail recovery, especially the high-frequency details of the landslide area are enhanced through the attention mechanism, which improves the performance of the reconstructed image in practical applications.
[0042] In terms of landslide identification, an improved U-Net model was used, combined with encoder feature extraction, multi-scale hole convolution and attention mechanism, to improve the network's perception of landslide areas, and optimized the segmentation accuracy of the model in the case of imbalance between landslide areas and background areas through the fusion of multiple loss functions. By comparing the performance of super-resolution reconstructed images and low-resolution images in landslide identification tasks, the effectiveness of super-resolution technology in improving landslide identification accuracy was verified, and the detail restoration and identification accuracy of landslide remote sensing images were effectively improved, providing a more accurate and reliable solution for landslide monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 The figure is a flow chart of the landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In order to achieve the above objectives, Figure 1 A structural schematic diagram of a landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology of the present invention is given, which specifically includes the following steps:
[0046] Collect multi-source remote sensing images and mark the landslide area, unify the image size and resolution, enhance and standardize the data, generate low-resolution images and save them;
[0047] Design a generative adversarial network and reconstruct high-resolution images through the generator, introduce visual perception loss and pixel-level loss to optimize detail restoration, and enhance the landslide area features according to the total loss function of the generator;
[0048] The improved U-Net model is used for landslide segmentation, and encoder feature extraction, multi-scale hole convolution and attention mechanism are combined to determine the details of the landslide area, and multiple loss functions are used to optimize the segmentation accuracy.
[0049] The difference in recognition accuracy between super-resolution reconstructed images and low-resolution images was compared, the segmentation accuracy was compared through evaluation indicators, and the improvement in the accuracy of landslide identification by super-resolution technology was verified.
[0050] Step 1: Dataset construction and preprocessing. Remote sensing image data are collected from multiple sources, including high-resolution satellite images (such as WHU-RS19), drone images, and public remote sensing datasets. All data need to be labeled with landslide areas to ensure the accuracy and representativeness of the training data. Similarly, the construction process of the landslide segmentation dataset includes the selection of remote sensing images of a specific period of time (such as images of the vegetation peak period) and the annotation of landslide areas. After the data collection is completed, all image data (including high-resolution satellite images, drone images, remote sensing images, etc.) need to be converted into a unified size and resolution;
[0051] Set each image in the landslide segmentation dataset as I i , where i represents the i-th image in the dataset. In the landslide segmentation dataset annotation process, the landslide area is annotated for each image. And record the corresponding mask M i , that is, M i ∈{0, 1} indicates whether the pixel belongs to the landslide area: I i =Image(Source, R i ), where Image(Source, R i ) means extracting the annotated area from the source image and returning the corresponding image data I i and landslide mask data M i ;
[0052] Data augmentation is to expand the size of the training data set by performing a series of spatial transformations on each image, such as rotation, flipping, cropping, brightness / contrast adjustment, etc., to generate multiple transformed images;
[0053] Set the enhancement operation to T aug , including rotation matrix, flip operation and brightness adjustment, etc. After each enhancement operation, the image I′ i and its corresponding mask M′ i A new image and mask pair I′ will be generated i and M′ i :I′ i =T aug (I i ), M′ i =T aug(M i ), wherein, through such enhancement operations, N new image pairs are generated;
[0054] Perform image cropping and standardization, crop each image and its corresponding mask into a fixed-size window (such as 512×512 or 128x128 pixels), and extract sub-images by sliding windows. Set the cropping operation to C crop , for image I″ i And mask M″, the cutting process is as follows: I″ i =C crop (I′ i ), M″=C crop (M′ i ), where C crop An image of size (512×512) and the corresponding mask sub-image will be returned. This process will generate several cropped sub-images I″ i and mask M″;
[0055] The cropped image is normalized and the normalization operation is set to S norm (I″), whose goal is to normalize the pixel values of each image to the interval [0, 1]: This process maps the minimum and maximum values of the image to between 0 and 1, respectively, to ensure that the numerical range of the training data is consistent;
[0056] In the field of remote sensing, high resolution generally refers to remote sensing images with a spatial resolution of less than 5 meters. In some standards, remote sensing images with a pixel size of less than 1 meter (for example, satellite images with a resolution of 0.5 meters) can be distinguished according to actual needs in different scenarios. In this embodiment, a resolution of 0.5 meters is used as a high-resolution image;
[0057] High-resolution satellite images (such as WorldView, GeoIQ, etc.) usually refer to images with a resolution between 0.3 meters and 2 meters. These images can clearly identify the details of ground objects, such as buildings, roads, landslides, etc. UAV remote sensing images can often provide sub-meter or even centimeter-level high-resolution images due to their low flight altitude. They are suitable for scenes with high detail accuracy requirements such as landslide monitoring and urban planning.
[0058] Align the low-resolution generation with the high-resolution generation, and set the operation of generating the low-resolution image to G LR , a low-resolution image is generated by downsampling the high-resolution image, the downsampling factor is set to α (for example, a downsampling factor of 4 will reduce the image resolution to 1 / 4 of the original), and the low-resolution image is generated by downsampling using the bicubic interpolation algorithm: I LR =G LR (I″′i , α), where α is the downsampling factor, which controls the generation method of low-resolution images. The high-resolution image is downsampled to a low-resolution version through the bicubic interpolation method. Then, the correspondence between the low-resolution image and the high-resolution image is kept consistent through the mask for subsequent super-resolution reconstruction and landslide identification training;
[0059] After the above preprocessing steps, all data sets will be stored in the following form:
[0060] High-resolution dataset: stores high-resolution images I i and the corresponding landslide mask M i ;
[0061] Low-resolution dataset: stores the generated low-resolution images I LR and the corresponding mask M i ;
[0062] Enhanced dataset: stores all images I′ generated by data augmentation i And the corresponding mask M′ i ;
[0063] The data will be saved in a format suitable for model training (such as TF Record format or HDF5 format) and can be processed in batches through the data loader;
[0064] Step 2, super-resolution reconstruction model design and optimization. In this scheme, super-resolution reconstruction is the core technology for improving the resolution of landslide remote sensing images. First, a generative adversarial network (GAN) architecture is constructed, and a deep convolutional generative adversarial network (SRGAN) is used for image super-resolution reconstruction. The network consists of a generator and a discriminator. The generator is responsible for reconstructing low-resolution images into high-resolution images, while the discriminator optimizes the generator by judging the authenticity of the image. In order to improve the image reconstruction effect, especially the restoration of details and textures, a visual perception loss (such as LPIPS loss) is introduced in combination with the Charbonnier loss to optimize the details and visual quality of the super-resolution image, and an attention mechanism is added to the generator to improve the perception of high-frequency details (such as boundaries and cracks) in the landslide area, thereby improving the performance of the reconstructed image in practical applications. The specific steps are as follows:
[0065] In the super-resolution reconstruction task, the commonly used deep learning architecture is the Generative Adversarial Network (GAN), especially the Deep Convolutional Generative Adversarial Network (SRGAN), which can effectively convert low-resolution images into high-resolution images. SRGAN consists of two main modules: the generator network (GSR) and the discriminator network (DSR);
[0066] Generator network design, the goal of the generator is to generate a low-resolution input image I LR Generate high-resolution reconstructed images In order to restore the details of the image, the generator adopts a multi-layer convolutional neural network (CNN) structure and uses residual modules and jump connections to enhance the feature extraction capability. In order to improve the details and texture of the generated image, a sub-pixel convolution layer is introduced in the generator, which can preserve the spatial details of the image while restoring high resolution. SR Output image It can be expressed by the following formula: Among them, θ G is a learnable parameter in the representation generator;
[0067] The discriminator network design aims to determine whether the image is a real high-resolution image or a fake image generated by the generator. The discriminator outputs a binary result. To represent the authenticity of the image, the loss function of the discriminator will train the generator and discriminator based on the combination of adversarial loss and pixel-level loss. The discriminator's prediction of the image can be expressed by the following formula: Among them, σ is the Sigmoid activation function, θ D are the learnable parameters of the discriminator;
[0068] In the super-resolution reconstruction task, the design of the loss function directly affects the reconstruction effect of the generator. The pixel-level Charbonnier loss, perceptual loss (LPIPS) and adversarial loss are used. The pixel-level loss is to minimize the difference between the generated image and the real high-resolution image. The Charbonnier loss is used to measure the pixel-level difference, which is defined as: Where N is the total number of pixels in the image, and I HR,i Generate images for and the real image I HR The i-th pixel value of , ∈ is a small constant (usually 0.001, used to avoid zero division errors when the error is small;
[0069] Charbonnier loss has a smoother gradient than L2 loss, which makes it better to avoid the problem of gradient vanishing when training deep networks. Compared with L2 loss, Charbonnier loss can provide more accurate gradients when processing high-contrast parts in images, which helps to preserve details.
[0070] Perceptual loss mainly considers the visual perception effect of the image, not just the pixel-level difference. LPIPS (Learned Perceptual Image Patch Similarity) loss calculates the high-level feature differences of the image through a convolutional neural network (CNN) to measure the visual similarity of the image. It is defined as: Among them, φ k represents the feature map of the k-th convolutional network, K is the total number of layers in the network, is the L2 norm;
[0071] The adversarial loss is used to train the generator to make the images it generates more realistic. The goal of the generator is to "fool" the discriminator so that the discriminator cannot distinguish between generated images and real images. The adversarial loss can be expressed as: Among them, P data represents data distribution, E is the expected value symbol, G SR is the image generated by the generator, D SR is the discriminator’s judgment on the generated image;
[0072] In summary, the total loss function of the generator is the weighted sum of pixel-level loss, perceptual loss, and adversarial loss, which can be expressed as: L total =λ 1 L MSE +λ 2 L LPIPS +λ 3 L adv , where λ 1 ,λ 2 , and λ 3 It is the weight coefficient of each loss item, which is used to balance the impact of each loss on the training process;
[0073] The generator's loss function tells the model what kind of error should be minimized when generating images so that its output is closer and closer to the real image, thereby improving the image's pixel-level accuracy, visual perception quality, and structural similarity. By guiding the generator to optimize various aspects of image generation, including pixel-level accuracy, visual perception effects, and structural features, the generator is helped to gradually improve the quality of its output images.
[0074] Conduct model training and optimize the strategy during the model training process. Use the Adam optimizer to update the parameters of the generator and discriminator. Perform discriminator training first. First, fix the parameters of the generator and train the discriminator. Improve the classification ability of the discriminator by minimizing the adversarial loss.
[0075] With the parameters of the discriminator fixed, the generator is trained to optimize the quality of the generated images by minimizing the total loss.
[0076] Step 3, design a landslide recognition model, that is, after completing the design and optimization of the super-resolution reconstruction model, design and optimize the landslide recognition model, and intelligently identify and segment the landslide area based on the high-resolution image after super-resolution reconstruction. The improved U-Net architecture is used. The improved U-Net adds a residual module to the original encoder-decoder structure. The network's perception of the landslide area is improved through multi-level feature extraction. It handles targets of different scales that may appear in the landslide image. A multi-scale hole convolution module is designed, which can process convolution operations with different expansion rates in parallel and capture the characteristics of landslide areas at different scales. In addition, an adaptive feature extraction module is used to enable the model to dynamically adjust the feature extraction strategy according to the complexity of the landslide area. Finally, in order to solve the imbalance problem between the landslide area and the background area, a multi-loss function fusion strategy is designed to combine the Dice loss with the cross entropy loss to improve the model's sensitivity to landslide targets during segmentation. The specific steps are as follows:
[0077] The landslide recognition model uses convolutional neural network (CNN) as the basic architecture, especially the improved U-Net architecture. U-Net is a classic network architecture for image segmentation, which is characterized by a symmetrical encoder (downsampling) and decoder (upsampling) structure, and effectively transmits high-resolution features through skip connections.
[0078] Construct an improved U-Net architecture, which includes an encoder, a decoder, and skip connections. The encoder extracts multi-level features of the image through a series of convolution operations and gradually reduces the spatial resolution of the image. The decoder restores the spatial resolution of the image through deconvolution operations while retaining the extracted deep features. The skip connection is used to splice the features of the corresponding layer in the encoder with the corresponding layer in the decoder to help the network retain detailed information.
[0079] In order to better handle the detail differences between the landslide area and the background area, the improved U-Net adds a multi-scale hole convolution module and an attention mechanism on this basis, as follows:
[0080] In order to capture the multi-scale features of the landslide area, the improved U-Net network introduces a dilated convolution. The dilated convolution expands the convolution kernel and increases the perception of large-scale features without increasing the amount of calculation, thereby effectively capturing the details of landslide areas of different scales. The dilated convolution module is applied at different levels of the encoder to extract multi-scale features at different levels: For the input feature map F encode , the feature map F obtained after the dilated convolution operation dilated It can be expressed as: F dilated =C Dilated (F encode; k, r), where k is the size of the convolution kernel and r is the dilation rate, which determines the size of the receptive field. Through multi-scale dilated convolution, the network can extract details and global features at the same time, thereby improving the detection capability of landslide areas;
[0081] In order to improve the model's attention to the details of the landslide area, the improved U-Net introduces an attention mechanism, which enables the network to automatically focus on the key parts of the landslide area in the image and suppress the interference of the background area. In each layer of the decoder, the attention mechanism is used to weight the feature map so that the network can focus more on the high-frequency details of the landslide area, that is, the feature map F obtained after the input hole convolution operation dilated , the weighted feature map calculated by the attention mechanism is expressed as: F attended =A Attention (F dilated θ A ), where θ A is the learnable parameter of the attention mechanism, F attended It is the feature map after attention weighting;
[0082] In the landslide identification task, the commonly used loss functions are Cross-Entropy Loss and Dice loss. The combination of the two can effectively deal with the pixel imbalance problem between the landslide area and the background area.
[0083] The cross entropy loss is used to measure the difference between the segmentation results predicted by the model and the true label, for each pixel x and its corresponding predicted category probability and the true label y x , cross entropy loss L CE Defined as: Where N is the total number of pixels in the image, is the probability predicted by the model, y x is the true label;
[0084] The Dice loss is used to measure the overlap between the landslide area predicted by the model and the actual landslide area. The Dice loss is:
[0085] The total loss function combines the cross entropy loss and the Dice loss to achieve a balanced training goal. The total loss function can be defined as: L total =λ 1 L CE +λ 2 L Dice , where λ 1 and λ 2 is a hyperparameter that adjusts the importance of the two losses;
[0086] After training, the performance of the model was evaluated using common segmentation evaluation metrics (e.g., mean intersection over union (mIoU) (mIoU is calculated as the average of the IoUs of the landslide area and the background), Dice coefficient, accuracy, and F1 score).
[0087] For example, suppose there is a low-resolution remote sensing image of a landslide. The landslide area in the image is small and surrounded by a large amount of cluttered vegetation in the background. The image is reconstructed through a super-resolution model to obtain a high-resolution image. At this time, the landslide boundary in the image becomes clearer. The high-resolution image is input into the improved U-Net. The model ensures that the details of the landslide boundary are preserved through the residual module, and uses the multi-scale hole convolution module to capture the large-scale landslide area and small crack features at the same time. The model finally outputs the landslide segmentation image R, in which the segmentation effect of the landslide area and the background area is significantly improved, especially in complex background and detail recovery.
[0088] By using a weighted combination of Dice loss and cross entropy loss, the model can avoid over-prediction of the background area and ensure accurate identification of the landslide area. Ultimately, the output landslide area segmentation result has higher accuracy, especially in the precise extraction of the landslide boundary.
[0089] In summary, by designing an improved U-Net network and combining it with a multi-scale dilated convolution module and an attention mechanism, an efficient landslide recognition model is constructed. The model can accurately identify the landslide area while suppressing the interference of the background area. Through the combination of cross entropy loss and Dice loss, the model shows high accuracy in dealing with the imbalance problem between the landslide area and the background area.
[0090] Step 4: Compare and verify the super-resolution reconstruction and recognition results. After completing the design of the super-resolution reconstruction model and the landslide recognition model, the effectiveness of the method is verified through experiments. The high-resolution image after super-resolution reconstruction is compared with the original low-resolution image to evaluate the improvement of the reconstructed image in detail recovery and visual effects. These super-resolution images are input into the landslide recognition model to segment the landslide area. By comparing with the recognition results on the low-resolution data set, the improvement effect of super-resolution technology on the landslide recognition accuracy is evaluated. The specific steps are as follows:
[0091] Compare the reconstruction results using the super-resolution model (generator G SR ) for low-resolution images (I LR ) to reconstruct and evaluate the reconstructed image Compared with the original high-resolution image (I HR ), the specific steps are as follows;
[0092] Use the trained super-resolution model to perform super-resolution reconstruction on the low-resolution image to obtain the predicted high-resolution image: Among them, θ G are the learnable parameters (weights) of the model;
[0093] Super-resolution images generated by comparison and the real high-resolution image I LR , a series of image quality evaluation indicators can be used for quantitative analysis to determine the effectiveness of the super-resolution reconstruction model in detail recovery, image quality improvement and visual perception effect. Commonly used evaluation indicators include PSNR (peak signal-to-noise ratio) and SSIM (structural similarity);
[0094] PSNR measures the overall error of image reconstruction quality and is defined as: Among them, I max is the maximum pixel value of the image, MSE is the mean square error;
[0095] SSIM measures the structural similarity of images, and the formula is: Among them, μ I and are the means of the real image and the reconstructed image, respectively. and is the variance between the real image and the reconstructed image, is the covariance of the real image and the reconstructed image, C 1 and C 2 is a constant to prevent the denominator from being zero;
[0096] The comparison between the reconstructed image and the original image helps evaluate the ability of the super-resolution model to recover details and structures. PSNR focuses on the global error of the image, while SSIM considers the similarity of the structure.
[0097] After the super-resolution image is generated, the performance evaluation of the landslide recognition model is carried out. The reconstructed high-resolution image is input into the landslide recognition model G Land , and the results are quantitatively evaluated:
[0098] Use the reconstructed image As input, the landslide identification model G is input Land , and the landslide segmentation result is obtained Among them, θ G are the learnable parameters (weights) of the model;
[0099] The following indicators are mainly used to evaluate the segmentation results of the model:
[0100] The mean intersection over union (mIoU) is a measure of the overlap between the predicted area and the true area. and the true mask R Land , the calculation formula of average intersection-over-union ratio is:
[0101] The Dice coefficient measures the similarity between the predicted area and the true area, and is particularly suitable for dealing with unbalanced category problems. The calculation formula is:
[0102] The F1 score is a comprehensive evaluation index calculated by comprehensively considering the precision (Pre) and recall rate (Recall). The calculation formula is:
[0103] By evaluating the landslide recognition model, we can quantify the effect of super-resolution images on the improvement of landslide segmentation accuracy. The mIoU and Dice coefficients can effectively measure the overlap and similarity between the predicted and true masks, and the F1 score combines the precision and recall of the model.
[0104] Once the effectiveness is determined, the landslide identification results of the super-resolution image are compared with those of the low-resolution image, and the differences between the two in various evaluation indicators are analyzed:
[0105] The low-resolution image and the high-resolution image are input into the landslide recognition model respectively to obtain the low-resolution image results and the high-resolution image results. The two sets of results are compared and analyzed, and the aforementioned evaluation indicators (mIoU, Dice coefficient, F1 score, etc.) are used for quantitative comparison;
[0106] By comparing the visualized images of the landslide identification results (such as the binary images after segmentation), we further analyze whether the super-resolution reconstruction has effectively improved the detection accuracy of the landslide area. Through visual comparison of the images, we evaluate the degree to which the super-resolution technology can restore the details of the landslide area. After completing the above comparative analysis, we summarize the results of all evaluation indicators and derive the final impact of super-resolution reconstruction on the performance of the landslide identification model, thereby improving data quality and providing a fast path for subsequent mission objectives.
[0107] It should be noted that the threshold information in this embodiment is pre-set by professionals and will not be explained in detail here. In the embodiments, some parameter English letters have the same meanings when used, and they are explained with different meanings, which will not be explained one by one here.
[0108] The present invention constructs a data set of uniform size and resolution by collecting multi-source remote sensing images and marking landslide areas, and ensures the diversity and consistency of training data through data enhancement and standardization. At the same time, low-resolution images are generated and saved, providing high-quality input for super-resolution reconstruction and landslide identification. In the super-resolution reconstruction process, a generative adversarial network is designed, and a generator is used to reconstruct low-resolution images into high-resolution images. Visual perception loss and pixel-level loss are introduced to optimize detail recovery, especially the high-frequency details of the landslide area are enhanced through the attention mechanism, which improves the performance of the reconstructed image in practical applications.
[0109] In terms of landslide identification, an improved U-Net model was used, combined with encoder feature extraction, multi-scale hole convolution and attention mechanism, to improve the network's perception of landslide areas, and optimized the segmentation accuracy of the model in the case of imbalance between landslide areas and background areas through the fusion of multiple loss functions. By comparing the performance of super-resolution reconstructed images and low-resolution images in landslide identification tasks, the effectiveness of super-resolution technology in improving landslide identification accuracy was verified, and the detail restoration and identification accuracy of landslide remote sensing images were effectively improved, providing a more accurate and reliable solution for landslide monitoring.
[0110] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0111] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0112] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0113] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0114] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0115] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology, characterized by: The steps include: Collect multi-source remote sensing images and mark the landslide area, unify the image size and resolution, enhance and standardize the data, generate low-resolution images and save them; Design a generative adversarial network and reconstruct high-resolution images through the generator, introduce visual perception loss and pixel-level loss to optimize detail restoration, and enhance the landslide area features according to the total loss function of the generator; The improved U-Net model is used for landslide segmentation, and encoder feature extraction, multi-scale hole convolution and attention mechanism are combined to determine the details of the landslide area, and multiple loss functions are used to optimize the segmentation accuracy. The difference in recognition accuracy between super-resolution reconstructed images and low-resolution images was compared, the segmentation accuracy was compared through evaluation indicators, and the improvement in the accuracy of landslide identification by super-resolution technology was verified.
2. The landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology according to claim 1 is characterized in that: Collect multi-source remote sensing images and mark the landslide area, unify the image size and resolution, enhance and standardize the data, generate low-resolution images and save them. The specific steps are as follows: Multi-source remote sensing images include high-resolution satellite images, drone images, and public remote sensing datasets. The acquired multi-source remote sensing images are used to construct a landslide segmentation dataset, and the landslide segmentation dataset is annotated to record the image as a mask corresponding to the landslide area and the image; Perform data enhancement on images in the landslide segmentation dataset, including rotation, flipping, cropping, and contrast adjustment; Perform image cropping and standardization, cropping each image and the corresponding mask into a fixed-size window; Normalize the cropped images and normalize the pixel values of each image to [0, 1]; interval; Downsampling the high-resolution image using the bicubic interpolation algorithm to generate a low-resolution image; All images generated by data augmentation and the corresponding masks are used as augmented data sets; The generated low-resolution images and corresponding masks are used as low-resolution datasets; The high-resolution images and the corresponding landslide masks are stored as a high-resolution dataset.
3. The landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology according to claim 2 is characterized in that: Design a generative adversarial network and reconstruct high-resolution images through the generator. Introduce visual perception loss and pixel-level loss to optimize detail recovery. The specific steps include: Use a deep convolutional generative adversarial network to convert low-resolution images into high-resolution images. The deep convolutional generative adversarial network includes a generator network and a discriminator network. The generator network uses a multi-layer convolutional neural network structure, adds a residual module and a skip connection generator, and introduces a sub-pixel convolution layer; The loss function of the discriminator network is based on a combination of adversarial loss and pixel-level loss to train the generator and discriminator; The Charbonnier loss is used to measure the pixel-level difference. The Charbonnier loss is defined as: Where N is the total number of pixels in the image, and I HR,i Generate images for and the real image I HR The i-th pixel value of ∈ is 0.001; Perceptual loss calculates the high-level feature differences of images through convolutional neural networks to determine the visual similarity of images. Perceptual loss is defined as: Among them, φ k represents the feature map of the k-th convolutional network, K is the total number of layers in the network, is the L2 norm; The adversarial loss is used to train the generator, and the adversarial loss is expressed as: Among them, P data represents data distribution, E is the expected value symbol, G SR is the image generated by the generator, D SR is the judgment of the discriminator on the generated image.
4. The landslide target remote sensing intelligent identification method based on super-resolution reconstruction technology according to claim 3 is characterized in that: The landslide area features are enhanced according to the total loss function of the generator. The specific steps are as follows: The total loss function of the generator is the weighted sum of pixel-level loss, perceptual loss, and adversarial loss, expressed as: L total =λ1L MSE +λ2L LPIPS +λ3L adv , where λ1, λ2, and λ3 are weight coefficients of each loss term, which are used to balance the impact of each loss on the training process.
5. The method for remote sensing intelligent identification of landslide targets based on super-resolution reconstruction technology according to claim 4 is characterized in that: The improved U-Net model is used for landslide segmentation. The encoder feature extraction, multi-scale hole convolution and attention mechanism are combined to determine the details of the landslide area. The segmentation accuracy is optimized using multiple loss functions, including the following steps: Add multi-scale hole convolution modules and attention mechanisms to the U-Net architecture; The dilated convolution module is applied in the encoder layer to extract multi-scale features at different levels: for the input feature map, the feature map obtained after the dilated convolution operation; The attention mechanism focuses on the key parts of the landslide area in the image and suppresses the interference of the background area. For the feature map obtained after the input hole convolution operation, the attention mechanism calculates the weighted feature map; Use cross entropy loss to measure the difference between the segmentation results predicted by the model and the actual labels; Dice loss is used to measure the overlap between the landslide area predicted by the model and the actual landslide area; The cross entropy loss and Dice loss are combined as the total loss function to balance the training objectives.
6. The method for remote sensing intelligent identification of landslide targets based on super-resolution reconstruction technology according to claim 5 is characterized in that: Compare the difference in recognition accuracy between super-resolution reconstructed images and low-resolution images, evaluate the segmentation accuracy through indicators, and verify the improvement of landslide recognition accuracy by super-resolution technology, including the following steps: Reconstruct low-resolution images using super-resolution models; Perform image analysis using peak signal-to-noise ratio and structural similarity on the reconstructed super-resolution images and the real high-resolution images to determine their effectiveness; When the validity is confirmed, the reconstructed high-resolution image is input into the landslide recognition model to obtain the landslide segmentation result; The low-resolution image and the high-resolution image are input into the landslide identification model respectively to obtain the low-resolution image results and the high-resolution image results. The two sets of results are compared and analyzed to determine the degree of restoration of the details of the landslide area.
Citation Information
Cited By
Low-quality image-oriented generalization method in field of collaborative optimization tool wear recognition
CN120747613A
Low-resolution remote sensing image enhancement method and system based on multiple time phases
CN121095064A