Visual perception guided multi-scale feature fusion landslide image super-resolution model

Through visual perception-guided multi-scale features fusion of landslide image super-resolution model, the problems of insufficient reduction of boundary details and texture features of landslide image in the prior art are solved, and efficient landslide image super-resolution reconstruction and recognition effects are achieved.

CN119991439APending Publication Date: 2025-05-13SICHUAN CHUANJIAO ROAD & BRIDGE +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411910530.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the case of high-resolution image acquisition cost and difficult data acquisition, existing landslide image processing methods are difficult to effectively restore the boundary details and texture characteristics of landslides, and there are problems of artifacts or excessive smoothing.

Method used

A visual perception-guided multi-scale feature fusion landslide image super-resolution model is proposed, including generator network, discriminator network and comprehensive loss function module. The generator network extracts and fuses image features through head convolution modules, feature enhancement modules, upsampling reconstruction modules and branch network modules; the discriminator network performs adversarial training through multi-layer convolutional structures; the comprehensive loss function combines content loss, adversarial loss and perceived loss to optimize the performance of the generator.

Benefits of technology

It significantly improves the super-resolution effect of landslide images, improves landslide details processing, improves image quality and recognition accuracy, reduces artifact phenomena, and the generated high-resolution images are clearer and more realistic, and conforms to the visual characteristics of real images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991439A_ABST
    Figure CN119991439A_ABST
Patent Text Reader

Abstract

The invention discloses a visual perception guided multi-scale feature fusion landslide image super-resolution model, which relates to the technical field of remote sensing image processing and comprises a generator network, a discriminator network and a comprehensive loss function module. The generator network comprises a head convolution module, a feature enhancement module, an up-sampling reconstruction module and a branch network module; the feature extraction module is used for extracting high-frequency details, realizing amplification of a feature map, enhancing important pixels and generating a weight coefficient for weighted fusion; the discriminator network is used for discriminating whether the input image is a real high-resolution image; and the comprehensive loss function module is used for setting the weighted sum of the content loss, the adversarial loss and the perception loss as total loss. Through the generator network, the discriminator network and the comprehensive loss function, the super-resolution effect of the landslide image is remarkably improved, high-frequency details and boundary information are effectively extracted, the artifact phenomenon is smoothed, and the generated high-resolution image is clearer and more vivid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing image processing, and in particular to a multi-scale feature fusion landslide image super-resolution model guided by visual perception. Background Art

[0002] Landslide is one of the main geological disasters that threatens human life and property safety. Its early identification and monitoring are of great significance to disaster prevention and control and engineering construction. High-resolution remote sensing images are the key data source for landslide identification. However, in practical applications, due to the high cost of obtaining high-resolution remote sensing images and the difficulty of data collection, only low-resolution image data can be obtained. Low-resolution images are insufficient in expressing the range and boundary information of landslide areas, which directly affects the accuracy and reliability of landslide identification.

[0003] Existing landslide image processing methods mainly include traditional image enhancement technology and image super-resolution reconstruction technology based on deep learning. Traditional technology usually enlarges the image through simple algorithms such as interpolation, but it cannot effectively restore the boundary details and texture features of the landslide; the generative adversarial network (GAN) based on deep learning shows good potential in image super-resolution reconstruction tasks, but the images it generates often have artifacts or over-smoothing problems, which are difficult to meet the needs of landslide identification.

[0004] In addition, the existing super-resolution methods are designed to focus more on the restoration effect of general images, and lack the ability to extract high-frequency features of landslide images, a special target. Especially when the vegetation coverage is high or the landslide area is relatively hidden, the boundary extraction ability and response degree of the existing models are low. Summary of the invention

[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a visual perception-guided multi-scale feature fusion landslide image super-resolution model to solve the above-mentioned technical problems.

[0006] To achieve the above-mentioned object, the present invention provides the following technical solutions: a multi-scale feature fusion landslide image super-resolution model guided by visual perception, comprising a generator network, a discriminator network and a comprehensive loss function module;

[0007] The generator network includes a head convolution module, a feature enhancement module, an upsampling reconstruction module and a branch network module; the head convolution module is used to extract high-frequency details in low-resolution landslide images, and a small-size convolution kernel is used to capture the boundary information of the landslide area; the feature enhancement module is used to extract landslide image features by combining deep and shallow layer features through a residual block embedded with an attention mechanism; the upsampling reconstruction module is used to realize the amplification of feature maps by alternating deconvolution layers and PA layers, and enhance important pixels, and the branch network module is used to perform feature fusion on the original input image and the image after interpolation and upsampling, and generate weight coefficients for weighted fusion;

[0008] The discriminator network is used to adapt to the training requirements of the generative adversarial network through a multi-layer convolution structure to discriminate whether the input image is a real high-resolution image;

[0009] The comprehensive loss function module is used to set the weighted sum of content loss, adversarial loss and perceptual loss as the overall loss; the content loss is used for image generation tasks to improve the model's sensitivity to landslide details; the adversarial loss includes generator loss and discriminator loss; the perceptual loss accurately evaluates the visual difference between super-resolution images and real images by learning the inverse mapping between generated images and real images.

[0010] The present invention is further configured such that the head convolution module includes a 3×3 small-size convolution kernel for capturing high-frequency edge details of the image;

[0011] The feature enhancement module is composed of 6 EDCA residual blocks and a convolution layer with a convolution kernel size of 3×3 in a stacked manner;

[0012] The up-sampling reconstruction module is composed of multiple Up-Blocks and PA layers alternately, wherein the Up-Block is used for up-sampling, and the PA layer is used to enhance important pixels. Each Up-Block includes a deconvolution layer and two convolution layers, and the deconvolution layer is used to up-sample the input feature map;

[0013] The branch network module consists of a convolutional layer, a leaky ReLU activation function and an adaptive average pooling layer, which is used to extract the features of the fused image. The adaptive average pooling layer is used to convert the feature map into a vector of length 64, and the vector is mapped to a vector of length 3 through a linear layer, and then normalized using a softmax layer. The result is the weight coefficient of the two images.

[0014] The present invention is further configured such that the discriminator network is composed of a 64-channel convolution layer with a step size of 1 and multiple repeated convolution structures. The step size of the convolution layer is 2. The length and width of the original input image are changed to 1 / 16 of the original image. The last channel is 512-dimensional. After the output, an adaptive average pooling and LReLU function activation function are passed, and then a convolution layer with a dimension of 1 is connected to obtain the output result.

[0015] The present invention is further configured to include a coordinate attention module, which captures spatial long-range correlation and channel global features by encoding along the horizontal and vertical coordinate directions, thereby enhancing the feature extraction capability of the landslide area.

[0016] The present invention is further configured to further include a multi-scale feature fusion module, and the multi-scale feature fusion module is implemented by the following steps:

[0017] The input image is feature extracted through convolutional layers and upsampling layers, which can capture information of different scales from local to global.

[0018] The extracted features are concatenated with the original image in the channel dimension, where concatenation in the channel dimension includes: using a fully connected layer to calculate the weights of features of different scales and normalizing them through a softmax function; performing weighted summation on the feature maps according to the weights to obtain a fused high-resolution feature representation.

[0019] The present invention is further configured such that the calculation logic of the content loss is: Among them, L char is the content loss, x is the difference between the predicted value and the true value, and ε is a small constant used to ensure the differentiability of the function when x is zero;

[0020] The calculation logic of the generator loss is: in, is the generator loss, Represents the real image distribution X r The expected value of Indicates the probability that the real image sample is more real than the generated sample, and the value tends to 1. Represents the generated image distribution X f The expected value of Indicates the probability that the generated image is less realistic than the real image, and the value is closer to 0;

[0021] The calculation logic of the discriminator loss is: in, is the discriminator loss;

[0022] The calculation logic of the perception loss is: Among them, L adv is the perceptual loss, d(x,x0) is the distance function, which measures the difference between the input image x and the target image x0 in a specific feature space, ∑ l To sum the features of the lth layer in the network, H l and W l Represents the height and width of the feature map of the lth layer, which is used to normalize the feature difference of each layer, ∑ h,w To sum the spatial dimensions of the feature map, which include height h and width w, calculate the overall difference on the feature map of a specific layer, w l is the weight of the l-th layer feature, which is used to adjust the contribution of different layer features to the final loss value. ⊙ represents element-by-element multiplication, that is, element-by-element weighted operation. and is the feature value of the h,wth position in the normalized feature map of the input image x and the target image x0 at the lth layer, is the square of the Euclidean norm, which is used to measure the difference between two sets of eigenvalues;

[0023] The calculation logic of comprehensive loss is: Among them, L G is the comprehensive loss, L per is the adversarial loss, α, and β are coefficients that balance the different loss terms.

[0024] The present invention is further configured to adopt a dual time scale and mixed precision training strategy, wherein the dual time scale and mixed precision training strategy includes dual time scale training and mixed precision training;

[0025] The dual time scale training uses two different learning rates to update the parameters of the network. The larger learning rate is used to update the low-level parameters in the network, quickly converge and learn the low-level features; the smaller learning rate is used to update the high-level parameters in the network and accurately adjust the high-level features.

[0026] The mixed precision training balances the training speed and the training accuracy of the final model by utilizing the advantages of both half-precision and full-precision floating-point numbers, and effectively reduces memory consumption and computational overhead by using half-precision floating-point numbers for calculations during forward propagation and back-propagation. Since the limited representation range of half-precision floating-point numbers may lead to numerical overflow and precision loss, key operations are calculated using full-precision floating-point numbers to ensure the accuracy of the model.

[0027] The present invention provides a visual perception-guided multi-scale feature fusion landslide image super-resolution model, comprising a generator network, a discriminator network and a comprehensive loss function module; the generator network comprises a head convolution module, a feature enhancement module, an upsampling reconstruction module and a branch network module; the head convolution module is used to extract high-frequency details in low-resolution landslide images, and a small-size convolution kernel is used to capture the boundary information of the landslide area; the feature enhancement module is used to extract landslide image features by combining deep and shallow layer features through a residual block embedded with an attention mechanism; the upsampling reconstruction module is used to realize feature map amplification by alternating deconvolution layers and PA layers, and enhance important pixels, and the branch network module is used to The original input image and the image after interpolation and upsampling are feature fused to generate weight coefficients for weighted fusion; the discriminator network is used to adapt to the training requirements of the generative adversarial network through a multi-layer convolution structure to determine whether the input image is a real high-resolution image; the comprehensive loss function module is used to set the weighted sum of content loss, adversarial loss and perceptual loss as the overall loss; the content loss is used for image generation tasks to improve the model's sensitivity to landslide details; the adversarial loss includes generator loss and discriminator loss; the perceptual loss accurately evaluates the visual difference between the super-resolution image and the real image by learning the inverse mapping between the generated image and the real image, and the beneficial effects produced include:

[0028] 1. The super-resolution effect is remarkable, and the details of the landslide are processed perfectly. The VGG network features trained in the ImageNet classification task show remarkable results when used in the loss function of image synthesis training. This metric standard promotes the generator to master the ability to reconstruct real images from synthetic images by learning the inverse mapping between generated images and real images, and gives priority to the similarity between the two at the perceptual level;

[0029] 2. High interpretability. Based on the comprehensive evaluation of model performance and the visual analysis of the landslide target response during the training process, a series of quantitative and qualitative experiments were conducted to comprehensively evaluate the performance of super-resolution reconstruction of landslide remote sensing images and landslide identification. The experimental results show that the proposed model outperforms traditional methods and benchmark models in multiple evaluation indicators, verifying the effectiveness and superiority of the model; and the training process of the two models was visualized and analyzed, especially the visualization comparison of the activation degree under different loss functions, realizing the visual interpretation of the comprehensive loss function for visual perception enhancement;

[0030] 3. Significantly improve the image quality. By introducing the head convolution module and the multi-scale feature fusion module, the high-frequency details and boundary information of the landslide image are effectively extracted, which solves the problem of insufficient detail restoration in the landslide area by the traditional interpolation method and the general super-resolution method. The generated high-resolution image is clearer and more realistic. By designing the EDCA residual block and introducing the coordinate attention mechanism in the generator network, the model can effectively capture the high-frequency features of the landslide boundary, improve the recognition accuracy of the landslide area and boundary, and provide more reliable data support for the monitoring of hidden landslides. The upsampling reconstruction module with alternating PA layers and deconvolution layers is used to effectively smooth the artifact phenomenon and improve the visual effect of the generated image. At the same time, a comprehensive loss function (including content loss, adversarial loss and perceptual loss) is used to ensure that the generated high-resolution image is visually closer to the real image.

[0031] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0033] Figure 1 A training flow chart of a multi-scale feature fusion landslide image super-resolution model guided by visual perception is shown as an exemplary embodiment of the present invention;

[0034] Figure 2 A network model framework diagram for generating a multi-scale feature fusion landslide image super-resolution model guided by visual perception according to an exemplary embodiment of the present invention;

[0035] Figure 3 A discriminant network model framework diagram of a multi-scale feature fusion landslide image super-resolution model guided by visual perception is shown as an exemplary embodiment of the present invention;

[0036] Figure 4 A structural diagram of a coordinate attention module of a multi-scale feature fusion landslide image super-resolution model guided by visual perception according to an exemplary embodiment of the present invention;

[0037] Figure 5The SK-Net model framework diagram of a multi-scale feature fusion landslide image super-resolution model guided by visual perception is shown as an exemplary embodiment of the present invention;

[0038] Figure 6 A multi-scale feature fusion module diagram of a landslide image super-resolution model guided by visual perception and fused with multi-scale features is shown as an exemplary embodiment of the present invention;

[0039] Figure 7 A flowchart of LPIPS index calculation of a multi-scale feature fusion landslide image super-resolution model guided by visual perception is shown as an exemplary embodiment of the present invention;

[0040] Figure 8 A detail comparison diagram of a reconstructed image of a multi-scale feature fusion landslide image super-resolution model guided by visual perception, shown as an exemplary embodiment of the present invention;

[0041] Fig. 9 A visualization result diagram of a generator before and after improvement of a multi-scale feature fusion landslide image super-resolution model guided by visual perception according to an exemplary embodiment of the present invention;

[0042] Fig.10 A visualization result diagram of the generator under different loss functions of the FFSRGAN of the multi-scale feature fusion landslide image super-resolution model guided by visual perception is shown as an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0043] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.

[0044] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0045] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0046] The multi-scale feature fusion landslide image super-resolution model guided by visual perception includes:

[0047] Generator network, discriminator network and comprehensive loss function module;

[0048] The generator network includes a head convolution module, a feature enhancement module, an upsampling reconstruction module and a branch network module; the head convolution module is used to extract high-frequency details in low-resolution landslide images, and a small-size convolution kernel is used to capture the boundary information of the landslide area; the feature enhancement module is used to extract landslide image features by combining deep and shallow layer features through a residual block embedded with an attention mechanism; the upsampling reconstruction module is used to realize the amplification of feature maps by alternating deconvolution layers and PA layers, and enhance important pixels, and the branch network module is used to perform feature fusion on the original input image and the image after interpolation and upsampling, and generate weight coefficients for weighted fusion;

[0049] The discriminator network is used to adapt to the training requirements of the generative adversarial network through a multi-layer convolution structure to discriminate whether the input image is a real high-resolution image;

[0050] The comprehensive loss function module is used to set the weighted sum of content loss, adversarial loss and perceptual loss as the overall loss; the content loss is used for image generation tasks to improve the model's sensitivity to landslide details; the adversarial loss includes generator loss and discriminator loss; the perceptual loss accurately evaluates the visual difference between super-resolution images and real images by learning the inverse mapping between generated images and real images.

[0051] Specifically, the generative adversarial network is a deep learning model consisting of two parts: the generator and the discriminator. The function of the generator is to generate pseudo images, while the function of the discriminator is to determine whether an image is real or fake. In the super-resolution reconstruction task, the low-resolution image (LR) is converted into a high-resolution image (SR) by the generator, and the features of the SR and the real high-resolution image are compared through the VGG network to obtain the G-Loss. At the same time, the discriminator compares the SR with the real high-resolution image to obtain the D-Loss. The generator updates the parameters according to the G-Loss and D-Loss through back propagation to improve the performance of the generator.

[0052] In the adversarial training between the cyclic generator and the discriminator, the generator continuously creates fake images and submits them to the discriminator for authenticity identification. Once the generated image quality is high enough to "confuse" the discriminator, the image produced by the generator can be used to perform super-resolution reconstruction tasks. However, methods that rely purely on generative adversarial networks may introduce a certain amount of noise and artifacts in the reconstruction process, which may have an adverse effect on the quality of super-resolution reconstructed images.

[0053] To solve this problem, SRGAN introduces the concepts of residual learning and perceptual loss. By introducing residual blocks in the generator, the generation of noise and artifacts can be reduced; by adding perceptual loss to G-Loss, the similarity between the generated SR image and the real high-resolution image can be guaranteed. The introduction of these technologies can improve the performance of the SRGAN generator, thereby improving the quality of super-resolution reconstructed images. The specific training process is as follows Figure 1 shown.

[0054] The present invention is further configured such that the head convolution module includes a 3×3 small-size convolution kernel for capturing high-frequency edge details of the image;

[0055] The feature enhancement module is composed of 6 EDCA residual blocks and a convolution layer with a convolution kernel size of 3×3 in a stacked manner;

[0056] The up-sampling reconstruction module is composed of multiple Up-Blocks and PA layers alternately, wherein the Up-Block is used for up-sampling, and the PA layer is used to enhance important pixels. Each Up-Block includes a deconvolution layer and two convolution layers, and the deconvolution layer is used to up-sample the input feature map;

[0057] The branch network module consists of a convolutional layer, a leaky ReLU activation function and an adaptive average pooling layer, which is used to extract the features of the fused image. The adaptive average pooling layer is used to convert the feature map into a vector of length 64, and the vector is mapped to a vector of length 3 through a linear layer, and then normalized using a softmax layer. The result is the weight coefficient of the two images.

[0058] Specifically, the present invention designs a FFSRGAN network to improve the insufficient edge detail information and "artifact" phenomenon in the reconstructed landslide image. The specific network framework is as follows: Figure 2 shown.

[0059] In the generative adversarial network in the field of super-resolution reconstruction, the core function of the generator is to receive low-resolution images as input and output the corresponding high-resolution version of the image through the processing flow of the neural network. The network structure of the FFSRGAN model generator can be divided into the head convolution module A, the feature enhancement module B, the upsampling reconstruction module C, and the branch network module. In the forward propagation process, the input image first passes through the head convolution layer, and then the feature extraction is performed through the main layer. Next, the features are upsampled through the upsampling layer, and then the generated high-resolution image is obtained through the final convolution layer and activation function. At the same time, the input image is fused with the interpolated upsampled image and passed to the branch network for feature extraction. The obtained feature vector is weighted by the linear layer and normalized by softmax. Finally, the weighted sum of the fused image and the weight is returned as the final generated result.

[0060] (1) Head convolution module A: This layer uses a small 3×3 convolution kernel, which is specially designed to capture the high-frequency edge details of the image. This design choice aims to effectively extract the detail information in the landslide image while optimizing the computational efficiency, so as to maximize the retention of the key visual features of the image while reducing the computational burden.

[0061] (2) Feature enhancement module B is composed of 6 EDCA residual blocks and a convolutional layer with a convolution kernel size of 3×3. Considering that more layers and connections can improve network performance, the stacking of residual blocks is conducive to deep feature extraction. At the same time, the residual connection method can effectively alleviate the "gradient vanishing" phenomenon caused by too many network layers.

[0062] (3) Upsampling reconstruction module C. The goal of the upsampling layer is to enlarge the low-resolution feature map to a high resolution through interpolation operations, so that the final generated image can be clearer and richer in details. Specifically, the upsampling layer is composed of multiple Up-Blocks and PA layers alternating, where Up-Block is used for upsampling and PA layer is used to enhance important pixels. Each Up-Block includes a deconvolution layer (i.e., transposed convolution layer) and two convolution layers, where the function of the deconvolution layer is to upsample the input feature map. Since the deconvolution operation will cause jagged artifacts in the image, a convolution layer is needed to smooth the feature map after deconvolution.

[0063] (4) Branch network module D. The branch network is a network structure used to extract features from two input images, calculate the feature similarity of the two images, and generate weight coefficients for weighted fusion of the two images. Specifically, the branch network consists of a convolution layer, a Leaky ReLU activation function, and an adaptive average pooling layer, which is used to extract the features of the fused image. The adaptive average pooling layer is used to convert the feature map into a vector of length 64, and finally a linear layer is used to map the vector to a vector of length 3, and then a softmax layer is used for normalization. The result is the weight coefficient of the two images.

[0064] The present invention is further configured such that the discriminator network is composed of a 64-channel convolution layer with a step size of 1 and multiple repeated convolution structures. The step size of the convolution layer is 2. The length and width of the original input image are changed to 1 / 16 of the original image. The last channel is 512-dimensional. After the output, an adaptive average pooling and LReLU function activation function are passed, and then a convolution layer with a dimension of 1 is connected to obtain the output result.

[0065] Specifically, the discriminator network structure of the FFSRGAN model is as follows: Figure 3 As shown in the figure, there is only one convolution layer with 64 channels and a stride of 1, as well as multiple repeated convolution structures. The stride of the convolution layer is 2. The length and width of the original input image become 1 / 16 of the original image. The last channel is 512-dimensional. After the output, it passes through an adaptive average pooling and LReLU function activation function, and then a convolution layer with a dimension of 1 is used to output the result. In order to simplify the calculation process and avoid introducing the maximum pooling layer in the network structure, LReLU is used as the activation function and adaptive average pooling. As the number of network model layers increases, the feature extraction capability is enhanced, the number of image features increases, and the convolution size of each extracted feature decreases. The LReLU activation function works better than the ReLU function when the value is negative, and sparse gradients can be avoided. Therefore, LReLU is selected as the activation function. Adaptive average pooling can also extract features more effectively, making image processing more accurate and saving a lot of time.

[0066] The present invention is further configured to include a coordinate attention module, which captures spatial long-range correlation and channel global features by encoding along the horizontal and vertical coordinate directions, thereby enhancing the feature extraction capability of the landslide area.

[0067] Specifically, the coordinate attention module reduces the consumption of computing resources by removing the BN layer, which is suitable for high-level computer vision tasks such as classification but not for low-level computer vision tasks such as SR. When performing super-resolution reconstruction tasks, the batch normalization layer normalizes the network features, which may limit the dynamic range of the residual module and increase the burden on GPU memory. In view of this, in order to optimize the generator structure, we adjusted the architecture and removed the BN layer. Inserting a CA module that focuses on position information improves feature extraction capabilities. The residual dense block structure is as follows Figure 4 (c) as shown.

[0068] The feature extraction module of SRGAN can only capture local relationships, but ignores overall relationships. In landslide image enhancement, the entire image often appears inconsistent. In areas where image information is severely lost, it is difficult to extract effective features to represent local information and then use the extracted features to restore the blurred image.

[0069] To solve the problem of insufficient feature extraction, Coordinate Attention is introduced into SRGAN, and its structure is as follows Figure 4 As shown in (d), for a given input, the coordinate attention block encodes each channel along the horizontal and vertical coordinates respectively, producing a pair of direction-aware feature maps. The above operation can capture long-range correlations along one spatial direction and retain position information along another spatial direction to help the network locate the object of interest more accurately; it also considers attention in the channel dimension and the spatial dimension at the same time, and can learn adaptive channel weights so that the model pays more attention to useful channel information. Channel information usually represents the global features of the image, such as the color, brightness, etc. of the image. Coordinate attention is used to capture the edge enhancement of landslide images, making full use of effective information, making up for the shortcomings of local extraction by convolution operations, and having the advantage of extracting global features.

[0070] Pixel attention is different from channel attention and spatial attention. Pixel attention can generate a 3D attention map. This attention strategy introduces fewer additional parameters, but can help generate better SR results. The network structure is also very simple, that is, 1x1 convolution followed by Sigmoid activation function. This attention can strengthen the weight of some important pixels at the pixel level to attract the model's attention.

[0071] The present invention is further configured to further include a multi-scale feature fusion module, and the multi-scale feature fusion module is implemented by the following steps:

[0072] The input image is feature extracted through convolutional layers and upsampling layers, which can capture information of different scales from local to global.

[0073] The extracted features are concatenated with the original image in the channel dimension, where concatenation in the channel dimension includes: using a fully connected layer to calculate the weights of features of different scales and normalizing them through a softmax function; performing weighted summation on the feature maps according to the weights to obtain a fused high-resolution feature representation.

[0074] Specifically, the multi-scale feature fusion module transforms the technology used in the SK-Net framework for adaptively adjusting the receptive field into an explicit feature fusion strategy. Its network framework is as follows: Figure 5 As shown in the figure. In SK-Net, the network adapts to features of different scales by dynamically selecting different convolution kernel sizes, while our method uses a fully connected layer to calculate the weight of each feature after extracting features at different levels of the network. These weights are processed by the softmax function to ensure their balance and complementarity during the feature fusion process. This fusion strategy utilizes the intrinsic connection between high-level information, making the generated HR images more realistic and accurate.

[0075] First, the input image is feature extracted through convolutional layers and upsampling layers. These layers can capture different scale information from local to global, such as Figure 6 As shown in Figure 2. These features are then concatenated with the original image in the channel dimension to retain more contextual information. Unlike the traditional simple feature addition method, this module exploits the intrinsic connections between high-level features and pays special attention to and strengthens these high-level features through the branch network, thereby achieving more effective feature fusion in the generative network and further refining high-level feature representations that can capture more abstract information.

[0076] In order to achieve weighted fusion of features, two fully connected layers in the module calculate the weight of each feature. These weights reflect the importance of different features in the final decision, and the softmax function ensures that their sum is 1, thus avoiding conflicts between weights. Finally, these weights are multiplied with the feature map and summed in the channel dimension to obtain the fused feature representation. This fusion not only improves the model's ability to understand complex image structures, but also enhances its generalization performance on diverse datasets.

[0077] In the super-resolution reconstruction task, restoring high-frequency details in low-resolution images is inherently uncertain, resulting in the possibility that the same scene corresponds to multiple different high-resolution reconstruction versions. In the framework of generative adversarial networks, the generator's loss function plays a key role in measuring the difference between the generated image and the real image, which directly affects the visual quality of the high-resolution image.

[0078] We use the weighted sum of content loss, adversarial loss, and perceptual loss as the overall loss, taking into account both the image visual effect and the level of indicators to judge the quality of model improvement. Therefore, the LPIPS image similarity metric replaces the traditional MSE as the perceptual loss.

[0079] The present invention is further configured such that the calculation logic of the content loss is: Among them, L char is the content loss, x is the difference between the predicted value and the true value, and ε is a very small constant used to ensure the differentiability of the function when x is zero. Specifically, using Charbonnier instead of conventional MSE as the content loss has the following advantages: insensitivity to outliers, better smoothness, faster training speed, more in line with human eye perception, and the ability to balance visual effects and indicators. Charbonnier Loss is a loss function used for image generation tasks. It is an approximation of L1 loss and is used to improve the performance of the model.

[0080] The calculation logic of the generator loss is: in, is the generator loss, Represents the real image distribution X r The expected value of Indicates the probability that the real image sample is more real than the generated sample, and the value tends to 1. Represents the generated image distribution X f The expected value of Indicates the probability that the generated image is less realistic than the real image, and the value is closer to 0;

[0081] The calculation logic of the discriminator loss is: in, is the discriminator loss; specifically, in the training of GAN, the adversarial losses of the generator and the discriminator are always changing in balance. In the early training stage: the discriminator has strong performance and the image quality generated by the generator is low. At this time, the discriminator's loss is small (it can easily distinguish between true and false), and the generator's loss is large (the generated image is easily found to be fake by the discriminator); in the mid-term training stage: the generator gradually generates images that are closer to the real thing, and the discriminator begins to be "confused". At this time, the discriminator's loss increases (gradually loses its ability to distinguish), and the generator's loss decreases (the generator's forging ability improves); ideal training state: the discriminator and the generator reach a Nash equilibrium, and the discriminator cannot accurately judge the authenticity of the generated image (the output probability is close to 0.5). At this time, the discriminator's loss and the generator's loss are close to the optimal value, and the model performance is balanced. The value logic of the adversarial loss reflects the state and progress of the adversarial loss between the generator and the discriminator.

[0082] The calculation logic of the perception loss is: Among them, L adv is the perceptual loss, d(x,x0) is the distance function, which measures the difference between the input image x and the target image x0 in a specific feature space, ∑ l To sum the features of the lth layer in the network, H l and W l Represents the height and width of the feature map of the lth layer, which is used to normalize the feature difference of each layer, ∑ h,w To sum the spatial dimensions of the feature map, which include height h and width w, calculate the overall difference on the feature map of a specific layer, w l is the weight of the l-th layer feature, which is used to adjust the contribution of different layer features to the final loss value. ⊙ represents element-by-element multiplication, that is, element-by-element weighted operation. and is the feature value of the h,wth position in the normalized feature map of the input image x and the target image x0 at the lth layer, It is the square of the Euclidean norm, which is used to measure the difference between two sets of eigenvalues. Specifically, LPIPS-loss is used instead of the traditional MSE as the perceptual loss in the generator loss function. LPIPS is used as the perceptual loss function in the generator. By learning perceptual image features, it can accurately evaluate the visual difference between super-resolution images and real images, taking into account the differences in human eye perception, and has high interpretability. The VGG network features trained in the ImageNet classification task have shown significant results when used in the loss function of image synthesis training. This metric encourages the generator to master the ability to reconstruct real images from synthetic images by learning the inverse mapping between generated images and real images, and gives priority to the similarity between the two at the perceptual level. The calculation process Figure 7 As shown;

[0083] The calculation logic of comprehensive loss is: Among them, L G is the comprehensive loss, L per is the adversarial loss, α, and β are coefficients that balance the different loss terms.

[0084] The present invention is further configured to adopt a dual time scale and mixed precision training strategy, wherein the dual time scale and mixed precision training strategy includes dual time scale training and mixed precision training;

[0085] The dual-time-scale training uses two different learning rates to update the parameters of the network. The larger learning rate is used to update the low-level parameters in the network, quickly converge and learn low-level features; the smaller learning rate is used to update the high-level parameters in the network and accurately adjust the high-level features. By using dual-time-scale training, the network can perform more effective parameter updates at different levels, thereby accelerating the training process. In generative adversarial networks, a smaller learning rate is often set for the generator to prevent the optimal solution from being missed in the process of slowly approaching the optimal solution, and more iterations are required to optimize the generation process. Therefore, a smaller learning rate is set to help the generator learn more stably; however, for the discriminator, considering that its task is only a binary classification, a larger value is set to make it converge quickly, and fast convergence can provide more targeted feedback.

[0086] The mixed precision training balances the training speed and the training accuracy of the final model by utilizing the advantages of both half-precision and full-precision floating-point numbers, and effectively reduces memory consumption and computational overhead by using half-precision floating-point numbers for calculations during forward propagation and back-propagation. Since the limited representation range of half-precision floating-point numbers may lead to numerical overflow and precision loss, key operations are calculated using full-precision floating-point numbers to ensure the accuracy of the model. Key operations include parameter updates. The advantage of mixed precision training is that it can significantly reduce training time and resource consumption while maintaining model performance.

[0087] The currently widely used perceptual indicators (PSNR and SSIM) are very simple functions that cannot well explain human perception. In addition, the current GAN, especially VAE and other generative models, have too smooth results. The specific results are as follows: Figure 8 As shown, (a) the red box is the low-resolution image before super-resolution and the blue box is the detail contrast range (b) the generated result details are too smooth and the details are lost (c) the image after reconstruction under ideal conditions (d) the low-resolution image details. The VGG network features trained in the ImageNet classification task showed significant results when used as the loss function for image synthesis training. This metric encourages the generator to master the ability to reconstruct real images from synthetic images by learning the inverse mapping between generated images and real images, and prioritizes the similarity between the two at the perceptual level.

[0088] Feature maps are generated by activation functions after the convolution and pooling operations of the model. By visualizing these feature maps, we can understand the intensity and spatial distribution of the features. Blue indicates a low response value, green indicates a medium response value, yellow indicates a high response value, and red indicates a very high response value. The brighter the color, the higher the activity level of the pixel area, while the darker the color, the lower the activity level of the pixel area. Through the visualization analysis of layer feature activation, we can deeply study the internal operation mechanism of the model and find out the advantages and characteristics of the model in dealing with the super-resolution task of landslide remote sensing imagery. This analysis method can better understand the working principle of the model and provide guidance and inspiration for improving the model. A high response value usually means that the convolution kernel at that location is relatively sensitive to certain features in the input image and gives a higher activation response at the location where this feature exists.

[0089] In the three sets of feature visualization, the three columns correspond to the visualization comparison of the generator feature extraction capabilities of SRGAN, EDCA-SRGAN, and FFSRGAN (under the constraint of visual perception enhanced comprehensive loss function), such as Fig. 9 As shown, it can be seen intuitively that the improved residual block and feature fusion have a significant promoting effect on the performance improvement of the generator. From the three groups of comparisons, from left to right, it can be clearly seen that ① the landslide boundary gradually becomes clearer and more separable; ② and the response degree of the landslide area range increases significantly. According to the comparison results, it can be found that EDCA-SRGAN and FFSRGAN (the third and fourth columns) are very similar in the response degree of the landslide boundary outward, both showing blue-green, but there are obvious differences in the response degree of the two models in the landslide area. Among them, the response degree of FFSRGAN in the landslide area is almost all the highest level of red, which shows that the designed feature fusion module has a very significant effect on identifying landslide areas. The response degree of EDCA-SRGAN in the landslide area is relatively scattered, and it is not as obviously concentrated in the landslide area as FFSRGAN. Therefore, it can be considered that the additional feature fusion module added by FFSRGAN is more effective than the simple enhancement of feature extraction capability of EDCA-SRGAN, and can more accurately identify landslide areas; ③ The feature extraction modules before and after the improvement have different responses to landslide images under different vegetation coverage, which depends on the concealment of the landslide, and the feature extraction modules before and after the improvement have different responses to landslide images under different vegetation coverage. It can be concluded that the improved model has certain effectiveness in discovering hidden landslides, and has improved the difficulty of landslide identification - boundary information extraction, which proves that the module we designed is effective in feature extraction for the target of landslides.

[0090] Secondly, this experiment visualizes the feature graphs of FFSRGAN under different loss functions, as shown below: Fig.10As shown in Figure 2, a careful comparison shows that in the generative adversarial model, different loss functions still have a significant impact on the feature extraction ability of the model under the same network framework, such as Fig.10 The range of the two black boxes in (b), Fig.10 The response degree in the landslide area in (b) is significantly higher than that in (c) and (d), and the area outside the landslide has a lower response. Therefore, a suitable loss function can not only help improve the feature extraction ability of the landslide, but also suppress the interference of the background.

[0091] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0092] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0093] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0094] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0095] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0096] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0097] In the several embodiments provided in the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0098] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0100] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0101] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. Visual perception-guided multi-scale feature fusion landslide image super-resolution model, characterized by: Includes generator network, discriminator network and comprehensive loss function module; The generator network includes a head convolution module, a feature enhancement module, an upsampling reconstruction module and a branch network module; the head convolution module is used to extract high-frequency details in low-resolution landslide images, and a small-size convolution kernel is used to capture the boundary information of the landslide area; the feature enhancement module is used to extract landslide image features by combining deep and shallow layer features through a residual block embedded with an attention mechanism; the upsampling reconstruction module is used to realize the amplification of feature maps by alternating deconvolution layers and PA layers, and enhance important pixels, and the branch network module is used to perform feature fusion on the original input image and the image after interpolation and upsampling, and generate weight coefficients for weighted fusion; The discriminator network is used to adapt to the training requirements of the generative adversarial network through a multi-layer convolution structure to discriminate whether the input image is a real high-resolution image; The comprehensive loss function module is used to set the weighted sum of content loss, adversarial loss and perceptual loss as the overall loss; the content loss is used for image generation tasks to improve the model's sensitivity to landslide details; the adversarial loss includes generator loss and discriminator loss; the perceptual loss accurately evaluates the visual difference between super-resolution images and real images by learning the inverse mapping between generated images and real images.

2. The multi-scale feature fusion landslide image super-resolution model guided by visual perception according to claim 1 is characterized in that: The head convolution module includes a small-size convolution kernel of 3×3, which is used to capture high-frequency edge details of the image; The feature enhancement module is composed of 6 EDCA residual blocks and a convolution layer with a convolution kernel size of 3×3 in a stacked manner; The up-sampling reconstruction module is composed of multiple Up-Blocks and PA layers alternately, wherein the Up-Block is used for up-sampling, and the PA layer is used to enhance important pixels. Each Up-Block includes a deconvolution layer and two convolution layers, and the deconvolution layer is used to up-sample the input feature map; The branch network module consists of a convolutional layer, a leaky ReLU activation function and an adaptive average pooling layer, which is used to extract the features of the fused image. The adaptive average pooling layer is used to convert the feature map into a vector of length 64, and the vector is mapped to a vector of length 3 through a linear layer, and then normalized using a softmax layer. The result is the weight coefficient of the two images.

3. The visual perception-guided multi-scale feature fusion landslide image super-resolution model according to claim 1 is characterized in that: The discriminator network consists of a 64-channel convolutional layer with a stride of 1 and multiple repeated convolutional structures. The stride of the convolutional layer is 2. The length and width of the original input image become 1 / 16 of the original image. The last channel is 512-dimensional. After the output, it passes through an adaptive average pooling and LReLU function activation function, and then a convolutional layer with a dimension of 1 is used to obtain the output result.

4. The multi-scale feature fusion landslide image super-resolution model guided by visual perception according to claim 1 is characterized in that: It also includes a coordinate attention module, which captures spatial long-range correlation and channel global features by encoding along the horizontal and vertical coordinate directions, enhancing the feature extraction ability of the landslide area.

5. The multi-scale feature fusion landslide image super-resolution model guided by visual perception according to claim 1 is characterized in that: It also includes a multi-scale feature fusion module, which is implemented by the following steps: The input image is feature extracted through convolutional layers and upsampling layers, which can capture information of different scales from local to global. The extracted features are concatenated with the original image in the channel dimension, where concatenation in the channel dimension includes: using a fully connected layer to calculate the weights of features of different scales and normalizing them through a softmax function; performing weighted summation on the feature maps according to the weights to obtain a fused high-resolution feature representation.

6. The visual perception-guided multi-scale feature fusion landslide image super-resolution model according to claim 1 is characterized in that: The calculation logic of the content loss is: Among them, L char is the content loss, x is the difference between the predicted value and the true value, and ε is a small constant used to ensure the differentiability of the function when x is zero; The calculation logic of the generator loss is: in, is the generator loss, Represents the real image distribution X r The expected value of Indicates the probability that the real image sample is more real than the generated sample, and the value tends to 1. Represents the generated image distribution X f The expected value of Indicates the probability that the generated image is less realistic than the real image, and the value is closer to 0; The calculation logic of the discriminator loss is: in, is the discriminator loss; The calculation logic of the perception loss is: Among them, L adv is the perceptual loss, d ( x,x 0) is the distance function, which measures the difference between the input image x and the target image x0 in a specific feature space, ∑ l To sum the features of the lth layer in the network, H l and W l Represents the height and width of the feature map of the lth layer, which is used to normalize the feature difference of each layer, ∑ h,w To sum the spatial dimensions of the feature map, which include height h and width w, calculate the overall difference on the feature map of a specific layer, w l is the weight of the l-th layer feature, which is used to adjust the contribution of different layer features to the final loss value. ⊙ represents element-by-element multiplication, that is, element-by-element weighted operation. and is the feature value of the h,wth position in the normalized feature map of the input image x and the target image x0 at the lth layer, is the square of the Euclidean norm, which is used to measure the difference between two sets of eigenvalues; The calculation logic of comprehensive loss is: Among them, L G is the comprehensive loss, L per is the adversarial loss, α, and β are coefficients that balance the different loss terms.

7. The visual perception-guided multi-scale feature fusion landslide image super-resolution model according to claim 1 is characterized in that: Adopting a dual time scale and mixed precision training strategy, wherein the dual time scale and mixed precision training strategy includes dual time scale training and mixed precision training; The dual time scale training uses two different learning rates to update the parameters of the network. The larger learning rate is used to update the low-level parameters in the network, quickly converge and learn low-level features. A smaller learning rate is used to update high-level parameters in the network and accurately adjust high-level features; The mixed precision training balances the training speed and the training accuracy of the final model by utilizing the advantages of both half-precision and full-precision floating-point numbers, and effectively reduces memory consumption and computational overhead by using half-precision floating-point numbers for calculations during forward propagation and back-propagation. Since the limited representation range of half-precision floating-point numbers may lead to numerical overflow and precision loss, key operations are calculated using full-precision floating-point numbers to ensure the accuracy of the model.

Citation Information

Cited By

  • Synthetic aperture radar image super-resolution method and device based on deep learning

    CN121258796A

  • Synthetic aperture radar image super-resolution method and device based on deep learning

    CN121258796B

  • Climate data downscaling method and device, equipment and storage medium

    CN121705723A