A super-resolution remote sensing image generation method and system

Through the combination of depth-separable convolutional layer and attention mechanism, combined with multiple deep feature extraction and residual connection, the existing super-resolution remote sensing image generation method is solved, and efficient and natural image generation and optimization are achieved.

CN119991527BActive Publication Date: 2025-06-24YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465507.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing super-resolution remote sensing image generation method is difficult to effectively utilize the high-frequency information and context information of the image, resulting in insufficient details of the generated image, easy to cause blur distortion problems, and large model parameters, increased calculation overhead, and limited adaptability.

Method used

The depth-separable convolutional layer and attention mechanism are adopted, and the image quality evaluation index without reference is generated through multiple deep feature extraction and residual connections, combined with two upsampling reconstruction processes, mapping and fusion features, and super-resolution images are generated, and the reference-free image quality evaluation index NIQE is introduced to optimize image quality.

Benefits of technology

Improve the resolution and visual quality of remote sensing images, avoid artifacts and details loss, and the generated high-resolution images are more natural and realistic, reducing the computing resource requirements of the model, and are suitable for real-time processing on devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991527B_ABST
    Figure CN119991527B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image data processing, and specifically provides a super-resolution remote sensing image generation method and system. The specific steps of the method are as follows: obtaining an original high-resolution image from a dataset, downsampling it to obtain a low-resolution image, performing shallow feature extraction through a first convolution to obtain a first feature, obtaining deep features through multiple deep feature extractions, performing residual connection with the first feature to obtain a fifth feature, mapping the fifth feature, and using frequency domain fusion to obtain a super-resolution image; the super-resolution image and the original high-resolution image are respectively processed through a depthwise separable convolution layer to obtain a sixth feature, and negative values are eliminated through an activation function; after passing through a fully connected layer to unify the dimension of the seventh feature vector and mapping it to a scalar, a judgment result on whether the super-resolution image is the original high-resolution image is obtained. The present invention can improve the resolution and visual quality of remote sensing images while maintaining computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and specifically to a method and system for generating super-resolution remote sensing images. Background Art

[0002] High-resolution remote sensing images record rich geographical environment information and are widely used in various fields, such as small target detection, semantic segmentation, land classification, etc.

[0003] Inferring a high-resolution image from a low-resolution image is an inverse problem. Early image super-resolution methods mostly relied on interpolation or reconstruction strategies to obtain a high-resolution image by upsampling the low-resolution image. However, these methods cannot fully utilize the high-frequency information and context information of the image and are prone to problems such as blurring and distortion. With the development of deep learning technology, significant breakthroughs have been made in image super-resolution methods. Many studies have been conducted to improve the ill-posedness, such as dense networks, but due to their excessive focus on pixel-level errors, the generated images lack details. Especially when dealing with high-frequency textures and details, a blurring effect is easily produced.

[0004] The application of generative adversarial networks in image super-resolution faces the problem of a large number of model parameters, which significantly increases the computational overhead of the training and inference processes. In addition, most current methods have limited adaptability to training datasets and application scenarios and are difficult to maintain stable performance on different types of input data. Summary of the Invention

[0005] The present invention provides a method and system for generating super-resolution remote sensing images.

[0006] The technical solution of the present invention is as follows:

[0007] A method for generating super-resolution remote sensing images, the specific steps are as follows:

[0008] S1. Obtain the original high-resolution image from the dataset, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform deep feature extraction multiple times to obtain deep features, perform residual connection with the first feature to obtain the fifth feature, and perform mapping and fusion on the fifth feature through the second convolution and two upsampling reconstruction processes to obtain a super-resolution image;

[0009] Each deep feature extraction operation is as follows: The depthwise separable convolution layer includes pointwise convolution and depthwise convolution. The depthwise separable convolution layer is used to extract features from the first feature, and the activation function is used to eliminate negative values. The cyclic feature extraction and negative value elimination operations output the second feature. The second feature and the first feature are subjected to residual connection to obtain the third feature, which is used as the input of the next depthwise separable convolution layer. The above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature;

[0010] S2. For the super-resolution image, it is respectively processed by the depthwise separable convolution layer to obtain the sixth feature, and the activation function is used to eliminate negative values; The seventh feature is obtained through multiple optimization processes. Each optimization process sequentially passes through the depthwise separable convolution layer processing, the activation function layer processing, and batch normalization; The dimension of the seventh feature vector is unified through the fully connected layer and mapped to a scalar. It is judged whether the scalar is greater than the preset threshold. If it is greater, it is determined as the target super-resolution remote sensing image.

[0011] The processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature. The specific operation is as follows:

[0012] The input feature is divided into two parts. Among them, one part of the input feature is pre-normalized to obtain the query, key, and value. By calculating the similarity between the query and the key, the association between different positions in the feature map is evaluated, and the attention weight is normalized for weighting the value to generate the weighted output feature. The weighted output feature and the other part of the input feature are connected to obtain the fourth feature.

[0013] Through the second convolution and two upsampling reconstruction processes, the fifth feature is mapped and fused to obtain the super-resolution image. Specifically:

[0014] After the second convolution for mapping, the first upsampling reconstruction process is performed to obtain the first reconstructed image, and then the second upsampling reconstruction process is performed to obtain the second reconstructed image. The first reconstructed image and the second reconstructed image are fused to obtain the super-resolution image.

[0015] When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated through the no-reference image quality assessment metric.

[0016] The deep feature is obtained by performing three deep feature extractions on the first feature.

[0017] The above process of obtaining the third feature is repeated ten times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature.

[0018] A super-resolution remote sensing image generation system includes:

[0019] Generator; It obtains the original high-resolution image from the dataset, downsamples it to obtain a low-resolution image, performs shallow feature extraction by the first convolution to obtain the first feature, performs deep feature extraction multiple times to obtain deep features, performs residual connection with the first feature to obtain the fifth feature, and through the second convolution and two upsampling reconstruction processes, maps and fuses the fifth feature to obtain a super-resolution image;

[0020] Each deep feature extraction operation is as follows: The depthwise separable convolution layer includes pointwise convolution and depthwise convolution. The depthwise separable convolution layer is used to extract features from the first feature, and negative values are eliminated through the activation function; The cyclic feature extraction and negative value elimination operations output the second feature, and the second feature and the first feature are subjected to residual connection to obtain the third feature, which is used as the input of the next depthwise separable convolution layer. The above process of obtaining the third feature is repeated several times, and attention processing is performed on the processing result of the last depthwise separable convolution layer to obtain the fourth feature;

[0021] Discriminator; The super-resolution image is respectively processed by the depthwise separable convolution layer to obtain the sixth feature, and negative values are eliminated through the activation function; The seventh feature is obtained through multiple optimization processes. Each optimization process sequentially passes through the depthwise separable convolution layer processing, the activation function layer processing, and batch normalization; The seventh feature vector dimension is unified through the fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than a preset threshold. If it is greater, it is determined as the target super-resolution remote sensing image.

[0022] The training method is as follows: The super-resolution image and the original high-resolution image are subjected to standardization processing; Feature extraction is performed on the two standardized images; The distance between the same feature in the two images is calculated as the similarity; The loss function is calculated based on the similarity, and the loss function is fed back to the generator to optimize the generator parameters. The super-resolution image output again is transmitted to the discriminator to train the discrimination ability of the discriminator. When after multiple trainings, the fluctuation of the loss function is less than the preset value or reaches the preset number of training times, the training stops.

[0023] The operation of feeding back the loss function to the generator and optimizing the generator parameters is specifically as follows:

[0024] The parameters of the generator include the weights and biases of each layer in the generator network, and the generator adjusts the weights and biases of each layer through the feedback of the loss function.

[0025] The beneficial effects of the present invention are as follows:

[0026] The method proposed by the present invention is used for super-resolution reconstruction of remote sensing images, aiming to improve the resolution and visual quality of remote sensing images while maintaining computational efficiency.

[0027] This method can improve the spatial resolution while avoiding artifacts and detail loss commonly found in traditional super-resolution methods, and the generated high-resolution images are more natural and realistic.

[0028] To further optimize the image quality, the present invention introduces a no-reference image quality assessment metric to ensure that during the training process, not only pixel-level errors are concerned, but also the naturalness and visual perception of the image are improved, avoiding unnatural or distorted images.

[0029] In addition, the present invention uses a technique to prevent mode collapse: an activation function, to ensure that the generator learns rich and diverse image features during the training process, thereby improving the training stability and avoiding the phenomenon of degraded or single-faceted generated image quality. This makes the method show strong stability and generalization ability during the training process.

[0030] The method of the present invention reduces the computational resource requirements of the model, can operate efficiently on devices with limited resources, and is applicable to the real-time processing of drone and satellite remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In the drawings:

[0032] Figure 1 It is a schematic diagram of the reconstruction results of remote sensing images in the AID dataset. Figure 1 In it, (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is a partial enlarged image of the self-supervised degradation-guided adaptive blind reconstruction network, (g) is the image obtained by the large kernel convolutional super-resolution model, (h) is the image obtained by the diffusion probability model, (i) is the image obtained by using the method of the present invention, and (j) is a partial enlarged image of the original high-resolution image;

[0033] Figure 2 It is a schematic diagram of the reconstruction results of remote sensing images in the AID dataset. Figure 2 In it, (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is a partial enlarged image of the self-supervised degradation-guided adaptive blind reconstruction network, (g) is the image obtained by the large kernel convolutional super-resolution model, (h) is the image obtained by the diffusion probability model, (i) is the image obtained by using the method of the present invention, and (j) is a partial enlarged image of the original high-resolution image;

[0034] Figure 3Schematic diagram of the reconstruction results of remote sensing images in the WHU-RS19 dataset Figure 3 Among them, (a) is the original high-resolution image, (b) is the enlarged local image of the image obtained by bicubic interpolation, (c) is the enlarged local image of the image obtained by the generative adversarial network model, (d) is the enlarged local image of the image obtained by the enhanced generative adversarial network model, (e) is the enlarged local image of the image obtained by the blind super-resolution reconstruction model, (f) is the image obtained by the large kernel convolution super-resolution model, (g) is the image obtained by the diffusion probability model, (h) is the enlarged local image of the image obtained by using the remote sensing image reconstruction method of this application, and (i) is the enlarged local image of the original high-resolution image;

[0035] Figure 4 Schematic diagram of the reconstruction results of remote sensing images in the WHU-RS19 dataset Figure 4 Among them, (a) is the original high-resolution image, (b) is the enlarged local image of the image obtained by bicubic interpolation, (c) is the enlarged local image of the image obtained by the generative adversarial network model, (d) is the enlarged local image of the image obtained by the enhanced generative adversarial network model, (e) is the enlarged local image of the image obtained by the blind super-resolution reconstruction model, (f) is the image obtained by the large kernel convolution super-resolution model, (g) is the image obtained by the diffusion probability model, (h) is the enlarged local image of the image obtained by using the remote sensing image reconstruction method of this application, and (i) is the enlarged local image of the original high-resolution image. Detailed implementation manners

[0036] Example 1

[0037] The technical solution of the present invention is as follows:

[0038] A method for generating super-resolution remote sensing images, the specific steps are as follows:

[0039] S1. Obtain the original high-resolution image from the dataset, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform deep feature extraction multiple times to obtain deep features, perform residual connection with the first feature to obtain the fifth feature, and perform mapping and fusion on the fifth feature through the second convolution and two upsampling reconstruction processes to obtain a super-resolution image;

[0040] Each deep feature extraction operation is as follows: The depthwise separable convolution layer includes depthwise convolution and pointwise convolution. The depthwise separable convolution layer is used to extract features from the first feature, and the negative values are eliminated through the activation function. The cyclic feature extraction and negative value elimination operations output the second feature. The second feature and the first feature are subjected to residual connection to obtain the third feature, which serves as the input to the next depthwise separable convolution layer. The above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature.

[0041] The attention processing is specifically as follows:

[0042] The input feature is divided into two parts. Among them, one part of the input feature undergoes pre-normalization to obtain the query, key, and value. By calculating the similarity between the query and the key, the association between different positions in the feature map is evaluated, and the attention weights are normalized for weighting the values to generate the weighted output feature. The weighted output feature and the other part of the input feature are connected to obtain the fourth feature.

[0043] The residual connection strategy effectively optimizes the feature extraction process, and at the same time, the integration of the attention mechanism and frequency domain features improves the efficiency and effect of super-resolution reconstruction.

[0044] Therefore, step S1 can be summarized into three sub-steps: shallow feature extraction, deep feature extraction, and fusion reconstruction.

[0045] In the experiment of the present invention, shallow feature extraction is realized through 3×3 convolution, and then through three times of deep feature extraction, the deep feature is obtained;

[0046] The mapping relationship of deep feature extraction is as follows:

[0047] ,

[0048] where, is the deep feature, n is the number of times of deep feature extraction, is the first feature, is the deep feature extraction.

[0049] The depthwise separable convolution layer and the activation function cooperate to perform deep feature extraction. The depthwise separable convolution includes depthwise convolution and pointwise convolution. While reducing the computational amount, it can maintain the expressive ability of the model. The activation function alleviates the gradient vanishing problem by allowing a small part of the negative slope for negative values to pass through, and improves the information transmission ability of the model during the training process.

[0050] In each process of deep feature extraction, the above process of obtaining the third feature is repeated ten times, and the processing result of the last depthwise separable convolutional layer is subjected to attention processing to obtain the fourth feature, and then the next deep feature extraction is performed until the deep feature is obtained through the third deep feature extraction.

[0051] The present invention performs deep feature extraction multiple times, aiming to effectively extract low-level and high-level features. Aiming to effectively extract low-level and high-level features.

[0052] In the real remote sensing image super-resolution reconstruction task, as the network depth increases, the problem of gradient disappearance or explosion often makes the training process unstable. In addition, traditional methods have limitations in the transmission of context information, and it is difficult for low-level features to be effectively transmitted to high levels, resulting in information loss. Especially in complex tasks, the detailed features of the input image cannot be fully captured, reducing the effect and accuracy of super-resolution reconstruction.

[0053] To solve this problem, residual connections are introduced in deep feature extraction, and the feature map after multiple processes is added to the first feature. Different from traditional feature addition, residual connections can effectively alleviate the problem of gradient disappearance in deep feature extraction and ensure the smooth propagation of information. This mechanism not only improves the optimization and convergence efficiency, but also can more accurately capture the subtle relationship between the input and output, thereby improving performance and accuracy. At the same time, residual connections help to reduce overfitting, reduce the over-reliance on complex features, and enhance the generalization ability of deep feature extraction.

[0054] Next, through two upsampling reconstruction processes, the high-level deep feature map is restored to a high-resolution image. This step uses the nearest neighbor upsampling operation to gradually restore the spatial resolution of the image. The nearest neighbor upsampling enhances the resolution while maintaining the image structure by simply copying adjacent pixel values, reducing the computational overhead. In this way, the high-frequency details of the image can be more accurately restored, gradually approaching the true structure of the original high-resolution image.

[0055] After each upsampling and convolution process, the spatial resolution of the feature map continuously increases. When the network gradually restores the details and textures of the image, it can effectively avoid the problem of image blurring or distortion. The mathematical expressions of the two upsamplings are as follows:

[0056] ,

[0057] ,

[0058] Among them, and represent the results of the upsampling operation, is the upsampling operation, and the multiple is .

[0059] Finally, the results obtained by the two upsampling modules are subjected to frequency-domain feature fusion, specifically as follows:

[0060] After the second convolution is mapped, the first upsampling reconstruction process is performed to obtain the first reconstructed image, and then the second upsampling reconstruction process is performed to obtain the second reconstructed image. The first reconstructed image and the second reconstructed image are fused to obtain the super-resolution image.

[0061] The reconstructed images obtained by the two upsampling processes are subjected to frequency-domain feature fusion. By fusing the high-level features of the low-resolution and the high-resolution features, while maintaining the overall image structure, the transmission of high-frequency information is effectively enhanced. This process helps to refine the detailed performance of the image, especially playing an important role in texture restoration and edge enhancement.

[0062] S2. The super-resolution image is respectively processed through a depthwise separable convolutional layer to obtain the sixth feature, and the negative values are eliminated through an activation function; the seventh feature is obtained through multiple optimization processes. Each optimization process sequentially passes through a depthwise separable convolutional layer process, an activation function layer process, and batch normalization; the dimension of the seventh feature vector is unified through a fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than a preset threshold. If it is greater, it is determined as the target super-resolution remote sensing image.

[0063] The main advantage is that the depthwise separable convolution effectively reduces the computational complexity of the model while maintaining a high feature extraction ability. Extract the detailed information of the image from different levels and different scales, thereby improving the discrimination accuracy and the sensitivity to subtle differences.

[0064] After each depthwise separable convolution, batch normalization and an activation function are introduced, aiming to improve the training stability of the network, accelerate the convergence process, and strengthen the expression ability of non-linear features. It helps to slow down the problem of gradient disappearance, and the activation function avoids the phenomenon of "dead neurons", enabling the network to capture image features more effectively.

[0065] The high-level features extracted are mapped to a scalar through a fully connected layer, representing the probability that the input image is a real image. The present invention discriminates whether the super-resolution image is the original high-resolution image by outputting 0 and 1. This network discriminates in a multi-scale manner, extracts the detailed information of the image from different levels and different scales, thereby improving the discrimination accuracy and the sensitivity to subtle differences.

[0066] When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated through a no-reference image quality assessment metric.

[0067] In the process of super-resolution image generation, in order to more accurately evaluate the visual quality of the super-resolution image output by the generator, the present invention introduces the NIQE (Natural Image Quality Evaluator) metric to evaluate the visual quality of the generated image. NIQE is a reference-free image quality assessment metric that can measure the quality of an image by analyzing the natural statistical characteristics of the image without relying on the original image or a reference image. Different from traditional image quality assessment methods, NIQE pays more attention to the visual perception effect of the human eye and can more accurately reflect the naturalness and realism of the image.

[0068] A super-resolution remote sensing image generation system, comprising:

[0069] A generator; obtaining an original high-resolution image from a dataset, downsampling it to obtain a low-resolution image, performing shallow feature extraction by a first convolution to obtain a first feature, performing multiple deep feature extractions to obtain deep features, performing a residual connection with the first feature to obtain a fifth feature, and performing mapping and fusion on the fifth feature through a second convolution and two upsampling reconstructions to obtain a super-resolution image;

[0070] Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes a depthwise convolution and a pointwise convolution, using the depthwise separable convolution layer to perform feature extraction on the first feature, and eliminating negative values through an activation function; looping the feature extraction and negative value elimination operations to output a second feature, performing a residual connection between the second feature and the first feature to obtain a third feature, which serves as the input to the next depthwise separable convolution layer, repeating the above process of obtaining the third feature several times, and performing an attention process on the processing result of the last depthwise separable convolution layer to obtain a fourth feature;

[0071] A discriminator; for the super-resolution image, respectively performing processing through a depthwise separable convolution layer to obtain a sixth feature, and eliminating negative values through an activation function; obtaining a seventh feature through multiple optimization processes, and each optimization process sequentially includes processing through a depthwise separable convolution layer, an activation function layer, and batch normalization; unifying the dimension of the seventh feature vector through a fully connected layer and mapping it to a scalar, and determining whether the scalar is greater than a preset threshold. If it is greater, it is determined to be the target super-resolution remote sensing image.

[0072] The training method for the generator and the discriminator is to perform normalization processing on the super-resolution image and the original high-resolution image; perform feature extraction on the two normalized images; calculate the distance between the same feature in the two images as the similarity; calculate a loss function based on the similarity, feedback the loss function to the generator, optimize the parameters of the generator, and transmit the super-resolution image output again to the discriminator to train the discrimination ability of the discriminator. When, after multiple trainings, the fluctuation of the loss function is less than a preset value or reaches the preset number of training times, stop training.

[0073] During the training process, a two-stage training strategy is adopted to ensure that the generator can produce high-quality remote sensing images. First, the generator is trained independently. The loss function optimizes the features of the generated images. By introducing the CLIP model to calculate the semantic similarity of the images, the loss function evaluates the generated images at the semantic level, calculates the similarity between the images and the high-level semantic features in the real scene, and helps generate high-quality images that are more in line with human perception. The cross-modal learning ability of the CLIP model between images and texts enables the generator to not only focus on the low-level features of the images (such as pixel-level similarity), but rather pay more attention to the high-level semantic content of the images, thus generating images that conform to human perception.

[0074] The input super-resolution image and the original high-resolution image are uniformly resized and normalized. Then, the CLIP model is selected to extract features from the normalized images. Finally, the distance between the extracted image features is calculated to measure the similarity between the images. The calculation method of the loss function is as follows:

[0075] ,

[0076] ,

[0077] ,

[0078] where the super-resolution image I and the original high-resolution image GT are encoded into corresponding feature vectors and , the image size after normalization, is the redefined image size. is the dimension of the feature vector, and respectively represent the th elements in the feature vectors of I and GT, is the loss function.

[0079] The total loss function is defined as:

[0080] ,

[0081] where represents the pixel loss, is the perceptual loss, is the adversarial loss. are the proportions of the three loss functions respectively.

[0082] The loss function of the present invention focuses on high-level semantics and structure, overcomes the limitations in the calculation of traditional loss functions, and avoids over-optimizing pixel-level details, which may cause the image to appear unnatural visually or lack semantic consistency.

[0083] The generator optimizes its own loss function to learn how to generate as realistic images as possible without being interfered by the discriminator's feedback. The specific operation is as follows:

[0084] The parameters of the generator include the weights and biases of each layer in the generator network. The generator adjusts the weights and biases of each layer through the feedback of the loss function.

[0085] This enables the generator to first learn to generate effective image features in the initial stage of training.

[0086] Then, the generator and the discriminator are alternately trained to further improve the image quality. The generator "deceives" the discriminator by generating more realistic images, making it unable to distinguish between real images and generated images. The discriminator, on the other hand, improves its ability to distinguish real and generated images. By enhancing the discriminator's ability, it promotes the generator to generate more natural and detailed images. The interaction between the two promotes each other. The mutual game between the generator and the discriminator during the training process continuously improves the quality of the generated images and effectively avoids the mode collapse problem, ensuring that the generated images can maintain high quality in multiple aspects (such as structure, texture, semantic consistency), ultimately improving the image quality and avoiding the mode collapse problem.

[0087] The design of the activation function layer processing and batch normalization aims to improve the training stability of the network, accelerate the convergence process, and enhance the expression ability of non-linear features.

[0088] During the optimization process, the generator not only needs to focus on pixel-level losses (such as L1, L2 losses), but also needs to consider the quality of the generated images in terms of natural statistical characteristics.

[0089] In the generative adversarial training, the generator is optimized according to the discriminator's feedback, continuously improving the quality of the generated images and gradually approaching the details of real images. The discriminator, by comparing the generated images with real images, provides accurate discrimination results to help the generator improve its generation strategy.

[0090] Through this adversarial training process, the discriminator not only promotes the generator to generate higher-quality images, but also improves the final effect of the super-resolution reconstruction task. The introduction of deep feature extraction enables the super-resolution generative network to more accurately restore the details of high-resolution images.

[0091] The NIQE (Natural Image Quality Evaluator) metric is introduced to evaluate the visual quality of the generated images. The introduction of this metric makes the optimization process more in line with the needs of human visual perception, avoiding the problem of visual distortion or unnaturalness of images caused by over-reliance on pixel-level losses. NIQE is a reference-free image quality assessment metric that can measure the quality of an image by analyzing its natural statistical properties without relying on the original image or a reference image. Different from traditional image quality assessment methods (such as PSNR, SSIM), NIQE pays more attention to the visual perception effect of the human eye and can more accurately reflect the naturalness and realism of the image. By monitoring the change of the NIQE metric of the model on the validation set, the similarity between the generated super-resolution image and the original high-resolution image is regularly evaluated to ensure that the generator can produce high-quality images closer to real images.

[0092] The remote sensing images in the AID dataset are reconstructed using the remote sensing image reconstruction method of the present invention. The reconstructed low-resolution images are obtained by bicubic downsampling, and the reconstruction results are respectively as Figure 1 , Figure 2 shown.

[0093] Among them, Figure 1 and Figure 2 in (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is the local enlarged image of the self-supervised degradation-guided adaptive blind reconstruction network, (g) is the image obtained by the large kernel convolutional super-resolution model, (h) is the image obtained by the diffusion probability model, (i) is the image obtained by using the remote sensing image reconstruction method of this application, and (j) is the local enlarged image of the original high-resolution image.

[0094] In addition, the WHU-RS19 dataset is also reconstructed using the remote sensing image reconstruction method of this application. The low-resolution images required for reconstruction are obtained by processing with a Gaussian degradation kernel of size 0.6, as Figure 3 , Figure 4As shown. Among them, (a) is the original high-resolution image, (b) is the local enlarged image of the image obtained by bicubic interpolation, (c) is the local enlarged image of the image obtained by the generative adversarial network model, (d) is the local enlarged image of the image obtained by the enhanced generative adversarial network model, (e) is the local enlarged image of the image obtained by the blind super-resolution reconstruction model, (f) is the image obtained by the large kernel convolutional super-resolution model, (g) is the image obtained by the diffusion probability model, (h) is the local enlarged image of the image obtained by using the remote sensing image reconstruction method of the present application, and (i) is the local enlarged image of the original high-resolution image.

[0095] Among the above, the image obtained by bicubic interpolation, compared with the present invention that does not use a lightweight attention mechanism and depthwise separable convolution, is only based on mathematical interpolation and cannot recover real high-frequency information.

[0096] The image obtained by the generative adversarial network model (compared with the present invention that does not use frequency domain feature fusion) learns the mapping relationship from low resolution to high resolution through the generative adversarial network model. Because it does not use frequency domain feature fusion, it will cause some image distortion or blurring.

[0097] The image obtained by the enhanced generative adversarial network model, compared with the present invention that does not use depthwise separable convolution, introduces a reconstruction residual block and a perceptual loss function, resulting in a high computational complexity and being difficult to be deployed to edge devices.

[0098] The image obtained by the blind super-resolution reconstruction model, compared with the method of the present invention, does not use feature extraction of multiple scales, but proposes a more complex but practical image degradation model, resulting in being overly dependent on the predefined degradation model and having limited generalization.

[0099] The image obtained by the self-supervised degradation-guided adaptive blind reconstruction network, compared with the present invention, requires a large amount of unlabeled data for pre-training; the present invention only requires single-image training, and the lightweight design can achieve efficient adaptation.

[0100] The image obtained by the large kernel convolutional super-resolution model, compared with the method of the present invention, does not use single-head attention. Through large convolutional and channel split-shuffle operations, not adding the single-head attention mechanism will result in a large computational overhead and high memory occupancy.

[0101] The image obtained by the diffusion probability model, compared with the one-way inference of the present invention, this method iterates multiple times, with too many iteration times and slow inference speed.

[0102] By Figure 1 , Figure 2 , Figure 3 , Figure 4It can be seen that, compared with other methods, the remote sensing image reconstruction method of the present application can more effectively simulate the degradation of real images and adaptively fuse information of different resolutions during the reconstruction process, thereby improving the clarity and accuracy of the reconstructed images.

Claims

1. A method for generating a super-resolution remote sensing image, characterized in that: The specific steps are: S1. Obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; S2, super-resolution image, is processed by a depthwise separable convolution layer to obtain the sixth feature, and negative values ​​are eliminated by an activation function; the seventh feature is obtained through multiple optimization processes, each of which is processed by a depthwise separable convolution layer, an activation function layer, and batch normalization in sequence; the seventh feature vector dimension is unified through a fully connected layer and mapped to a scalar, and it is determined whether the scalar is greater than a preset threshold. If so, it is determined to be a target super-resolution remote sensing image.

2. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The processing result of the last depth-separable convolutional layer is subjected to attention processing to obtain the fourth feature. The specific operation is: The input features are divided into two parts, where one part of the input features is pre-normalized to obtain the query, key and value. The similarity between the query and the key is calculated, the association between different positions in the feature map is evaluated, the attention weights are normalized to weight the values, and the weighted output features are generated. The weighted output features are connected with the other part of the input features to obtain the fourth feature.

3. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The fifth feature is mapped and fused to obtain a super-resolution image through the second convolution and two upsampling reconstruction processes, specifically: After the second convolution is mapped, a first upsampling reconstruction process is performed to obtain a first reconstructed image, and then a second upsampling reconstruction process is performed to obtain a second reconstructed image, and the first reconstructed image and the second reconstructed image are fused to obtain a super-resolution image.

4. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated by a no-reference image quality assessment metric.

5. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The deep feature is obtained by extracting the first feature three times.

6. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: Repeat the above process of obtaining the third feature ten times, perform attention processing on the processing result of the last depth-separable convolutional layer, and obtain the fourth feature.

7. A super-resolution remote sensing image generation system, characterized in that: include: Generator; The original high-resolution image is obtained from the data set, and the low-resolution image is obtained by downsampling. The first convolution is used to extract shallow features to obtain the first feature. After multiple deep feature extractions, the deep feature is obtained. The fifth feature is obtained by residual connection with the first feature. The fifth feature is mapped and fused through the second convolution and two upsampling reconstruction processes to obtain a super-resolution image. Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; Discriminator; the super-resolution image is processed by the depthwise separable convolution layer to obtain the sixth feature, and the negative value is eliminated by the activation function; the seventh feature is obtained through multiple optimization processes, and each optimization process is sequentially processed by the depthwise separable convolution layer, the activation function layer and batch normalization; the seventh feature vector dimension is unified through the fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than the preset threshold. If it is greater, it is judged as the target super-resolution remote sensing image.

8. A super-resolution remote sensing image generation system according to claim 7, characterized in that: The training method is to standardize the super-resolution image and the original high-resolution image; extract features from the two standardized images; calculate the distance between the same feature in the two images as the similarity; calculate the loss function based on the similarity, feed the loss function back to the generator, optimize the generator parameters, and transmit the super-resolution image output again to the discriminator to train the discriminator's discrimination ability. When the fluctuation of the loss function is less than the preset value or reaches the preset number of training times after multiple trainings, the training is stopped.

9. A super-resolution remote sensing image generation system according to claim 8, characterized in that: The loss function is fed back to the generator to optimize the generator parameters. The specific operation is: The parameters of the generator include the weights and biases of each layer in the generator network. The generator adjusts the weights and biases of each layer through feedback from the loss function.

Citation Information

Patent Citations

  • Super-resolution reconstruction method based on attention mechanism

    CN113706386A

  • Remote sensing image super-resolution reconstruction method and system based on deep and shallow feature fusion

    CN117830100A