Super-resolution remote sensing image generation method and system
Through the combination of depth-separable convolutional layer and attention mechanism, combined with multiple deep feature extraction and residual connection, the problems of blur distortion and computational overhead in the existing super-resolution remote sensing image generation methods are solved, and high-quality and efficient super-resolution image generation is achieved.
Patent Information
- Application Number
- CN202510465507.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing super-resolution remote sensing image generation method is difficult to make full use of the high-frequency information and context information of the image, and is prone to fuzzy distortion problems, and the model parameters are large, the calculation overhead is increased, and the adaptability is limited.
The depth-separable convolutional layer and attention mechanism are adopted, and multiple deep feature extraction and residual connections are combined with two upsampling reconstruction processes to generate super-resolution images, and reference-free image quality evaluation indicators such as NIQE are introduced to optimize image quality.
It improves the resolution and visual quality of remote sensing images, avoids artifacts and details loss, and the generated images are more natural and realistic, reducing the computing resource requirements, and is suitable for devices with limited resources.
Smart Images

Figure CN119991527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a method and system for generating a super-resolution remote sensing image. Background Art
[0002] High-resolution remote sensing images record rich geographic environment information and are widely used in various fields, such as small target detection, semantic segmentation, land classification, etc.
[0003] Inferring a high-resolution image from a low-resolution image is an inverse problem. Early image super-resolution methods mostly relied on interpolation or reconstruction strategies to obtain high-resolution images by upsampling low-resolution images. However, these methods cannot fully utilize the high-frequency information and contextual information of the image, and are prone to blurring and distortion. With the development of deep learning technology, image super-resolution methods have made significant breakthroughs. Many studies have been made to improve the ill-posedness, such as dense networks, but because they focus too much on pixel-level errors, the details of the generated images are insufficient. Especially when dealing with high-frequency textures and details, blurriness is prone to occur.
[0004] The application of generative adversarial networks in image super-resolution faces the problem of large number of model parameters, which significantly increases the computational overhead of training and inference. In addition, most current methods have limited adaptability to training datasets and application scenarios, and it is difficult to maintain stable performance on different types of input data. Summary of the invention
[0005] The present invention provides a method and system for generating a super-resolution remote sensing image.
[0006] The technical solution of the present invention is as follows: A super-resolution remote sensing image generation method, the specific steps are: S1. Obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; S2, super-resolution image, is processed by a depthwise separable convolution layer to obtain the sixth feature, and negative values are eliminated by an activation function; the seventh feature is obtained through multiple optimization processes, each of which is processed by a depthwise separable convolution layer, an activation function layer, and batch normalization in sequence; the seventh feature vector dimension is unified through a fully connected layer and mapped to a scalar, and it is determined whether the scalar is greater than a preset threshold. If so, it is determined to be a target super-resolution remote sensing image.
[0007] Perform attention processing on the processing result of the last depth-separable convolutional layer to obtain the fourth feature. The specific operation is: The input features are divided into two parts, where one part of the input features is pre-normalized to obtain the query, key and value. The similarity between the query and the key is calculated, the association between different positions in the feature map is evaluated, the attention weights are normalized to weight the values, and the weighted output features are generated. The weighted output features are connected with the other part of the input features to obtain the fourth feature.
[0008] Through the second convolution and two upsampling reconstruction processes, the fifth feature is mapped and fused to obtain a super-resolution image, specifically: After the second convolution is mapped, a first upsampling reconstruction process is performed to obtain a first reconstructed image, and then a second upsampling reconstruction process is performed to obtain a second reconstructed image, and the first reconstructed image and the second reconstructed image are fused to obtain a super-resolution image.
[0009] When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated by a no-reference image quality assessment metric.
[0010] The deep feature is obtained by extracting the first feature three times.
[0011] Repeat the above process of obtaining the third feature ten times, perform attention processing on the processing result of the last depth-separable convolutional layer, and obtain the fourth feature.
[0012] A super-resolution remote sensing image generation system, comprising: Generator: obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; Discriminator; the super-resolution image is processed by the depthwise separable convolution layer to obtain the sixth feature, and the negative value is eliminated by the activation function; the seventh feature is obtained through multiple optimization processes, and each optimization process is sequentially processed by the depthwise separable convolution layer, the activation function layer and batch normalization; the seventh feature vector dimension is unified through the fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than the preset threshold. If it is greater, it is judged as the target super-resolution remote sensing image.
[0013] The training method is to standardize the super-resolution image and the original high-resolution image; extract features from the two standardized images; calculate the distance between the same feature in the two images as the similarity; calculate the loss function based on the similarity, feed the loss function back to the generator, optimize the generator parameters, and transmit the super-resolution image output again to the discriminator to train the discriminator's discrimination ability. When the fluctuation of the loss function is less than the preset value or reaches the preset number of training times after multiple trainings, the training is stopped.
[0014] The loss function is fed back to the generator to optimize the generator parameters. The specific operation is: The parameters of the generator include the weights and biases of each layer in the generator network. The generator adjusts the weights and biases of each layer through feedback from the loss function.
[0015] The beneficial effects of the present invention are: The method proposed in the present invention is used for super-resolution reconstruction of remote sensing images, aiming to improve the resolution and visual quality of remote sensing images while maintaining computational efficiency.
[0016] This method can improve the spatial resolution while avoiding the artifacts and detail loss common in traditional super-resolution methods, and the generated high-resolution images are more natural and realistic.
[0017] In order to further optimize the image quality, the present invention introduces a no-reference image quality evaluation index to ensure that during the training process, not only the pixel-level error is focused on, but also the naturalness and visual perception of the image are improved to avoid unnatural or distorted images.
[0018] In addition, the present invention uses a technology to prevent mode collapse: activation function, which ensures that the generator learns rich and diverse image features during the training process, thereby improving the training stability and avoiding the phenomenon of quality degradation or uniformity of generated images. This makes the method show strong stability and generalization ability during the training process.
[0019] The method of the present invention reduces the computing resource requirements of the model, can run efficiently on devices with limited resources, and is suitable for real-time processing of unmanned aerial vehicles and satellite remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In the attached picture: Figure 1 This is a schematic diagram of the reconstruction results of the remote sensing image in the AID dataset. Figure 1 (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is the locally enlarged image of the adaptive blind reconstruction network guided by self-supervised degradation, (g) is the image obtained by the large kernel convolution super-resolution model, (h) is the image obtained based on the diffusion probability model, (i) is the image obtained by the method of the present invention, and (j) is the locally enlarged image of the original high-resolution image; Figure 2 This is a schematic diagram of the reconstruction results of the remote sensing image in the AID dataset. Figure 2 (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is the locally enlarged image of the adaptive blind reconstruction network guided by self-supervised degradation, (g) is the image obtained by the large kernel convolution super-resolution model, (h) is the image obtained based on the diffusion probability model, (i) is the image obtained by the method of the present invention, and (j) is the locally enlarged image of the original high-resolution image; Figure 3 This is a schematic diagram of the reconstruction results of the remote sensing images in the WHU-RS19 dataset. Figure 3(a) is the original high-resolution image, (b) is a locally enlarged image of the image obtained by bicubic interpolation, (c) is a locally enlarged image of the image obtained by the generative adversarial network model, (d) is a locally enlarged image of the image obtained by the enhanced generative adversarial network model, (e) is a locally enlarged image of the image obtained by the blind super-resolution reconstruction model, (f) is an image obtained by the large kernel convolution super-resolution model, (g) is an image obtained based on the diffusion probability model, (h) is a locally enlarged image of the image obtained by the remote sensing image reconstruction method of the present application, and (i) is a locally enlarged image of the original high-resolution image; Figure 4 This is a schematic diagram of the reconstruction results of the remote sensing images in the WHU-RS19 dataset. Figure 4 Among them, (a) is the original high-resolution image, (b) is the locally enlarged image of the image obtained by bicubic interpolation, (c) is the locally enlarged image of the image obtained by the generative adversarial network model, (d) is the locally enlarged image of the image obtained by the enhanced generative adversarial network model, (e) is the locally enlarged image of the image obtained by the blind super-resolution reconstruction model, (f) is the image obtained by the large kernel convolution super-resolution model, (g) is the image obtained based on the diffusion probability model, (h) is the locally enlarged image of the image obtained by the remote sensing image reconstruction method of the present application, and (i) is the locally enlarged image of the original high-resolution image. DETAILED DESCRIPTION
[0021] Example 1 The technical solution of the present invention is as follows: A super-resolution remote sensing image generation method, the specific steps are: S1. Obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, and the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, and the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer. The above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature.
[0022] Attention processing, the specific operations are: The input features are divided into two parts, where one part of the input features is pre-normalized to obtain the query, key and value. The similarity between the query and the key is calculated, the association between different positions in the feature map is evaluated, the attention weights are normalized to weight the values, and the weighted output features are generated. The weighted output features are connected with the other part of the input features to obtain the fourth feature.
[0023] The residual connection strategy effectively optimizes the feature extraction process, and at the same time integrates the attention mechanism and frequency domain features to improve the efficiency and effect of super-resolution reconstruction.
[0024] Therefore, step S1 can be summarized into three sub-steps: shallow feature extraction, deep feature extraction and fusion reconstruction.
[0025] In the experiment of the present invention, shallow feature extraction is achieved through 3×3 convolution, and then deep feature extraction is performed three times to obtain deep features; The mapping relationship of deep feature extraction is as follows: , in, is the deep feature, n is the number of times the deep feature is extracted, The first feature is For deep feature extraction.
[0026] The deep separable convolution layer and activation function work together to extract deep features. The deep separable convolution includes channel-by-channel convolution and point-by-point convolution. It reduces the amount of calculation while maintaining the expressiveness of the model. The activation function alleviates the gradient vanishing problem by allowing negative values to pass through a small part of the negative slope, improving the model's ability to transmit information during training.
[0027] In each deep feature extraction process, the above process of obtaining the third feature is repeated ten times, and the processing result of the last depth-separable convolutional layer is subjected to attention processing to obtain the fourth feature, and then the next deep feature extraction is performed until the third deep feature extraction obtains the deep feature.
[0028] The present invention performs deep features multiple times, aiming to effectively extract low-level and high-level features. Aims to effectively extract low-level and high-level features.
[0029] In the task of super-resolution reconstruction of real remote sensing images, as the network depth increases, the gradient vanishing or exploding problem often makes the training process unstable. In addition, traditional methods have limitations in the transmission of contextual information, and low-level features are difficult to effectively transmit to high-level features, resulting in information loss. Especially in complex tasks, it is impossible to fully capture the detailed features of the input image, reducing the effect and accuracy of super-resolution reconstruction.
[0030] To solve this problem, residual connections are introduced in deep feature extraction to add the feature map after multiple processing to the first feature. Unlike traditional feature addition, residual connections can effectively alleviate the gradient vanishing problem in deep feature extraction and ensure that information can be transmitted smoothly. This mechanism not only improves the optimization and convergence efficiency, but also more accurately captures the subtle relationship between input and output, thereby improving performance and accuracy. At the same time, residual connections help alleviate overfitting, reduce excessive reliance on complex features, and enhance the generalization ability of deep feature extraction.
[0031] Next, two upsampling reconstruction processes are used to restore the high-level deep feature map to a high-resolution image. This step uses the nearest neighbor upsampling operation to gradually restore the spatial resolution of the image. The nearest neighbor upsampling simply copies the values of neighboring pixels to increase the resolution while maintaining the image structure and reduce the computational overhead. In this way, the high-frequency details of the image can be restored more accurately, gradually approaching the true structure of the original high-resolution image.
[0032] After each upsampling and convolution process, the spatial resolution of the feature map continues to increase, and the network can effectively avoid image blur or distortion when gradually restoring the details and texture of the image. The mathematical expression of two upsampling is as follows: , , in, and represents the result of the upsampling operation, is an upsampling operation, with a multiple of .
[0033] Finally, the results obtained by the two upsampling modules are fused in the frequency domain, specifically: After the second convolution is mapped, a first upsampling reconstruction process is performed to obtain a first reconstructed image, and then a second upsampling reconstruction process is performed to obtain a second reconstructed image, and the first reconstructed image and the second reconstructed image are fused to obtain a super-resolution image.
[0034] The frequency domain feature fusion of the two reconstructed images obtained by upsampling is carried out. By fusing the low-resolution high-level features and high-resolution features, the transmission of high-frequency information is effectively enhanced while maintaining the overall image structure. This process helps to refine the details of the image, especially playing an important role in texture recovery and edge enhancement.
[0035] S2, super-resolution image, is processed by a depthwise separable convolution layer to obtain the sixth feature, and negative values are eliminated by an activation function; the seventh feature is obtained through multiple optimization processes, each of which is processed by a depthwise separable convolution layer, an activation function layer, and batch normalization in sequence; the seventh feature vector dimension is unified through a fully connected layer and mapped to a scalar, and it is determined whether the scalar is greater than a preset threshold. If so, it is determined to be a target super-resolution remote sensing image.
[0036] The main advantage is that the depthwise separable convolution effectively reduces the computational complexity of the model while maintaining a high feature extraction capability. It extracts detailed information of the image from different levels and scales, thereby improving the discrimination accuracy and sensitivity to subtle differences.
[0037] After each depth-wise separable convolution, batch normalization and activation functions are introduced to improve the training stability of the network, accelerate the convergence process, and enhance the expression of nonlinear features. It helps to alleviate the gradient vanishing problem, while the activation function avoids the "dead neuron" phenomenon, allowing the network to capture image features more effectively.
[0038] The extracted high-level features are mapped to a scalar through the fully connected layer, indicating the probability that the input image is a real image. The present invention determines whether the super-resolution image is the original high-resolution image by outputting 0 and 1. The network extracts detailed information of the image from different levels and scales through a multi-scale discrimination method, thereby improving the discrimination accuracy and sensitivity to subtle differences.
[0039] When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated by a no-reference image quality assessment metric.
[0040] In the process of super-resolution image generation, in order to more accurately evaluate the visual quality of the super-resolution image output by the generator, the present invention introduces the NIQE (Natural Image Quality Evaluator) indicator to evaluate the visual quality of the generated image. NIQE is a reference-free image quality evaluation indicator that can measure the quality of an image by analyzing the natural statistical characteristics of the image without relying on the original image or reference image. Unlike traditional image quality evaluation methods, NIQE pays more attention to the visual perception effect of the human eye and can more accurately reflect the naturalness and realism of the image.
[0041] A super-resolution remote sensing image generation system, comprising: Generator: obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; Discriminator; the super-resolution image is processed by the depthwise separable convolution layer to obtain the sixth feature, and the negative value is eliminated by the activation function; the seventh feature is obtained through multiple optimization processes, and each optimization process is sequentially processed by the depthwise separable convolution layer, the activation function layer and batch normalization; the seventh feature vector dimension is unified through the fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than the preset threshold. If it is greater, it is judged as the target super-resolution remote sensing image.
[0042] The training method of the generator and the discriminator is to standardize the super-resolution image and the original high-resolution image; extract features from the two standardized images; calculate the distance between the same feature in the two images as the similarity; calculate the loss function based on the similarity, feed the loss function back to the generator, optimize the generator parameters, and transmit the super-resolution image output again to the discriminator to train the discriminator's discrimination ability. When the fluctuation of the loss function is less than the preset value or reaches the preset number of training times after multiple trainings, the training is stopped.
[0043] During the training process, a two-stage training strategy is adopted to ensure that the generator can produce high-quality remote sensing images. First, the generator is trained independently, and the loss function optimizes the generated image features. The loss function calculates the semantic similarity of the image by introducing the CLIP model, evaluates the generated image at the semantic level, and calculates the similarity between the image and the high-level semantic features in the real scene, helping to generate high-quality images that are more in line with human perception. The cross-modal learning ability of the CLIP model between images and texts enables the generator to focus not only on the low-level features of the image (such as pixel-level similarity), but also on the high-level semantic content of the image, thereby generating images that conform to human perception.
[0044] The input super-resolution image and the original high-resolution image are uniformly resized and standardized. Then, the CLIP model is selected to extract features from the standardized image. Finally, the distance between the extracted image features is calculated to measure the similarity between images. The loss function is calculated as follows: , , , Among them, the super-resolution image I and the original high-resolution image GT are encoded as corresponding feature vectors and , The image size after normalization is is the resized image. is the dimension of the feature vector, and Represents the first elements, is the loss function.
[0045] Total loss function Defined as: , in, represents pixel loss, is the perceived loss, To combat losses. They are the proportions of the three loss functions respectively.
[0046] The loss function of the present invention focuses on high-level semantics and structure, overcomes the limitations of traditional loss function calculations, and avoids over-optimization of pixel-level details, which causes images to appear visually unnatural or lack semantic consistency.
[0047] The generator learns how to generate images that are as realistic as possible by optimizing its own loss function without being disturbed by the discriminator's feedback. The specific operations are: The parameters of the generator include the weights and biases of each layer in the generator network. The generator adjusts the weights and biases of each layer through feedback from the loss function.
[0048] This allows the generator to learn to generate effective image features early in training.
[0049] Next, the generator and discriminator are trained alternately to further improve the image quality. The generator "tricks" the discriminator by generating more realistic images, making it unable to distinguish between real images and generated images, while the discriminator improves its ability to distinguish between real and generated images. By enhancing the ability of the discriminator, the generator is pushed to generate more natural and detailed images. The two promote each other. The mutual game between the generator and the discriminator during the training process continuously improves the quality of the generated images and effectively avoids the mode collapse problem, ensuring that the generated images can maintain high quality in multiple aspects (such as structure, texture, and semantic consistency), ultimately improving image quality and avoiding the mode collapse problem.
[0050] The design of activation function layer processing and batch normalization aims to improve the training stability of the network, accelerate the convergence process, and enhance the expression ability of nonlinear features.
[0051] During the optimization process, the generator not only needs to pay attention to pixel-level losses (such as L1 and L2 losses), but also needs to consider the quality of the generated images in terms of natural statistical characteristics.
[0052] In generative adversarial training, the generator optimizes based on the feedback from the discriminator, continuously improving the quality of the generated images and gradually approaching the details of the real images. The discriminator provides accurate discrimination results by comparing the generated images with the real images, helping the generator improve its generation strategy.
[0053] Through this adversarial training process, the discriminator not only promotes the generator to generate higher quality images, but also improves the final effect of the super-resolution reconstruction task. The introduction of deep feature extraction enables the super-resolution generation network to restore the details of high-resolution images more accurately.
[0054] The NIQE (Natural Image Quality Evaluator) indicator is introduced to evaluate the visual quality of the generated images. The introduction of this indicator makes the optimization process more in line with the needs of human visual perception and avoids the problem of visual distortion or unnaturalness of the image caused by over-reliance on pixel-level loss. NIQE is a reference-free image quality evaluation indicator that can measure the quality of an image by analyzing the natural statistical characteristics of the image without relying on the original image or reference image. Unlike traditional image quality assessment methods (such as PSNR and SSIM), NIQE pays more attention to the visual perception effect of the human eye and can more accurately reflect the naturalness and realism of the image. By monitoring the changes in the NIQE indicator of the model on the validation set, the similarity between the generated super-resolution image and the original high-resolution image is regularly evaluated to ensure that the generator can produce high-quality images that are closer to the real image.
[0055] The remote sensing image reconstruction method of the present invention is used to reconstruct the remote sensing images in the AID data set. The reconstructed low-resolution image is obtained by bicubic downsampling. The reconstruction results are as follows: Figure 1 , Figure 2 shown.
[0056] in, Figure 1 and Figure 2 (a) is the original high-resolution image, (b) is the image obtained by bicubic interpolation, (c) is the image obtained by the generative adversarial network model, (d) is the image obtained by the enhanced generative adversarial network model, (e) is the image obtained by the blind super-resolution reconstruction model, (f) is the locally enlarged image of the adaptive blind reconstruction network guided by self-supervised degradation, (g) is the image obtained by the large kernel convolution super-resolution model, (h) is the image obtained based on the diffusion probability model, (i) is the image obtained by the remote sensing image reconstruction method of the present application, and (j) is the locally enlarged image of the original high-resolution image.
[0057] In addition, the remote sensing image reconstruction method of this application is also used to reconstruct the WHU-RS19 dataset. The low-resolution image required for reconstruction is obtained by processing with a Gaussian degenerate kernel of size 0.6, such as Figure 3 , Figure 4 As shown. Among them, (a) is the original high-resolution image, (b) is the local enlarged image of the image obtained by bicubic interpolation, (c) is the local enlarged image of the image obtained by the generative adversarial network model, (d) is the local enlarged image of the image obtained by the enhanced generative adversarial network model, (e) is the local enlarged image of the image obtained by the blind super-resolution reconstruction model, (f) is the image obtained by the large kernel convolution super-resolution model, (g) is the image obtained based on the diffusion probability model, (h) is the local enlarged image of the image obtained by the remote sensing image reconstruction method of the present application, and (i) is the local enlarged image of the original high-resolution image.
[0058] In the above, the image obtained by bicubic interpolation, compared with the present invention, does not use a lightweight attention mechanism and deep separable convolution, but is only based on mathematical interpolation and cannot restore the real high-frequency information.
[0059] The image is obtained through a generative adversarial network model (compared to the present invention, frequency domain feature fusion is not used), and the mapping relationship from low resolution to high resolution is learned through the generative adversarial network model. Because frequency domain feature fusion is not used, some images may be distorted or blurred.
[0060] The image is obtained by the enhanced generative adversarial network model. Compared with the present invention, the present invention does not use depthwise separable convolution, but introduces a reconstructed residual block and a perceptual loss function, which results in high computational complexity and makes it difficult to deploy on edge devices.
[0061] The images obtained by the blind super-resolution reconstruction model, compared with the method of the present invention, do not use feature extraction at multiple scales, but propose a more complex but practical image degradation model, which leads to over-reliance on predefined degradation models and limited generalization.
[0062] Compared with the present invention, the image obtained by the adaptive blind reconstruction network guided by self-supervised degradation requires a large amount of unlabeled data pre-training; the present invention only requires single image training, and the lightweight design can achieve efficient adaptation.
[0063] The image obtained by the large kernel convolution super-resolution model, compared with the method of the present invention, does not use single-head attention. Through large convolution and channel split-shuffle operations, not adding a single-head attention mechanism will result in high computational overhead and high memory usage.
[0064] The image obtained based on the diffusion probability model is compared with the one-way reasoning of the present invention. This method has multiple iterations, too many iterations, and slow reasoning speed.
[0065] Depend on Figure 1 , Figure 2 , Figure 3 , Figure 4 It can be seen that compared with other methods, the remote sensing image reconstruction method of the present application can more effectively simulate the degradation of real images and adaptively fuse information of different resolutions during the reconstruction process, thereby improving the clarity and accuracy of the reconstructed image.
Claims
1. A method for generating a super-resolution remote sensing image, characterized in that: The specific steps are: S1. Obtain the original high-resolution image from the data set, downsample it to obtain a low-resolution image, perform shallow feature extraction by the first convolution to obtain the first feature, perform multiple deep feature extractions to obtain the deep feature, perform residual connection with the first feature to obtain the fifth feature, and perform the second convolution and two upsampling reconstruction processes to map and fuse the fifth feature to obtain a super-resolution image; Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; S2, super-resolution image, is processed by a depthwise separable convolution layer to obtain the sixth feature, and negative values are eliminated by an activation function; the seventh feature is obtained through multiple optimization processes, each of which is processed by a depthwise separable convolution layer, an activation function layer, and batch normalization in sequence; the seventh feature vector dimension is unified through a fully connected layer and mapped to a scalar, and it is determined whether the scalar is greater than a preset threshold. If so, it is determined to be a target super-resolution remote sensing image.
2. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The processing result of the last depth-separable convolutional layer is subjected to attention processing to obtain the fourth feature. The specific operation is: The input features are divided into two parts, where one part of the input features is pre-normalized to obtain the query, key and value. The similarity between the query and the key is calculated, the association between different positions in the feature map is evaluated, the attention weights are normalized to weight the values, and the weighted output features are generated. The weighted output features are connected with the other part of the input features to obtain the fourth feature.
3. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The fifth feature is mapped and fused to obtain a super-resolution image through the second convolution and two upsampling reconstruction processes, specifically: After the second convolution is mapped, a first upsampling reconstruction process is performed to obtain a first reconstructed image, and then a second upsampling reconstruction process is performed to obtain a second reconstructed image, and the first reconstructed image and the second reconstructed image are fused to obtain a super-resolution image.
4. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: When outputting multiple super-resolution remote sensing images, the visual quality of the generated images is evaluated by a no-reference image quality assessment metric.
5. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: The deep feature is obtained by extracting the first feature three times.
6. The method for generating a super-resolution remote sensing image according to claim 1, characterized in that: Repeat the above process of obtaining the third feature ten times, perform attention processing on the processing result of the last depth-separable convolutional layer, and obtain the fourth feature.
7. A super-resolution remote sensing image generation system, characterized in that: include: Generator; The original high-resolution image is obtained from the data set, and the low-resolution image is obtained by downsampling. The first convolution is used to extract shallow features to obtain the first feature. After multiple deep feature extractions, the deep feature is obtained. The fifth feature is obtained by residual connection with the first feature. The fifth feature is mapped and fused through the second convolution and two upsampling reconstruction processes to obtain a super-resolution image. Each deep feature extraction operation is as follows: the depthwise separable convolution layer includes channel-by-channel convolution and point-by-point convolution, the depthwise separable convolution layer is used to extract the first feature, and the negative value is eliminated by the activation function; the second feature is output by the cyclic feature extraction and negative value elimination operation, the second feature is residually connected with the first feature to obtain the third feature as the input of the next depthwise separable convolution layer, the above process of obtaining the third feature is repeated several times, and the processing result of the last depthwise separable convolution layer is subjected to attention processing to obtain the fourth feature; Discriminator; the super-resolution image is processed by the depthwise separable convolution layer to obtain the sixth feature, and the negative value is eliminated by the activation function; the seventh feature is obtained through multiple optimization processes, and each optimization process is sequentially processed by the depthwise separable convolution layer, the activation function layer and batch normalization; the seventh feature vector dimension is unified through the fully connected layer and mapped to a scalar, and it is judged whether the scalar is greater than the preset threshold. If it is greater, it is judged as the target super-resolution remote sensing image.
8. A super-resolution remote sensing image generation system according to claim 7, characterized in that: The training method is to standardize the super-resolution image and the original high-resolution image; extract features from the two standardized images; calculate the distance between the same feature in the two images as the similarity; calculate the loss function based on the similarity, feed the loss function back to the generator, optimize the generator parameters, and transmit the super-resolution image output again to the discriminator to train the discriminator's discrimination ability. When the fluctuation of the loss function is less than the preset value or reaches the preset number of training times after multiple trainings, the training is stopped.
9. A super-resolution remote sensing image generation system according to claim 8, characterized in that: The loss function is fed back to the generator to optimize the generator parameters. The specific operation is: The parameters of the generator include the weights and biases of each layer in the generator network. The generator adjusts the weights and biases of each layer through feedback from the loss function.
Citation Information
Patent Citations
Super-resolution reconstruction method based on attention mechanism
CN113706386A
Remote sensing image super-resolution reconstruction method based on generative adversarial network
CN117114984A
Remote sensing image super-resolution reconstruction method and system based on deep and shallow feature fusion
CN117830100A
Remote sensing image super-resolution reconstruction method and system based on detail recovery
CN118780987A
Generation method, system and apparatus capable of visual resolution enhancement, and storage medium
WO2022242029A1
Cited By
Super-resolution remote sensing image reconstruction method, system and equipment based on frequency domain enhancement
CN121481853A
Multi-modal medical image fusion method, model training method and related device
CN121616473A
Remote sensing image processing method and device based on dynamic degradation modulation
CN121961849A
A remote sensing image processing method and device based on dynamic degradation modulation
CN121961849B