Remote sensing image super-resolution reconstruction method

By constructing an adversarial network model of multi-dimensional interaction and conditional identification, the problems of unreal reconstruction effects and insufficient detail texture in the remote sensing image super-resolution reconstruction method are solved, and a higher quality remote sensing image reconstruction is achieved.

CN120070182AActive Publication Date: 2025-05-30CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510139750.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The reconstruction effect of the existing remote sensing image super-resolution reconstruction method is unreal and the details and textures are insufficient.

Method used

A remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and conditional identification is adopted to construct an adversarial network model including a generator and a discriminator. The generator uses the initial module, a low-dimensional feature guide group, a high-dimensional feature guide group, a multi-dimensional interaction module and a reconstruction module. The discriminator uses the loss function to train the generator and discriminator network through gradient information branches, semantic information branches and gradient-semantic fusion blocks.

Benefits of technology

The details and clarity of the remote sensing image are significantly improved, and the generated images are more realistic and the details are richer, solving the problems of unreal reconstruction effects and insufficient detailed texture in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070182A_ABST
    Figure CN120070182A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image super-resolution reconstruction method, and relates to the technical field of image processing. Comprising the following steps: preparing a data set; constructing a network model; training a network model; finely adjusting the network model; and solidifying the network model. According to the technical scheme, a multi-dimensional interaction and condition identification structure and a training mode are designed, high-dimension and low-dimension feature interaction is introduced into a generator, generation of images with richer details is promoted, gradient and semantic information is introduced into a discriminator, the generator is guided to learn gradient and semantic perception textures with finer grit, and the image quality is improved. Therefore, the generator can generate a more vivid image, and the quality of the super-resolution reconstructed image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a super-resolution reconstruction method for remote sensing images. Background Art

[0002] The super-resolution (SR) technology of remote sensing images is a technology that can convert low-resolution (LR) images into high-resolution (HR) images. Its purpose is to make the images clearer, more detailed and more real by increasing the resolution of the images. Compared with general natural images, remote sensing images have a wider field of view, can capture more extensive ground information, and have richer texture information. However, due to the huge amount of remote sensing image data, the processing process requires a large amount of computing resources and time. In addition, in some cases, limited by sensor performance, data transmission and storage, etc., the resolution of remote sensing images may be low, making it difficult to present sufficient details and clarity, thus limiting their practical application scope. Through the super-resolution technology of remote sensing images, high-resolution images can be comprehensively generated from multiple low-resolution images, thereby significantly improving the details and clarity of remote sensing images. This technology is of great significance for improving the analysis accuracy, object recognition ability and monitoring effect of remote sensing images, and can be widely applied to fields such as land cover classification, building detection and 3D reconstruction.

[0003] The Chinese patent publication number is "CN113034361A", and the name is "A Super-resolution Reconstruction Method for Remote Sensing Images Based on Improved ESRGAN". This technology makes improvements to the super-resolution reconstruction network for remote sensing images, which includes a generation network and a discriminator network: the generation network consists of 64 3×3 convolutional layers, 23 RRDB modules and LeakyReLU activation functions; the discriminator network consists of 6 layers, using a fully convolutional network with an even-sized convolutional kernel, and adding BN layers and LeakyReLU activation layers between these layers for construction; the first layer of the discriminator network receives the original low-resolution remote sensing image realA, the image enlarged by bicubic interpolation, and the image after channel merging of the fakeB image output by the generation network as input. By alternately training the generation network and the discriminator network and updating their parameters, the improvement of the super-resolution reconstruction network model for remote sensing images is finally realized. This method uses a single convolutional network, resulting in a lack of realism in the reconstruction effect and insufficient richness in the texture details of the generated images.

[0004] In summary, how to solve the problems of unrealistic reconstruction effect and insufficient detail texture in the current super-resolution reconstruction method for remote sensing images is an urgent problem for those skilled in the art at present. Summary of the Invention

[0005] The technical solution of the present invention to solve the above technical problems is to provide a remote sensing image super-resolution reconstruction method, including the following steps:

[0006] Step 1, prepare the data set: Obtain a remote sensing image data set and divide it into a training set, a validation set, and a test set according to a ratio;

[0007] Step 2, construct a network model: Construct an adversarial network model including a generator and a discriminator; the generator includes an initial module, a low-dimensional feature guidance group, a high-dimensional feature guidance group, a multi-dimensional interaction module, and a reconstruction module. The low-dimensional feature guidance group includes a two-branch local module and a max pooling layer. The high-dimensional feature guidance group includes a global attention block and a sequential upsampling block; the discriminator includes a gradient information branch, a semantic information branch, a gradient-semantic fusion block, and multiple convolutional layers for evaluating the authenticity of the generated image;

[0008] Step 3, train the network model: Use a loss function to train the generator and discriminator networks until the number of training times reaches an initial set threshold or the value of the loss function reaches a preset range, then the network model training is completed.

[0009] Further, in step 2, the generator includes an initial module, three low-dimensional feature guidance groups, three high-dimensional feature guidance groups, three multi-dimensional interaction modules, and a reconstruction module. There is a skip connection between the low-dimensional feature guidance group and the multi-dimensional interaction module, and a global residual connection is introduced before the reconstruction module.

[0010] Further, the initial module includes an upsampling operation and a 3×3 convolution. The upsampling operation uses the bicubic interpolation method. The low-dimensional feature guidance group includes three two-branch local modules and a max pooling layer, and uses local residual connections to fuse features for extracting low-dimensional multi-scale features of the low-resolution image; the two-branch local module uses channel splitting technology to divide the features into two branch features for processing; branch one consists of a 3×3 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an L-type function, a channel concatenation operation, and a 1×1 convolutional layer, and fuses features through local residual connections. The 3×3 convolutional layer, the 5×5 convolutional layer, and the 7×7 convolutional layer are used to obtain multi-scale information of the input features, and the 1×1 convolutional layer is used for in-depth feature extraction; branch two consists of a 3×3 convolutional layer, a channel attention layer, and a spatial attention layer, and fuses features through local residual connections for extracting local details in the channel dimension and the spatial dimension; after the output features of branch one and branch two pass through the channel concatenation operation, they pass through a 1×1 convolution as the output of the final two-branch local feature module; the max pooling layer is used to output the low-dimensional component of the features.

[0011] Furthermore, the high-dimensional feature guidance group includes a global attention block and a sequence upsampling block; the global attention block includes a spatial self-attention mechanism and a channel self-attention mechanism, which are used to extract global global features. Through the first dimension conversion layer, the dimension is converted from C×H×W to (H W)×C, and dimensionality reduction is performed under the action of a 1×1 convolutional layer to generate Q (query), K (key), and V (value) feature matrices. The Q feature matrix is multiplied by the K feature matrix;

[0012] The attention weight matrix in the spatial dimension is calculated through the softmax function, and the attention weight matrix in the spatial dimension is multiplied by the V feature matrix to obtain the V1 feature matrix;

[0013] Then, the Q, K, and V1 feature matrices pass through the second dimension conversion layer (converting the dimension from (H W)×C to C×(H W)) to obtain Q′, K′, and V′ feature matrices, and the Q′ feature matrix is multiplied by the K′ feature matrix,

[0014] The attention weight matrix in the channel dimension is calculated using softmax, and the attention weight matrix in the channel dimension is multiplied by the V feature matrix to obtain the attention output sequence; the sequence upsampling block includes a multi-layer perceptron, a dimension conversion layer, and a pixel rearrangement layer, which are used to output the high-dimensional components of the features. The multi-layer perceptron doubles the channel dimension by introducing a non-linear transformation. The dimension conversion layer converts the dimension from C×(H W) to C×H×W, and the pixel rearrangement layer uses the PixelShuffle upsampling method;

[0015] The multi-dimensional interaction module includes a 1×1 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an average pooling layer, a max pooling layer, a channel concatenation operation, a channel splitting operation, a Sigmoid function, and a pixel-level multiplication operation. The 1×1 convolutional layer is used to reduce the number of channels and reduce the computational cost of this process. The features pass through the 5×5 convolutional layer and the 7×7 convolutional layer respectively through the channel concatenation operation, and then through pooling and the Sigmoid function calculation to learn the weights of the mixed features. The weights represent the importance of the features under different receptive fields. The calculated weights are multiplied by the low-dimensional guidance and high-dimensional guidance features respectively, and the selected dimensional features are obtained from the mixed features. Finally, the features of the two branches are added to achieve the adaptive fusion interaction of the low-dimensional features and the high-dimensional features;

[0016] The reconstruction module includes a 3×3 convolutional layer and an L-type function, which integrate global residual features and better reconstruct the details and textures of the image.

[0017] Furthermore, in step 2, the discriminator includes a 4×4 convolutional layer, an L-type function, a gradient information branch, a semantic information branch, and three gradient-semantic fusion blocks;

[0018] The gradient information branch includes a gradient extraction block, a 4×4 convolutional layer, and an L-type function for extracting gradient feature information; the 4×4 convolutional layer is used to adjust the scale of the gradient features, and the gradient extraction block uses gradient operators in the vertical and horizontal directions to obtain gradient information;

[0019] The semantic information branch consists of a semantic extraction block, and the pre-trained CLIP "RN50" is used as a semantic extractor to obtain semantic information;

[0020] For the gradient-semantic fusion block, the features obtained by passing the input features through the first convolutional layer, the L-type function, and the first dimension transformation layer are used as the content context information in the cross-attention mechanism, the features obtained by passing the semantic features through the semantic feature optimization block are used as the semantic features in the cross-attention mechanism, the features obtained by passing the gradient features through the gradient feature optimization block are used as the gradient features in the cross-attention mechanism. The content context information and the semantic features are used to obtain distorted semantic perception image features through the cross-attention mechanism, the content context information and the gradient features are used to obtain distorted gradient perception image features through the cross-attention mechanism, the distorted gradient perception image features and the distorted semantic perception image features are used to obtain gradient-semantic perception image features through the cross-attention mechanism, and the perception image features obtained through the layer normalization layer, the G-type function, and the third dimension transformation layer. The perception image features are concatenated with the original enhanced features, and finally the features are output by the second convolutional layer;

[0021] The gradient feature optimization block includes a 3×3 convolutional layer, a spatial attention layer, a dimension transformation layer, a first layer normalization layer, a residual self-attention mechanism, and a second layer normalization layer. Through the spatial attention layer and the residual self-attention mechanism, the gradient information of the global spatial features is extracted;

[0022] The semantic feature optimization block includes a group normalization layer, a dimension transformation layer, a first layer normalization layer, a residual self-attention mechanism, and a second layer normalization layer. The global semantic features are extracted through the residual self-attention mechanism.

[0023] Furthermore, in step 3, the following steps are also included:

[0024] Select a loss function and determine evaluation metrics: Use pixel loss, perceptual loss, and adversarial loss as loss functions, and use value signal-to-noise ratio, structural similarity, and perceptual image similarity as evaluation metrics.

[0025] Furthermore, the following steps are also included:

[0026] Step 4, fine-tune the network model: Use the validation set to adjust the network model and optimize the network model parameters;

[0027] Step 5, freeze the network model: After fine-tuning is completed, freeze the fine-tuned network parameters to determine the final remote sensing image super-resolution network parameters.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. The present invention generates a super-resolution image through a generator, and then inputs the super-resolution image and the real image into a discriminator to identify the authenticity of the image. Using a generative adversarial network in the field of remote sensing super-resolution solves the problem of the unrealistic super-resolution image.

[0030] 2. The present invention designs a gradient-semantic fusion block. Through the cross-attention mechanism, semantic perception and gradient texture are distorted from the image into the discriminator, thereby guiding the generator to learn finer-grained gradients and semantic perception textures, enabling the generator to generate more realistic images.

[0031] 3. The present invention designs a double-branch local module. By extracting rich local feature information through multi-scale and channel-spatial attention double branches, the local details of the reconstructed image are more abundant.

[0032] 4. The present invention designs a multi-dimensional interaction module. Through convolution and pooling operations, low-dimensional features and high-dimensional features are adaptively obtained to achieve interaction and fusion of different dimensions, obtaining richer features, making the quality of the reconstructed image higher and the details more abundant.

[0033] 5. The present invention designs a composite loss function, which consists of three parts: adversarial loss, pixel loss, and perceptual loss. Compared with the single pixel loss, this composite loss makes up for the problem of the overly smooth image caused by the pixel loss. It can also learn the implicit relationship between the low-resolution image and the high-resolution image, helping to promote the network to generate high-quality super-resolution images. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0035] Figure 1 It is the step flowchart of the remote sensing image super-resolution reconstruction method described in the present invention;

[0036] Figure 2 It is the structural schematic diagram of the generator in the method of the present invention;

[0037] Figure 3 It is the structural schematic diagram of the initial module in the method of the present invention;

[0038] Figure 4Schematic diagram of the dual-branch local module in the method of the present invention;

[0039] Figure 5 Schematic diagram of the global attention block in the method of the present invention;

[0040] Figure 6 Schematic diagram of the sequence upsampling block in the method of the present invention;

[0041] Figure 7 Schematic diagram of the multi-dimensional interaction module in the method of the present invention;

[0042] Figure 8 Schematic diagram of the reconstruction module in the method of the present invention;

[0043] Figure 9 Schematic diagram of the discriminator in the method of the present invention;

[0044] Figure 10 Schematic diagram of the gradient-semantic fusion block in the method of the present invention;

[0045] Figure 11 Schematic diagram of the cross-attention mechanism in the method of the present invention;

[0046] Figure 12 Schematic diagram of the gradient feature optimization block in the method of the present invention;

[0047] Figure 13 Schematic diagram of the residual self-attention mechanism in the method of the present invention;

[0048] Figure 14 Schematic diagram of the semantic feature optimization block in the method of the present invention. Detailed implementation manners

[0049] The present invention proposes a remote sensing image super-resolution reconstruction method, aiming to solve the problems of untrue reconstruction effect and insufficient detail texture in the current remote sensing image super-resolution reconstruction methods.

[0050] The remote sensing image super-resolution reconstruction method proposed by the present invention will be described in the following specific embodiments:

[0051] Embodiment 1:

[0052] A remote sensing image super-resolution reconstruction method, as Figure 1 shown, includes the following steps:

[0053] Step 1, prepare the dataset: Obtain the remote sensing image dataset and divide it into a training set, a validation set, and a test set according to a ratio;

[0054] Obtain remote sensing image datasets one and two; divide the data in dataset one into a training set and a validation set according to a certain ratio, while dataset two is used as a test set. Then, perform bicubic downsampling on the high-resolution images in the two datasets to generate low-resolution images, thereby constructing the required image pairs. Finally, perform augmentation preprocessing on the obtained image pair dataset to expand the dataset;

[0055] Step 2, construct a network model: Construct an adversarial network model including a generator and a discriminator; the generator includes an initial module, a low-dimensional feature guidance group, a high-dimensional feature guidance group, a multi-dimensional interaction module, and a reconstruction module. The low-dimensional feature guidance group includes a two-branch local module and a max pooling layer, and the high-dimensional feature guidance group includes a global attention block and a sequential upsampling block; the discriminator includes a gradient information branch, a semantic information branch, a gradient-semantic fusion block, and multiple convolutional layers, which are used to evaluate the authenticity of the generated images;

[0056] Step 3, train the network model: Select a loss function and determine evaluation metrics: Use pixel loss, perceptual loss, and adversarial loss as loss functions, and use peak signal-to-noise ratio, structural similarity, and perceptual image similarity as evaluation metrics;

[0057] Use the loss function to train the generator and discriminator networks until the number of training times reaches the initial set threshold or the value of the loss function reaches the preset range, then the network model training is completed;

[0058] Step 4, fine-tune the network model: Use the validation set to adjust the network model and optimize the network model parameters;

[0059] Step 5, solidify the network model: After the fine-tuning is completed, solidify the fine-tuned network parameters to determine the final remote sensing image super-resolution network parameters.

[0060] Furthermore, in step 2, the generator includes one initial module, three low-dimensional feature guidance groups, three high-dimensional feature guidance groups, three multi-dimensional interaction modules, and one reconstruction module. There is a skip connection between the low-dimensional feature guidance group and the multi-dimensional interaction module, and a global residual connection is introduced before the reconstruction module.

[0061] Furthermore, the initial module includes an upsampling operation and a 3×3 convolution. The upsampling operation adopts the bicubic interpolation method. The low-dimensional feature guidance group includes three double-branch local modules and a max pooling layer, and uses local residual connections to fuse features for extracting low-dimensional multi-scale features of the low-resolution image. The double-branch local module uses the channel splitting technique to divide the features into two branch features for processing. Branch one consists of a 3×3 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an L-shaped function, a channel concatenation operation, and a 1×1 convolutional layer, and fuses features through local residual connections. The 3×3 convolutional layer, 5×5 convolutional layer, and 7×7 convolutional layer are used to obtain multi-scale information of the input features, and the 1×1 convolutional layer is used for in-depth feature extraction. Branch two consists of a 3×3 convolutional layer, a channel attention layer, and a spatial attention layer, and fuses features through local residual connections for extracting local details in the channel dimension and spatial dimension. After the output features of branch one and branch two pass through the channel concatenation operation, they pass through a 1×1 convolution as the output of the final double-branch local feature module. The max pooling layer is used to output the low-dimensional components of the features.

[0062] Furthermore, the high-dimensional feature guidance group includes a global attention block and a sequential upsampling block. The global attention block includes a spatial self-attention mechanism and a channel self-attention mechanism for extracting global global features. Through the first dimension conversion layer, the dimension is converted from C×H×W to (H W)×C, and dimensionality reduction is performed under the action of a 1×1 convolutional layer to generate Q (query), K (key), and V (value) feature matrices, and the Q feature matrix is multiplied by the K feature matrix.

[0063] The attention weight matrix in the spatial dimension is calculated through the softmax function, and the attention weight matrix in the spatial dimension is multiplied by the V feature matrix to obtain the V1 feature matrix.

[0064] Then, the Q, K, and V1 feature matrices pass through the second dimension conversion layer to convert the dimension from (H W)×C to C×(H W) to obtain Q′, K′, and V′ feature matrices, and the Q′ feature matrix is multiplied by the K′ feature matrix.

[0065] The attention weight matrix in the channel dimension is calculated using softmax, and the attention weight matrix in the channel dimension is multiplied by the V feature matrix to obtain the attention output sequence. The sequential upsampling block includes a multi-layer perceptron, a dimension conversion layer, and a pixel shuffling layer for outputting the high-dimensional components of the features. The multi-layer perceptron doubles the channel dimension by introducing a non-linear transformation. The dimension conversion layer converts the dimension from C×(H W) to C×H×W, and the pixel shuffling layer uses the PixelShuffle upsampling method.

[0066] The multi-dimensional interaction module includes a 1×1 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an average pooling layer, a max pooling layer, a channel concatenation operation, a channel splitting operation, a Sigmoid function, and a pixel-wise multiplication operation. The 1×1 convolutional layer is used to reduce the number of channels and lower the computational cost of this process. The features pass through the 5×5 convolutional layer and the 7×7 convolutional layer respectively, and then through the channel concatenation operation. After that, through pooling and Sigmoid function calculations, the weights of the mixed features are learned. The weights represent the importance of features under different receptive fields. The calculated weights are multiplied with the low-dimensional guidance and high-dimensional guidance features respectively to obtain the selected dimensional features from the mixed features. Finally, the features of the two branches are added together to achieve the adaptive fusion interaction of low-dimensional and high-dimensional features.

[0067] The reconstruction module includes a 3×3 convolutional layer and an L-type function, which integrates global residual features to better reconstruct the details and textures of the image.

[0068] Further, in step 2, the discriminator includes a 4×4 convolutional layer, an L-type function, a gradient information branch, a semantic information branch, and three gradient-semantic fusion blocks.

[0069] The gradient information branch includes a gradient extraction block, a 4×4 convolutional layer, and an L-type function, which are used to extract gradient feature information. The 4×4 convolutional layer is used to adjust the scale of the gradient features. The gradient extraction block uses vertical and horizontal gradient operators to obtain gradient information.

[0070] The semantic information branch consists of a semantic extraction block, and the pre-trained CLIP "RN50" is used as the semantic extractor to obtain semantic information.

[0071] For the gradient-semantic fusion block, the features obtained by the input features passing through the convolutional layer one, the L-type function, and the dimension conversion layer one are used as the content context information in the cross-attention mechanism. The features obtained by the semantic features passing through the semantic feature optimization block are used as the semantic features in the cross-attention mechanism. The features obtained by the gradient features passing through the gradient feature optimization block are used as the gradient features in the cross-attention mechanism. The content context information and the semantic features obtain the distorted semantic perception image features through the cross-attention mechanism. The content context information and the gradient features obtain the distorted gradient perception image features through the cross-attention mechanism. The distorted gradient perception image features and the distorted semantic perception image features obtain the gradient-semantic perception image features through the cross-attention mechanism, and the perception image features obtained through the layer normalization layer, the G-type function, and the dimension conversion layer three. The perception image features are connected with the original enhanced features, and finally the features are output by the convolutional layer two.

[0072] The gradient feature optimization block includes a 3×3 convolutional layer, a spatial attention layer, a dimensionality transformation layer, a layer normalization layer 1, a residual self-attention mechanism, and a layer normalization layer 2. Through the spatial attention layer and the residual self-attention mechanism, the gradient information of the global spatial features is extracted;

[0073] The semantic feature optimization block includes a group normalization layer, a dimensionality transformation layer, a layer normalization layer 1, a residual self-attention mechanism, and a layer normalization layer 2. The global semantic features are extracted through the residual self-attention mechanism.

[0074] Example 2:

[0075] A remote sensing image super-resolution reconstruction method specifically includes the following steps:

[0076] Step 1, prepare the dataset;

[0077] Prepare remote sensing dataset 1 as the AID dataset, which contains 10,000 images covering 30 types of scenes (airport, bare land, baseball field, beach, bridge, center, church, commercial, dense residential, desert, farmland, forest, industrial, grassland, medium residential, mountain, park, parking lot, playground, pond, port, railway station, resort, river, school, sparse residential, square, stadium, storage tank, and viaduct), with approximately 200 - 420 images for each type. The scale size of the images is 600 pixels × 600 pixels. Clean the dataset. The cleaned dataset contains 5,000 images. Divide the dataset into a training set and a validation set according to a ratio of 7:3. Prepare remote sensing dataset 2 as the UCMLU dataset with an image scale of 256 pixels × 256 pixels, containing 21 types of scenes, 100 images for each type, a total of 2,100 images. Clean the dataset. The cleaned dataset contains 1,500 images and use this dataset as the test set. Perform 4-fold bicubic downsampling on the cleaned dataset to obtain low-resolution images and construct the paired datasets required for training, validation, and testing.

[0078] Perform image augmentation on the images in the dataset in Step 1. For the same pair of images, randomly flip and crop the high-resolution image to a size of 256×256, and perform the same operation on the low-resolution image to obtain an input image size of 64×64 as the input to the entire network; the random size and position can be implemented through software algorithms; processing the images in the dataset through image augmentation is to enhance the robustness of the network and improve the network generalization ability;

[0079] Step 2, construct the network model; construct a remote sensing image super-resolution reconstruction network including a generator and a discriminator;

[0080] The generator network model, as Figure 2As shown in the figure, the generator consists of an initial module, a low-dimensional feature guidance group 1, a low-dimensional feature guidance group 2, a low-dimensional feature guidance group 3, a high-dimensional feature guidance group 1, a multi-dimensional interaction module 1, a high-dimensional feature guidance group 2, a multi-dimensional interaction module 2, a high-dimensional feature guidance group 3, a multi-dimensional interaction module 3, and a reconstruction module;

[0081] The structure of the initial module is as Figure 3 shown. First, the image is enlarged to the target size through 4-fold upsampling, and then the expression ability of the model is increased through dimensionality elevation, enabling the network to better learn the complex features of the input data. It consists of an upsampling operation and a 3×3 convolutional layer. The upsampling operation uses the bicubic interpolation method, and the convolutional kernel size of the convolutional layer is 3×3, the stride is 1, and the padding is 1;

[0082] The low-dimensional feature guidance group, as Figure 2 shown in the low-dimensional feature guidance group 1 in the middle, can efficiently extract low-dimensional features of various scales. It consists of three double-branch local modules and a max pooling layer. Among them, the specific structure of the double-branch local module is as Figure 4 shown. First, the feature is divided into two branch features using the channel splitting technique and processed; Branch 1 consists of a 3×3 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an L-type function, a channel connection operation, and a 1×1 convolutional layer, and fuses the features through local residual connections. The 3×3 convolutional layer, 5×5 convolutional layer, and 7×7 convolutional layer are used to obtain multi-scale information of the input feature, and the 1×1 convolutional layer is used for in-depth feature extraction; Branch 2 consists of a 3×3 convolutional layer, a channel attention layer, and a spatial attention layer, and fuses the features through local residual connections; The output features of Branch 1 and Branch 2 are concatenated through channels and finally passed through a 1×1 convolutional layer as the output of the final double-branch local feature module; among them, the convolutional kernel size of the 3×3 convolutional layer is 3×3, the stride is 1, and the padding is 1; the convolutional kernel size of the 5×5 convolutional layer is 5×5, the stride is 1, and the padding is 2; the convolutional kernel size of the 7×7 convolutional layer is 7×7, the stride is 1, and the padding is 3; the convolutional kernel size of the 1×1 convolutional layer is 1×1, the stride is 1, and the padding is 0; this process can be expressed by the formula:

[0083] X in =[X 1 ,X 2 ,

[0084] F out =Conv 1×1 ([Branch1(X 1 ),Branch2(X 2 )]),

[0085] where represents the feature of input branch 1, Represents the features of input branch two, Represents the input features of the dual-branch local module, Represents the output features of the dual-branch local module. Branch1(·) (Branch one) represents the convolution, L-type function, and all other operations within Branch1(·) (Branch one) on branch one. Branch2(·) (Branch two) represents the convolution, channel attention, spatial attention, and all other operations within Branch2(·) (Branch two) on branch two. Conv 1×1 Represents a 1×1 convolution operation;

[0086] The high-dimensional feature guidance group, such as Figure 2 Shown in the medium-high dimensional feature guidance group one, can efficiently extract high-dimensional features within the entire domain, consisting of a global attention block and a sequence upsampling block. Among them, the structure of the global attention block is as Figure 5 Shown. The input features enter the global attention block. First, they pass through the dimension conversion layer one (converting the dimension from C×H×W to (H W)×C), and under the action of the 1×1 convolution layer, dimensionality reduction is performed to generate the Q (query), K (key), and V (value) feature matrices. The Q feature matrix is multiplied by the K feature matrix, and then the softmax function is used to calculate the spatial attention weight matrix. The spatial attention weight matrix is multiplied by the V feature matrix to obtain the V1 feature matrix. Then, the Q, K, and V1 feature matrices pass through the dimension conversion layer two (converting the dimension from (H W)×C to C×(H W)) to obtain the Q′, K′, and V′ feature matrices. The Q′ feature matrix is multiplied by the K′ feature matrix, and then softmax is used to calculate the channel attention weight matrix. The channel attention weight matrix is multiplied by the V′ feature matrix to obtain the attention output sequence. The sequence upsampling module consists of a multi-layer perceptron, a dimension conversion layer, and a pixel rearrangement layer, and is used to output the high-dimensional components of the features. The multi-layer perceptron doubles the channel dimension by introducing a non-linear transformation. The dimension conversion layer converts the dimension from C×(H W) to C×H×W, and the pixel rearrangement layer uses the PixelShuffle upsampling method. The spatial self-attention and channel self-attention can be designed as:

[0087]

[0088] Among them, is the scaling factor, K T is the transpose matrix of the K matrix, K′ T is the transpose matrix of the K′ matrix;

[0089] The sequence upsampling block, such as Figure 6As shown in the figure, it consists of a multi-layer perceptron, a dimension conversion layer, and a pixel rearrangement layer, and is used to output the high-dimensional components of features. The multi-layer perceptron doubles the channel dimension by introducing a non-linear transformation. The dimension conversion layer converts the dimension from C×(H W) to C×H×W. The pixel rearrangement layer uses the PixelShuffle upsampling method;

[0090] The structure of the multi-dimensional interaction module is as Figure 7 shown in the figure, and it consists of a 1×1 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an average pooling layer, a max pooling layer, a feature concatenation operation, a feature segmentation operation, a Sigmoid function, and a multiplication operation. The 1×1 convolutional layer is used to reduce the number of channels and reduce the computational cost of this process. The features pass through the 5×5 convolutional layer and the 7×7 convolutional layer respectively, and are concatenated through the channel dimension operation, and then through pooling and Sigmoid function calculations to learn the weights of the mixed features. The weights represent the importance of features under different receptive fields. Multiply the calculated weights with the low-dimensional guidance and high-dimensional guidance features respectively to obtain the selected dimensional features from the mixed features. Finally, add the features of the two branches to achieve the adaptive fusion interaction of low-dimensional features and high-dimensional features; among them, the convolutional kernel size of the 1×1 convolutional layer is 1×1, the stride is 1, and the padding is 0; the convolutional kernel of the 5×5 convolutional layer is 5×5, the stride is 1, and the padding is 2; the convolutional kernel of the 7×7 convolutional layer is 7×7, the stride is 1, and the padding is 3;

[0091] The structure of the reconstruction module is as Figure 8 shown in the figure, and it consists of a 3×3 convolutional layer and an L-shaped function. The convolutional kernel size is 3×3, the stride is 1, and the padding is 1. The reconstruction module can better reconstruct the details and textures of the image;

[0092] The discriminator network model is as Figure 9 shown in the figure, and it consists of a 4×4 convolutional layer, an L-shaped activation function, a gradient extraction branch, a semantic extraction branch, and three gradient-semantic fusion blocks. Among them, the convolutional kernel size of the 4×4 convolutional layer is 4×4, the stride is 2, and the padding is 1. By introducing gradient and semantic information through the gradient information branch and the semantic information branch, the discriminator can better focus on the detailed information of the image and improve the performance of the generator, and can generate more realistic and detailed images;

[0093] The gradient information branch consists of a gradient extraction block, a 4×4 convolutional layer, and an L-shaped function, and is used to extract gradient feature information. The 4×4 convolutional layer is used to adjust the scale size of the gradient features. The convolutional kernel size is 4×4, the stride is 2, and the padding is 1; the gradient extraction block uses the gradient operators in the vertical and horizontal directions to obtain gradient information. The horizontal convolution and vertical convolution used are as follows,

[0094] The horizontal convolutional kernel is:

[0095]

[0096] The vertical convolution kernel is:

[0097]

[0098] Among them, the formulas used for gradient calculation in the horizontal direction, gradient calculation in the vertical direction, and calculation of gradient magnitude are as follows:

[0099]

[0100] Among them, i ∈ [0, 1, 2] represents the three input channels, and ∈ is a small constant to prevent division by zero errors.

[0101] The semantic information branch consists of semantic extraction blocks, and the pre-trained CLIP "RN50" is used as the semantic extractor to obtain semantic information;

[0102] The structure of the gradient-semantic fusion block is as Figure 10 shown, and it is composed of a first convolutional layer, an L-type function, a first dimension transformation layer, a cross-attention mechanism, a gradient feature optimization block, a semantic feature optimization block, a layer normalization layer, a G-type function, and a second dimension transformation layer; among them, the convolutional kernel size of the first convolutional layer is 4×4, the stride is 2, and the padding is 1, and the convolutional kernel size of the second convolutional layer is 3×3, the stride is 1, and the padding is 1; the features obtained by passing the input features through the first convolutional layer, the L-type function, and the first dimension transformation layer (converting the dimension from C×H×W to (H W)×C) are used as the content context information in the cross-attention mechanism, the features obtained by passing the semantic features through the semantic feature optimization block are used as the semantic features in the cross-attention mechanism, the features obtained by passing the gradient features through the gradient feature optimization block are used as the gradient features in the cross-attention mechanism, the content context information and the semantic features pass through the cross-attention mechanism to obtain distorted semantic perception image features, the content context information and the gradient features pass through the cross-attention mechanism to obtain distorted gradient perception image features, the distorted gradient perception image features and the distorted semantic perception image features pass through the cross-attention mechanism to obtain gradient-semantic perception image features, and the perception image features obtained through layer normalization, the G-type function, and the second dimension transformation layer (converting the dimension from (H W)×C to C×H×W) are concatenated with the original enhanced features, and finally the features are output by the second convolutional layer; the structure of the cross-attention mechanism is as Figure 11 shown, the semantic / gradient features are reduced in dimension under the action of a 1×1 convolutional layer to generate a Q (query) feature matrix, the context input features are reduced in dimension under the action of a 1×1 convolutional layer to generate a K (key) and a V (value) feature matrix, the Q matrix and the K matrix are multiplied and then the softmax function is used to calculate the attention weight matrix in the spatial dimension, and the attention weight matrix in the spatial dimension is multiplied by the V feature matrix to obtain the context output feature matrix;

[0103] The structure of the gradient feature optimization block is as follows Figure 12 shown, which consists of a 3×3 convolutional layer, a spatial attention layer, a dimensionality transformation layer, a layer normalization layer 1, a residual self-attention mechanism, and a layer normalization layer 2. Among them, the convolutional kernel size of the 3×3 convolutional layer is 3×3, the stride is 1, and the padding is 1. Through the spatial attention layer and the residual self-attention mechanism, the gradient information of the global spatial features is extracted; among them, the structure of the residual self-attention mechanism is as follows Figure 13 shown. The input features are reduced in dimension under the action of a 1×1 convolutional layer to generate Q (query), K (key), and V (value) feature matrices. The Q matrix and the K matrix are multiplied and then the softmax function is used to calculate the spatial attention weight matrix. The spatial attention weight matrix is multiplied by the V feature matrix to obtain the V2 feature matrix. After the V2 feature matrix passes through a 1×1 convolutional layer, it is multiplied by the original input features to obtain the final output features;

[0104] The structure of the semantic feature optimization block is as follows Figure 14 shown, which consists of a group normalization layer, a dimensionality transformation layer, a layer normalization layer 1, a residual self-attention mechanism, and a layer normalization layer 2. The global semantic features are extracted through the residual self-attention mechanism;

[0105] To ensure the robustness of the network, retain more structural information, and fully extract image features, the present invention uses two sets of activation functions, namely the L-type function and the G-type function. The L-type function and the G-type function are defined as follows:

[0106]

[0107] Step 3, train the network model; First, select an appropriate loss function and determine the evaluation index: According to the network model in step 2, select an appropriate loss function to minimize the difference between the super-resolution image reconstructed by the network and the high-resolution image, and determine the evaluation index to evaluate the performance of the network; The loss function adopted by the generator consists of three types: pixel loss, adversarial loss, and perceptual loss; The discriminator only adopts adversarial loss;

[0108] In the supervised image super-resolution task, in order to make the generated image (SR) as close as possible to the real high-resolution image (HR), use L 1 loss to calculate the error between the values of the corresponding pixel positions of SR and HR, that is, the pixel loss. The formula is as follows:

[0109]

[0110] Among them, x is the input low-resolution image, y is the high-resolution image, and G(·) is the generator.

[0111] To help the generator network improve the quality of super-resolution reconstructed images, correct reconstruction details, and learn the implicit relationship between super-resolution images and real images, adversarial loss is adopted, and the formula is as follows:

[0112]

[0113] Among them, x is the input low-resolution image, y is the high-resolution image, and G(·) and D(·) are the generator and discriminator respectively.

[0114] The generator improves the image quality by capturing the semantic information and structural features of the image, and uses the pre-trained VGG-19 network to calculate the feature difference between the generated image and the real image, that is, the perceptual loss, and the formula is as follows:

[0115]

[0116] Among them, l represents the convolutional layer of the VGG network, λ l represents the weight of the l-th layer, M l , N l respectively represent the sizes of the feature maps of the l-th layer. φ l (G(x)) and φ l (y) represent the feature representations of the generated image and the real image at the l-th layer, and i, j represent the pixel indices in the feature map.

[0117] The total loss function of the generator is as follows:

[0118] L G = L 1 + λ a L adv + λ p L per ,

[0119] Among them, λ a , λ p respectively represent the weights of different loss shares in the total loss of the generator. The setting of the weights is based on preliminary experiments on the training dataset. In this experiment, λ a = 0.1, λ p = 0.5. By optimizing the final overall loss function, it helps the network learn clearer edges and more detailed textures, making the reconstructed super-resolution images more realistic and having better visual effects.

[0120] The discriminator loss adopts adversarial loss, and its formula is as follows:

[0121]

[0122] Among them, x is the input low-resolution image, y is the high-resolution image, and G(·) and D(·) are the generator and discriminator respectively.

[0123] The evaluation metrics selected are peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and perceptual image similarity (LPIPS); PSNR is used to measure the difference between the super-resolved image and the real image, and the higher the PSNR, the better the quality of the super-resolved image; SSIM measures the contour retention during the super-resolution process by calculating the structural similarity between the super-resolved image and the real image, and the higher the SSIM, the smaller the difference; LPIPS focuses on the perceptual similarity between the super-resolved image and the real image, and the lower the LPIPS, the more similar the super-resolved image and the real image are. The definitions of PSNR, SSIM, and LPIPS are as follows:

[0124]

[0125] where μ x , μ y represent the means of images x and y respectively, and represent the standard deviations of images x and y respectively, σ xy represents the covariance of images x and y, C 1 and C 2 are constants, w l is a trainable weight parameter;

[0126] Secondly, train the remote sensing super-resolution network model: start training the network using the loss function until the number of training times reaches the set threshold or the value of the loss function falls within the set range, then the model parameters are considered to be trained, and save the model parameters; select the test dataset of dataset one to test the network model, and the model performance is evaluated by the above-selected evaluation metrics;

[0127] All experiments were carried out on the AutoDL cloud server, and the NVIDIA RTX 3080Ti GPU with 12GB was used for algorithm acceleration. The training period was set to 200 epochs, and the learning rates of the generator and the discriminator were both set to 1e-4; the Adam optimizer was selected as the network parameter optimizer. Its main advantages are simple implementation, high computational efficiency, low memory requirements, and the parameter updates are not affected by the scaling transformation of the gradient, making the parameters relatively stable; when the ability of the discriminator to judge fake images and the ability of the generator to generate images that can deceive the discriminator reach a balance, the network is considered to be basically trained;

[0128] Step 4, fine-tune the network model; use the validation set of dataset one in step 1 to adjust the network, optimize the network model parameters, and further improve the performance of the super-resolution reconstruction network;

[0129] Step 5, solidify the network model; after the fine-tuning in Step 4 is completed, solidify the fine-tuned network parameters to determine the final super-resolution network parameters; when subsequent work requires remote sensing image super-resolution reconstruction, directly use the low-resolution remote sensing image as the input of the super-resolution network to obtain a better super-resolution result.

[0130] Among them, the implementations of convolution, activation function, multi-layer perceptron, feature splicing operation, max pooling operation, average pooling operation, CLIP RN50 semantic extraction operation, gradient operator operation, instance normalization, layer normalization, channel attention, spatial attention, and group normalization are algorithms well-known to those skilled in the art, and the specific processes and methods can be found in corresponding textbooks or technical documents.

[0131] The present invention constructs a remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and conditional discrimination, which can directly super-resolve a low-resolution remote sensing image into a high-resolution image; under the same conditions, by calculating the relevant metrics of the images obtained by the existing methods, the feasibility and superiority of this method are further verified; the existing technical solution is a generative adversarial network; the generator adopts the RRDB structure, that is, dense residual connections, and at the same time, both ends of the residual edge are connected in a Concat manner, and is composed of a convolutional layer and 23 RRDBs; the discriminator adopts the SeD structure, that is, a semantic-aware discriminator, and is composed of 3 semantic-aware fusion blocks and 2 convolutional layers; this method comes from the semantic-aware discriminator SeD proposed by the University of Science and Technology of China for image super-resolution. The paper is SeD: Semantic-Aware Discriminator for Image Super-Resolution. The comparison of the relevant metrics of the existing technology and the method proposed in the present invention is shown as follows;

[0132]

[0133] As mentioned above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A remote sensing image super-resolution reconstruction method, characterized in that: The following steps are involved: Step 1, prepare the data set: obtain the remote sensing image data set and divide it into training set, validation set, and test set in proportion; Step 2, building a network model: building an adversarial network model including a generator and a discriminator; the generator includes an initial module, a low-dimensional feature guidance group, a high-dimensional feature guidance group, a multi-dimensional interaction module, and a reconstruction module; the low-dimensional feature guidance group includes a dual-branch local module and a maximum pooling layer; the high-dimensional feature guidance group includes a global attention block and a sequence upsampling block; the discriminator includes a gradient information branch, a semantic information branch, a gradient-semantic fusion block, and multiple convolutional layers, which are used to evaluate the authenticity of the generated image; Step 3, training the network model: Use the loss function to train the generator and discriminator networks until the number of training times reaches the initial set threshold or the value of the loss function reaches the preset range, then the network model training is completed.

2. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 1 is characterized in that: In step 2, the generator includes an initial module, three low-dimensional feature guidance groups, three high-dimensional feature guidance groups, three multi-dimensional interaction modules and a reconstruction module. A jump connection is used between the low-dimensional feature guidance group and the multi-dimensional interaction module, and a global residual connection is introduced before the reconstruction module.

3. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 2 is characterized in that: The initial module includes an upsampling operation and a 3×3 convolution, the upsampling operation adopts a bicubic interpolation method, the low-dimensional feature guidance group includes three dual-branch local modules and a maximum pooling layer, and uses local residual connection fusion features to extract low-dimensional multi-scale features of low-resolution images; the dual-branch local module uses channel segmentation technology to divide the features into two branch features and process them; Branch 1 consists of a 3×3 convolutional layer, a 5×5 convolutional layer, a 7×7 convolutional layer, an L-type function, a channel splicing operation, and a 1×1 convolutional layer, and features are fused through local residual connections. The 3×3 convolutional layer, the 5×5 convolutional layer, and the 7×7 convolutional layer are used to obtain multi-scale information of the input features, and the 1×1 convolutional layer is used to extract deep features; Branch 2 consists of a 3×3 convolutional layer, a channel attention layer, and a spatial attention layer, and fuses features through local residual connections to extract local details in the channel and spatial dimensions; the output features of branch 1 and branch 2 are concatenated through the channel and then subjected to a 1×1 convolution as the output of the final dual-branch local feature module; the maximum pooling layer is used to output the low-dimensional components of the features.

4. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 2, characterized in that: The high-dimensional feature guidance group includes a global attention block and a sequence upsampling block; the global attention block includes a spatial self-attention mechanism and a channel self-attention mechanism, which are used to extract global features of the entire domain, and convert the dimension from C×H×W to (HW)×C through a dimension conversion layer 1, and perform dimensionality reduction under the action of a 1×1 convolution layer to generate Q, K and V feature matrices, and the Q feature matrix is ​​multiplied by the K feature matrix; The attention weight matrix of the spatial dimension is calculated by the softmax function, and the V1 feature matrix is ​​obtained by multiplying the attention weight matrix of the spatial dimension with the V feature matrix; Then, the Q, K, and V1 feature matrices are transformed through the second dimension conversion layer, and the dimensions are converted from (HW)×C to C×(HW) to obtain the Q′, K′, and V′ feature matrices. The Q′ feature matrix is ​​multiplied by the K′ feature matrix. The attention weight matrix of the channel dimension is obtained by using softmax calculation, and the attention weight matrix of the channel dimension is multiplied by the V feature matrix to obtain the attention output sequence; the sequence upsampling block includes a multi-layer perceptron, a dimension conversion layer and a pixel rearrangement layer, which are used to output the high-dimensional components of the features. The multi-layer perceptron doubles the channel dimension by introducing nonlinear transformation, the dimension conversion layer converts the dimension from C×(HW) to C×H×W, and the pixel rearrangement layer uses the PixelShuffle upsampling method; The multi-dimensional interaction module includes a 1×1 convolution layer, a 5×5 convolution layer, a 7×7 convolution layer, an average pooling layer, a maximum pooling layer, a channel splicing operation, a channel segmentation operation, a Sigmoid function, and a pixel-level multiplication operation. The 1×1 convolution layer is used to reduce the number of channels and reduce the computational cost of the process. The features are respectively passed through a 5×5 convolution layer and a 7×7 convolution layer through a channel splicing operation, and then through pooling and Sigmoid function calculation to learn the weights of the mixed features. The weights represent the importance of the features under different receptive fields. The calculated weights are respectively multiplied with the low-dimensional guide and high-dimensional guide features to obtain the selected dimensional features from the mixed features. Finally, the features of the two branches are added to realize the adaptive fusion interaction of the low-dimensional features and the high-dimensional features. The reconstruction module includes a 3×3 convolutional layer and an L-type function, which integrates global residual features to better reconstruct image details and textures.

5. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 1, characterized in that: In step 2, the discriminator includes a 4×4 convolutional layer, an L-type function, a gradient information branch, a semantic information branch, and three gradient-semantic fusion blocks; The gradient information branch includes a gradient extraction block, a 4×4 convolution layer and an L-type function, which are used to extract gradient feature information; the 4×4 convolution layer is used to adjust the scale of the gradient feature, and the gradient extraction block uses vertical and horizontal gradient operators to obtain gradient information; The semantic information branch is composed of semantic extraction blocks, and uses the pre-trained CLIP "RN50" as a semantic extractor to obtain semantic information; In the gradient-semantic fusion block, the input feature is obtained through the convolution layer 1, the L-type function and the dimension conversion layer 1 as the content context information in the cross attention mechanism, the semantic feature is obtained through the semantic feature optimization block as the semantic feature in the cross attention mechanism, the gradient feature is obtained through the gradient feature optimization block as the gradient feature in the cross attention mechanism, the content context information and the semantic feature are used to obtain the distorted semantic perception image feature through the cross attention mechanism, the content context information and the gradient feature are used to obtain the distorted gradient perception image feature through the cross attention mechanism, the distorted gradient perception image feature and the distorted semantic perception image feature are used to obtain the gradient-semantic perception image feature through the cross attention mechanism, and the perceived image feature is obtained through the layer normalization layer, the G-type function and the dimension conversion layer 3, the perceived image feature is connected with the original enhanced feature, and finally the feature is output by the convolution layer 2; The gradient feature optimization block includes a 3×3 convolution layer, a spatial attention layer, a dimension conversion layer, a layer normalization layer 1, a residual self-attention mechanism and a layer normalization layer 2. The gradient information of the global spatial features is extracted through the spatial attention layer and the residual self-attention mechanism. The semantic feature optimization block includes a group normalization layer, a dimension conversion layer, a layer normalization layer 1, a residual self-attention mechanism and a layer normalization layer 2, and extracts global semantic features through the residual self-attention mechanism.

6. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 1, characterized in that: In step 3, the following steps are also included: Select the loss function and determine the evaluation index: use pixel loss, perceptual loss and adversarial loss as the loss function, and use value signal-to-noise ratio, structural similarity and perceptual image similarity as the evaluation index.

7. The remote sensing image super-resolution reconstruction method based on multi-dimensional interaction and condition identification according to claim 1, characterized in that: The following steps are also included: Step 4, fine-tune the network model: use the validation set to adjust the network model and optimize the network model parameters; Step 5, solidify the network model: after fine-tuning is completed, solidify the fine-tuned network parameters and determine the final remote sensing image super-resolution network parameters.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method based on improved ESRGAN

    CN113034361A

  • Image super-resolution reconstruction method based on generative adversarial network

    CN111429355A

  • Remote sensing image reconstruction system based on attention mechanism

    CN113269848A

  • Adversarial hyperspectral and multispectral remote sensing fusion method

    CN116468645A

  • Satellite remote sensing image super-resolution reconstruction technology based on double-channel generative adversarial network

    CN116523742A

Cited By

  • Remote sensing image fine-grained target identification method based on local multi-scale super-division

    CN120808177A

  • Internet of vehicles semantic perception method based on multi-source data coevolution and dynamic verification

    CN121278317A

  • A Semantic Perception Method for Vehicle Networks Based on Multi-Source Data Co-evolution and Dynamic Verification

    CN121278317B

  • Abdominal tumor angiography magnetic resonance image generation method based on cross semantic perception

    CN121582392A