Model self-fine-tuning remote sensing image super-resolution reconstruction method, device and medium
Through a two-stage training method, the L1 loss function and the visual adversarial loss function are used to fine-tune the remote sensing image super-resolution network, which solves the problem of difficulty in learning high-frequency detail information in remote sensing image super-resolution reconstruction and achieves clearer high-resolution image reconstruction.
Patent Information
- Application Number
- CN202510633885.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-16
AI Technical Summary
In the super-resolution reconstruction of remote sensing images, existing deep learning methods find it difficult to learn the high-frequency detail information of the real high-resolution image when the difference between the input and the true value is too large, resulting in blurred super-resolution reconstruction results.
A two-stage training method is adopted. In the first stage, the super-resolution network is trained through the L1 loss function to learn the restoration information of high-resolution images. In the second stage, the first stage model is used as the feature loss function and fine-tuned in combination with visual loss and adversarial loss to obtain high-frequency feature information.
The clarity of super-resolution reconstruction of remote sensing images is improved, and more realistic high-resolution image results are obtained.
Smart Images

Figure CN120707383A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image reconstruction, and in particular to a remote sensing image super-resolution reconstruction method, equipment and medium with self-fine-tuning model. Background Art
[0002] Remote sensing images are a type of detection system installed on flying platforms such as satellites and aircraft to obtain information about land objects on the earth's surface. After being interpreted through a series of technical means, they can effectively provide effective image support for applications that require relevant land object information.
[0003] Remote sensing images are usually surface information taken by satellites in outer space. The resolution of their images is usually limited by the long detection distance. The information of detailed ground objects is usually small and difficult to distinguish in detail, which greatly hinders their application in related fields.
[0004] To address these issues, current mainstream methods often use deep learning to super-resolution remote sensing imagery. This approach, by increasing the resolution of remote sensing images, improves the ability to resolve relevant ground object information, providing effective technical support for applications in related fields. Typically, existing deep learning methods use a one-stage training approach to train the model. After achieving a good fit, the network model with the highest PSNR and SSIM metrics is selected as the trained model. This approach has achieved good super-resolution upscaling results in practical applications. However, this training approach is affected by the discrepancy between the true values and the input in the dataset. If the discrepancy between the true values and the input is too large, the one-stage training approach makes it difficult for the model to learn the high-frequency details in the true high-resolution image, resulting in blurred details in the super-resolution reconstruction results. Summary of the Invention
[0005] The purpose of the present invention is to propose a method, device and medium for super-resolution reconstruction of remote sensing images with self-fine-tuning of the model, so as to solve the problem that the existing super-resolution reconstruction network of remote sensing images has difficulty in learning the high-frequency detail information of the real high-resolution image due to the large difference between the input and the true value, thereby resulting in blurred details in the super-resolution reconstruction results.
[0006] Specifically, the present invention provides a method, device, and medium for super-resolution reconstruction of remote sensing images using self-fine-tuning models, the method comprising the following steps: S1. Build a remote sensing image super-resolution reconstruction network; S2, inputting the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for the first stage training to obtain the first stage super-resolution reconstruction result; S3. Import the training parameters of the remote sensing image super-resolution reconstruction network in the first stage into the second stage, train the remote sensing image super-resolution reconstruction network in the second stage, fine-tune the parameters of the remote sensing image super-resolution reconstruction network, and obtain the super-resolution reconstruction results in the second stage.
[0007] A storage medium stores instructions and data for implementing a remote sensing image super-resolution reconstruction method with self-fine-tuning model.
[0008] A remote sensing image super-resolution reconstruction device with self-fine-tuning model comprises: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a remote sensing image super-resolution reconstruction method with self-fine-tuning model.
[0009] The beneficial effects provided by the present invention are as follows: the present invention proposes a remote sensing image super-resolution reconstruction method of a two-stage self-fine-tuning model, which learns as much restoration information of high-resolution remote sensing images as possible during the first-stage training process, and then uses the first-stage model as a feature loss function in the second stage and combines visual loss and adversarial loss as the overall loss function to fine-tune the super-resolution network trained in the first stage, further obtaining high-frequency feature information that cannot be learned in the first stage, thereby obtaining more realistic high-resolution remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a simple flow chart of the method of the present invention; Figure 2 It is a schematic diagram of the super-resolution network structure; Figure 3 It is a schematic diagram of the deep feature extraction module; Figure 4 It is a schematic diagram of the deep feature extraction function; Figure 5 It is a schematic diagram of super-resolution reconstruction results; Figure 6 It is a schematic diagram of the working of the hardware device of an embodiment of the present invention. DETAILED DESCRIPTION
[0011] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0012] Before formally explaining the present invention, the scheme of the present invention is first generally explained for easy understanding.
[0013] Please refer to Figure 1 The present invention provides a remote sensing image super-resolution reconstruction method with model self-fine-tuning, comprising: S1. Build a remote sensing image super-resolution reconstruction network; It should be noted that the remote sensing image super-resolution reconstruction network in this invention is based on the SRNet architecture. Figure 2 , including: a shallow feature extraction module, multiple sequentially connected deep feature extraction modules and an upsampling module.
[0014] S2, inputting the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for the first stage training to obtain the first stage super-resolution reconstruction result; It should be noted that the specific process of obtaining the super-resolution reconstruction result of the first stage in step S2 is as follows: S21, inputting the low-resolution remote sensing image into a shallow feature extraction module to obtain a high-dimensional shallow feature vector; S22, inputting the high-dimensional shallow feature vector into a plurality of sequentially connected deep feature extraction modules to obtain a sufficient high-dimensional feature vector required for super-resolution reconstruction; S23, amplifying the high-dimensional feature vector required for super-resolution reconstruction through an upsampling module to obtain high-resolution residual information; S24. After bilinear interpolation and amplification of the low-resolution remote sensing image, it is fused with the high-resolution residual information to obtain the first-stage super-resolution reconstruction result.
[0015] As a specific embodiment, in the first stage of training, only the L1 loss function is used for gradient back propagation to achieve network fitting, so that the super-resolution network can learn as much super-resolution restoration information as possible. The overall method flow is as follows: Figure 1 Here, the low-resolution remote sensing image is input into the super-resolution network to obtain the first super-resolution reconstruction result of the first stage, as shown below: (1) in, 、 They are low-resolution remote sensing images, the first super-resolution reconstruction results of the first stage, The super-resolution network used is as follows: Figure 2 shown.
[0016] Next, the super-resolution network First, the low-resolution remote sensing image is input into the shallow feature extraction module to convert it into a high-dimensional shallow feature vector, as shown below: (2) in, To extract the high-dimensional shallow feature vector, It is a 3*3 convolution operation (shallow feature extraction module), with 3 input channels and 64 output channels.
[0017] Next, the shallow feature vector extracted Input into N deep feature extraction modules to extract deep feature vectors. In the embodiment of the present invention, N is set to N=32. Correspondingly, the schematic diagram of the deep feature extraction module is as follows Figure 3 shown.
[0018] First, the shallow feature vector Input to the 1*1 convolution layer as follows: (3) in, To extract the high-dimensional deep feature vector, It is a 1*1 convolution operation.
[0019] Next, the high-dimensional deep feature vector Input to the 3*3 convolution layer and activation function, specifically: (4) in, In order to further extract the high-dimensional deep feature vector, It is a 1*1 convolution operation. is the torch.nn.ReLU function in pytorch. Then, the high-dimensional deep feature vector After inputting into the 1*1 convolution layer, add the shallow feature vector input by this module , and obtain the fully fused deep feature vector as follows: (5) in, is the fully fused deep feature vector, is a 1*1 convolution operation. After convolution operation, activation function and residual connection again, we can further obtain the required deep feature vector. Specifically, first Input to the 1*1 convolution layer and activation function as follows: (6) in, To extract the high-dimensional deep feature vector, It is a 1*1 convolution operation. is the torch.nn.ReLU function in pytorch. Then, After inputting to the 1*1 convolution layer, add , and obtain the fully fused deep feature vector as follows: (7) in, is the deep feature vector output by the deep feature extraction module, It is a 1*1 convolution operation.
[0020] Next, the deep feature vector extracted by the deep feature extraction module is , and then input it into the next deep feature extraction module again, and repeat the operations of formula (3)-(7) for N=32 times, so that the deep features are fully integrated and contain the image information required for super-resolution reconstruction.
[0021] Through a series of deep feature extraction modules, we obtain the high-dimensional feature vectors required for super-resolution reconstruction. Next, we pass this through the upsampling module to obtain high-resolution residual information, which is then added to the interpolation and amplification results of the original low-resolution remote sensing image to obtain the final super-resolution reconstruction result, as shown below: (8) in, It is a 3*3 convolution operation. is the torch.nn.PixelShuffle function in pytorch, and bic is the bilinear interpolation amplification operation. This is the first super-resolution reconstruction result of the first stage.
[0022] The above is a detailed description of the super-resolution network. , which can transform low-resolution remote sensing images into high-resolution remote sensing images. Here, the training process of the first stage super-resolution network is described in detail. In the first stage of training, the super-resolution network To learn as much restoration information as possible from high-resolution remote sensing images, the network must be able to maximize its learning capacity. Therefore, in the first stage, only the L1 pixel loss is used to perform gradient backpropagation on the super-resolution network, and the corresponding network parameters are output during model fitting. The specific L1 loss is shown below: (9) in, 、 are low-resolution remote sensing images and high-resolution remote sensing images used in training. is the L1 norm, The loss function used in the one-stage training process.
[0023] Formula (9) is used as the loss function during network training. When the number of iterations is The model is considered to have reached a fitting state and the current network parameters are output Used as a second-stage fine-tuning.
[0024] S3. Import the training parameters of the remote sensing image super-resolution reconstruction network in the first stage into the second stage, train the remote sensing image super-resolution reconstruction network in the second stage, fine-tune the parameters of the remote sensing image super-resolution reconstruction network, and obtain the super-resolution reconstruction results in the second stage.
[0025] It should be noted that during the second stage of training in step S3, the upsampling module and residual connection in the remote sensing image super-resolution reconstruction network trained in the first stage are removed, and the remaining part is used as a deep feature extraction function. , used as the feature loss function in the fine-tuning stage, the deep feature extraction function The corresponding parameters of are inherited from the network parameters trained in the first stage, and only the gradient is updated using the following loss function:
[0026] in, 、 are low-resolution remote sensing images and high-resolution remote sensing images used in training. A network trained in two stages The remaining part of the desampling module and residual connection in ; It uses the classic VGG pre-trained visual loss. It is the classic UNetDiscriminatorSN adversarial loss, are the weight parameters of the corresponding loss function parts; is the L1 norm.
[0027] As an embodiment, through a one-stage training process, the network parameters of the one-stage training are obtained. , and use the training parameters to carry out the second stage of fine-tuning process to obtain high-frequency feature information that cannot be learned in the first stage, thereby obtaining more realistic high-resolution remote sensing images.
[0028] The specific process is to train the super-resolution network in one stage The sampling module and residual connection in are removed, and the remaining part is used as a deep feature extraction function , used as the feature loss function in the model fine-tuning stage, the corresponding structure is shown in Figure 4 As shown. The deep feature extraction function The corresponding parameters are the network parameters trained in one stage Inherited, no gradient update changes are performed during the second stage fine-tuning process. During the second stage training process, the super-resolution network trained in the stage Continue training based on the loss function and update only the gradient as follows: (10) in, is the low-resolution remote sensing image used in training, A network trained in two stages The remaining part of the removal sampling module and residual connection (detailed structure Figure 4 shown); It uses the classic VGG pre-trained visual loss. It is the classic UNetDiscriminatorSN adversarial loss, They are the weight parameters of the corresponding loss function parts, which are set here as .
[0029] Formula (10) is used as the loss function in the two-stage fine-tuning process. When the number of iterations is The model is considered to have reached a fitting state and the current network parameters are output Output as the final fine-tuning result.
[0030] Finally, using the fine-tuning , which can transform low-resolution remote sensing images into clear high-resolution remote sensing images, as shown below: (11) in, 、 They are low-resolution remote sensing images and super-resolution reconstruction results after two-stage network fine-tuning. It is a two-stage fine-tuned super-resolution network.
[0031] Please refer to Figure 5 , Figure 5 Schematic diagram of the results of the embodiment of the present invention. Figure 5 (a) in the figure represents a low-resolution remote sensing image; (b) represents the super-resolution reconstruction result of the first stage; (c) represents the super-resolution reconstruction result of the second stage. Figure 5 It can be seen from the results that the present invention learns as much restoration information of high-resolution remote sensing images as possible during the first stage of training, and then uses the first stage model as the feature loss function in the second stage and combines visual loss and adversarial loss as the overall loss function to fine-tune the super-resolution network trained in the first stage, further acquiring high-frequency feature information that cannot be learned in the first stage, thereby obtaining more realistic high-resolution remote sensing images.
[0032] See Figure 6 , Figure 6 4 is a schematic diagram of the working of the hardware device of an embodiment of the present invention, wherein the hardware device specifically comprises: a remote sensing image super-resolution reconstruction device 401 with self-fine-tuning model, a processor 402 and a storage medium 403.
[0033] A remote sensing image super-resolution reconstruction device 401 with self-fine-tuning model: The remote sensing image super-resolution reconstruction device 401 with self-fine-tuning model implements the remote sensing image super-resolution reconstruction method with self-fine-tuning model.
[0034] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the remote sensing image super-resolution reconstruction method with self-fine-tuning model.
[0035] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the remote sensing image super-resolution reconstruction method with self-fine-tuning of the model.
[0036] In general, the beneficial effects of the present invention are as follows: the present invention proposes a remote sensing image super-resolution reconstruction method with a two-stage self-fine-tuning model, which learns as much restoration information of high-resolution remote sensing images as possible during the first stage of training, and then in the second stage uses the first stage model as the feature loss function and combines visual loss and adversarial loss as the overall loss function to fine-tune the super-resolution network trained in the first stage, further obtaining high-frequency feature information that cannot be learned in the first stage, thereby obtaining more realistic high-resolution remote sensing images.
[0037] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A remote sensing image super-resolution reconstruction method with model self-fine-tuning, characterized by: The following steps are involved: S1. Build a remote sensing image super-resolution reconstruction network; S2, inputting the low-resolution remote sensing image into the remote sensing image super-resolution reconstruction network for the first stage training to obtain the first stage super-resolution reconstruction result; S3. Import the training parameters of the remote sensing image super-resolution reconstruction network in the first stage into the second stage, train the remote sensing image super-resolution reconstruction network in the second stage, fine-tune the parameters of the remote sensing image super-resolution reconstruction network, and obtain the super-resolution reconstruction results in the second stage.
2. The remote sensing image super-resolution reconstruction method with model self-fine-tuning according to claim 1, characterized in that: The remote sensing image super-resolution reconstruction network in step S1 is based on the SRNet architecture, including: a shallow feature extraction module, a plurality of sequentially connected deep feature extraction modules and an upsampling module.
3. The remote sensing image super-resolution reconstruction method with model self-fine-tuning according to claim 2, characterized in that: The specific process of obtaining the super-resolution reconstruction result of the first stage in step S2 is as follows: S21, inputting the low-resolution remote sensing image into a shallow feature extraction module to obtain a high-dimensional shallow feature vector; S22, inputting the high-dimensional shallow feature vector into a plurality of sequentially connected deep feature extraction modules to obtain a sufficient high-dimensional feature vector required for super-resolution reconstruction; S23, amplifying the high-dimensional feature vector required for super-resolution reconstruction through an upsampling module to obtain high-resolution residual information; S24. After bilinear interpolation and amplification of the low-resolution remote sensing image, it is fused with the high-resolution residual information to obtain the first-stage super-resolution reconstruction result.
4. The remote sensing image super-resolution reconstruction method with model self-fine-tuning according to claim 1, characterized in that: In step S2, the first stage of training only uses the L1 loss function to train the remote sensing image super-resolution reconstruction network.
5. The remote sensing image super-resolution reconstruction method with model self-fine-tuning according to claim 1, characterized in that: Each deep feature extraction module includes the following connected in sequence: a first 1*1 convolutional layer, a 3*3 convolutional layer, an activation function and a second 1*1 convolutional layer.
6. The remote sensing image super-resolution reconstruction method with model self-fine-tuning according to claim 1, characterized in that: During the second stage of training in step S3, the upsampling module and residual connection in the remote sensing image super-resolution reconstruction network trained in the first stage are removed, and the remaining part is used as a deep feature extraction function , used as the feature loss function in the fine-tuning stage, the deep feature extraction function The corresponding parameters of are inherited from the network parameters trained in the first stage, and only the gradient is updated using the following loss function: in, 、 are low-resolution remote sensing images and high-resolution remote sensing images used in training. A network trained in two stages The remaining part of the desampling module and residual connection in ; It uses the classic VGG pre-trained visual loss. It is the classic UNetDiscriminatorSN adversarial loss, are the weight parameters of the corresponding loss function parts; is the L1 norm.
7. A storage medium, characterized in that: The storage medium stores instructions and data for implementing a remote sensing image super-resolution reconstruction method with model self-fine-tuning as described in any one of claims 1 to 6.
8. A remote sensing image super-resolution reconstruction device with self-fine-tuning model, characterized by: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement a remote sensing image super-resolution reconstruction method with self-fine-tuning of a model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Deep feature aggregation real-time visual target tracking method based on twin network
CN111489361A
Landslide remote sensing image set super-resolution reconstruction method and system
CN115187463A
Remote sensing image super-resolution reconstruction method and device and storage medium
CN118134767A
Remote sensing image super-division reconstruction method based on convolution scale uncertainty
CN119722456A
Generative Adversarial Network for Joint Light Field Super-resolution and Deblurring and its Operation Method
KR102334730B1