A Short-Time Delay Super-Resolution Method for Natural Disaster Satellite Remote Sensing Images Based on Spatial Shift Convolution

Through the short-term delay super-segment method of natural disaster satellite remote sensing images based on spatial shift convolution, the problem of insufficient resolution of satellite remote sensing images and large demand for deep learning models is solved, and the generation of high-resolution images and significant improvement of disaster characteristics is achieved, and the real-time requirements are met.

CN118967446BActive Publication Date: 2025-05-30CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411097421.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2025-05-30
Estimated Expiration
2044-08-12

AI Technical Summary

Technical Problem

The limitations of the hardware facilities for the acquisition of existing satellite remote sensing images lead to insufficient resolution and difficulty in capturing detailed disaster features. In addition, deep learning models require large computing resources and storage space, making it difficult to deploy and operate in resource-constrained environments.

Method used

The short-term delay super-segment method of remote sensing images of natural disaster satellites based on spatial shift convolution is adopted. The spatial shift operation is realized through a 1×1 convolution layer, local feature aggregation, and the time delay super-segment model is trained to generate high-resolution images.

Benefits of technology

The resolution of satellite remote sensing images is improved, the degree of significance of disaster characteristics is increased, and the model structure design is lightweight to meet real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967446B_ABST
    Figure CN118967446B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of satellite remote sensing image processing, and particularly to a short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution. The method includes: obtaining a low-resolution satellite remote sensing image with a resolution lower than a preset value; implementing a spatial shift operation through a 1×1 convolution layer to complete local feature aggregation; using paired low-resolution and high-resolution satellite remote sensing images to train a time-delay super-resolution model; and processing the low-resolution remote sensing image with the trained time-delay super-resolution model to complete super-resolution conversion and generate a high-resolution image. Compared with the super-resolution methods based on bilinear interpolation, cubic interpolation, and nearest neighbor interpolation, the PSNR value and SSIM value generated by the present invention are both higher, and the resolution and clarity of the image can be more effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite remote sensing image processing, and particularly to a short-time lag super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution. Background Art

[0002] The monitoring of natural disasters based on satellite remote sensing images can monitor and respond to various disasters more comprehensively, timely and efficiently. Due to the limitations of the hardware facilities for obtaining existing satellite remote sensing images, such as the fidelity of imaging and the bandwidth limitation of data transmission, the resolution of satellite remote sensing images is insufficient, making it difficult to capture detailed disaster features, such as ground cracks, hurricanes, etc., thus restricting the detection performance of disaster models. In real-time application scenarios such as satellite remote sensing image disaster detection, the requirement of short-time lag is relatively important. However, deep learning models usually require a large amount of computing resources and storage space, and larger models may lead to slower inference speed and even be difficult to deploy and run in resource-constrained environments. To address the above problems, the present invention attempts to design a dedicated super-resolution model for satellite remote sensing images in order to improve the resolution of satellite remote sensing images and increase the significance of disaster features. Summary of the Invention

[0003] The present invention discloses a short-time lag super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution. The specific method is as follows:

[0004] Obtain a low-resolution satellite remote sensing image with a resolution lower than a preset value, and preprocess the low-resolution satellite remote sensing image;

[0005] Implement spatial shift operation with a 1×1 convolutional layer to process the low-resolution satellite remote sensing image and complete local feature aggregation;

[0006] Use paired low-resolution and high-resolution satellite remote sensing images to train a time-lag super-resolution model;

[0007] Use the trained time-lag super-resolution model to process the low-resolution remote sensing image to complete super-resolution conversion and generate a high-resolution image.

[0008] Furthermore, preprocessing the low-resolution satellite remote sensing image includes normalization processing, and the specific formula is:

[0009]

[0010] where x is the original pixel value, x min and x max are the minimum and maximum pixel values respectively, and x normalized is the normalized pixel value.

[0011] Furthermore, the local feature aggregation is specifically as follows:

[0012] Given the input feature f of the remote sensing image, first evenly divide it into n groups of feature maps along the channel dimension, where n = S, and obtain n thinner tensors f i , denoted as

[0013]

[0014] where f i represents the i-th thinner tensor, H is the height, W is the width, and C latent represents the channel dimension of the latent features, and n represents the number of groups after segmentation;

[0015] Each separated feature group is displaced according to the given step parameter to obtain the displaced feature f shift , f shift Each pixel feature in contains local features along the channel dimension and is filled with a constant 0 by default.

[0016] Furthermore, the displacement is performed according to the given step parameter, and the specific method is as follows:

[0017] Calculate the number of shift groups. According to the number of channels C of the input feature map and the number of shift steps steps, calculate the number of shift groups shift_groups, and the number of shift groups is equal to the number of channels divided by the number of shift step quantities;

[0018] Calculate the number of channels in each group. The number of channels in each group is group_dim, that is, the total number of channels C divided by the number of shift groups shift_groups;

[0019] Pad the input feature map. Use the F.pad() function to pad the input feature map f, and the padding size is pad;

[0020] Initialize the output feature map. Initialize a zero tensor output with the same shape as the input feature map f

[0021] Circular shift operation. For each shift step in steps, obtain the corresponding index idx, row shift amount s_h, and column shift amount s_w of the current shift step; divide the padded feature map f_pad into shift_groups sub-blocks, and the number of channels in each sub-block is group_dim; perform shift and clipping on each sub-block;

[0022] Return the output feature map. Return the shifted feature map output;

[0023] Among them, the displacement is to move the upper left corner of the sub-block to the lower right corner of the output feature map, and the specific position is determined by (idx + 1) * group_dim;

[0024] Crop it into cropped sub - blocks to make it adapt to the position of the output feature map. The cropping size is (pad + s_h, pad + s_H) and (pad + s_w, pad + W) to ensure that the size of the output feature map remains unchanged.

[0025] Furthermore, the time - lag super - resolution model is constructed as follows:

[0026] Use the output of the 1×1 convolutional layer as the input layer of the time - lag super - resolution model to map the low - resolution satellite remote - sensing image to the latent high - dimensional space. The shallow feature f in the latent high - dimensional space head is described as

[0027]

[0028] where X is the input low - resolution satellite remote - sensing image, C latent represents the channel dimension of the latent feature, and N head represents the input layer of the time - lag super - resolution model;

[0029] The backbone part N of the time - lag super - resolution model backbone is stacked by SC - ResBlocks. N backbone takes the shallow feature f head as the input and outputs the deep features of the low - resolution satellite remote - sensing image extracted, which can be expressed as

[0030] f backbone = N backbone (f head )

[0031] Based on the deep feature f backbone , use the up - sampling module to reconstruct the high - resolution satellite remote - sensing image Z;

[0032] The up - sampling module N rec includes an SC layer, ReLU, 1×1 convolution, and pixel - shuffling operations;

[0033] Use 1×1 convolution to map the multi - channel features obtained by up - sampling to a 3 - channel high - resolution satellite remote - sensing image Z, and introduce bilinear interpolation to assist the super - resolution process. The whole process can be expressed as

[0034] Z = N rec (f backbone ) + Bilinear(X)

[0035] where X is the low - resolution satellite remote - sensing image, Z is the high - resolution satellite remote - sensing image, f backbone is the deep feature of the low - resolution satellite remote - sensing image, N rec is the up - sampling module, and Bilinear is bilinear interpolation.

[0036] Furthermore, the training of the time-delay super-resolution model is carried out as follows:

[0037] Divide the original dataset into a training set, a validation set, and a test set

[0038] Compress the resolution of each image, and then crop it together with the original Figure 1 into 256×256-sized fragments and store them in pairs in the final training data set;

[0039] Set the initial parameters for the training model. In the training configuration, the maximum number of iterations is set to 50, the batch size is set to 32, and the Adam optimizer is used to adjust the model parameters, with the learning rate set to 0.001;

[0040] According to the convergence effect of the mean squared error loss function.

[0041] Furthermore, the method also includes testing and validation. The accuracy and generalization of the model are evaluated by testing the performance of the model in actual application scenarios. The specific method is as follows:

[0042] Use the peak signal-to-noise ratio and the structural similarity index measure as indicators to comprehensively evaluate the performance of the time-delay super-resolution model;

[0043] The formula for calculating the peak signal-to-noise ratio is as follows:

[0044]

[0045] where MAX I is the maximum pixel value of the image;

[0046] The formula for calculating the structural similarity index measure is as follows:

[0047] SSIM(x,y) = [lx,y)] α ·[c(x,y)] β ·[s(x,y)] γ

[0048]

[0049] where x and y represent the output image of the super-resolution model and the real satellite remote sensing super-resolution image respectively, l, c, and s represent the similarities of the three components of brightness, contrast, and structure respectively, and α, β, and γ are weighting factors, usually taking the value of 1 to ensure equal contribution of each part, but can also be adjusted according to actual applications. μ x and μ y are the average brightness of images x and y respectively, σ x and σ y are the standard deviations of images x and y respectively, σ xyis the covariance of images x and y, C 1 , C 2 , C 3 are constants.

[0050] Furthermore, after generating the high-resolution image, inverse normalization processing is performed to inverse normalize each pixel value back to its original range. The specific formula is as follows:

[0051] x original = x normalized *(x max - x min ) + x min

[0052] where x original is the original value, x normalized is the value after normalization, x max is the maximum value in the original values, and x min is the minimum value in the original values.

[0053] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:

[0054] 1. From the visual effect, the present invention can better retain the details and textures of the original image during the interpolation process. Compared with traditional interpolation methods, the method of the present invention has significantly improved clarity and detail performance in the interpolated image, and has a good performance in the super-resolution of satellite remote sensing images, and can effectively improve the resolution and clarity of the image.

[0055] 2. The structural design of the model of the present invention is based on lightweight modules, which can meet the real-time requirements of the super-resolution model in system applications.

[0056] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The brief description of the drawings of the present invention is as follows.

[0058] Figure 1 is a schematic diagram of the architecture of the short-time lag super-resolution algorithm for satellite remote sensing images;

[0059] Figure 2 is a schematic diagram of the structural comparison between the traditional 3×3 convolutional residual module and the 1×1 shifted convolutional residual module;

[0060] Figure 3 is a schematic diagram of the overall architecture of the short-time lag super-resolution model for satellite remote sensing images. Detailed implementation mode

[0061] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0062] A short-time lag super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution, and the schematic diagram of the algorithm architecture is as Figure 1 shown. Given a low-resolution satellite remote sensing image X, it is represented as a matrix where H and W are the height and width of the image respectively, and C is the number of channels. The lightweight satellite remote sensing image super-resolution model F sr learns a mapping function to generate a high-resolution image Z. F sr is constructed by using a lightweight deep convolutional neural network. The satellite remote sensing image super-resolution process can be expressed as:

[0063] Z = F sr (X; θ sr )

[0064] where θ sr is the weight parameter of the lightweight satellite remote sensing image super-resolution model. The mean square error MSE loss is used to measure the reconstruction error to optimize the weight parameter θ sr , and this process can be expressed as:

[0065]

[0066] where, is the true high-resolution image, H represents the height of the image, W represents the width of the image, and Z ij represent the pixel values of the true image and the output image of the super-resolution model at the position (i, j) respectively. The network parameters are updated through the backpropagation algorithm to minimize the loss function L MSE (θ sr ).

[0067] After passing through the lightweight super-resolution model, a high-resolution image is obtained, where H' and W' are the height and width of the image Z respectively, and C' is the number of channels of the image.

[0068] The specific steps are as follows:

[0069] S1. Obtain a low-resolution satellite remote sensing image with a resolution lower than a preset value, and preprocess the low-resolution satellite remote sensing image.

[0070] There are mainly the following ways to obtain satellite remote sensing images: government and research institutions, such as the United States Geological Survey (USGS) providing Landsat series data, the European Space Agency (ESA) providing Sentinel series data, and the China National Space Administration's Center for Resources Satellite Data and Application providing high-resolution series satellite data; online map services, such as Google Earth and Microsoft Bing Maps; open-source data platforms, such as Copernicus Open Access Hub and Sentinel Hub, providing free access to Sentinel satellite data; and data markets and trading platforms, such as Amazon Web Services Marketplace and GBDX, etc. Select the appropriate way according to the needs to obtain the images.

[0071] Read the satellite remote sensing image data and apply the normalization formula to map the values between 0 and 1.

[0072] The normalization formula is:

[0073]

[0074] where x is the original pixel value, x min and x max are the minimum and maximum pixel values respectively, and x normalized is the normalized pixel value.

[0075] S2. Implement spatial shift operations with 1×1 convolutional layers to process low-resolution satellite remote sensing images and complete local feature aggregation.

[0076] In step S2, on the basis of traditional convolutional operations, parameter-free spatial shift operations are introduced to balance the local feature aggregation ability and lightweight requirements of convolutional neural networks. The structural comparison between the 3×3 convolutional residual module and the 1×1 shifted convolutional residual module is as Figure 2 shown. Specifically, first, based on the 1×1 convolutional layer, spatial shift operations are designed for local feature aggregation. On this basis, inspired by the structure of residual neural networks, a shifted convolutional residual module is constructed as the backbone network of the lightweight satellite remote sensing image super-resolution model.

[0077] 1×1 convolution is also called pointwise convolution. The main difference from conventional convolution is that the size of its convolution kernel is 1×1, that is, the convolution kernel matrix has only one element, indicating that the computational complexity of this type of convolution is much lower than that of conventional convolution. At the same time, the nonlinear transformation ReLU is introduced to increase the deep feature extraction ability of the convolutional network.

[0078] Based on the 1×1 convolution, local features are explicitly aggregated through spatial shifting. To facilitate the formal description of this process, first, the displacement direction is represented as d ∈ {1, 0, -1}, where 1 represents moving upward or rightward, 0 represents no movement, and -1 represents moving downward or leftward. Additionally, use d h and d w to represent the vertical and horizontal displacement directions respectively. Correspondingly, the displacement step sizes are represented as s h and s w . On this basis, the displacement direction and step size form a spatial displacement operation step step = (d h × s h , d w × s w ). The set composed of spatial displacement steps is represented as

[0079] S = {step i , i = 1, …, n}

[0080] where n is the number of aggregated features, and step i represents the operation step of the i-th local pixel feature.

[0081] To achieve the aggregation of the features of 8 local pixels similar to the standard 3×3 convolution, the set of spatial displacement operation steps can be defined as (0, 1), (0, -1), (1, 0), (1, 1), (1, -1), (-1, 0), (-1, 1), (-1, -1). The steps of the spatial shifting operation are described in detail as follows:

[0082] S21. Declare parameters, f: the input feature map, of type torch.Tensor, with shape (B, C, H, W), representing B batches, C channels, height H, and width W. steps: a list of tuples representing the spatial displacement step sizes in each direction (row and column). pad: the padding size used to ensure that the output size is the same as the input.

[0083] S22. Calculate the number of shifting groups. According to the number of channels C of the input feature map and the number of displacement step sizes steps, calculate the number of shifting groups shift_groups, where the number of shifting groups is equal to the number of channels divided by the number of displacement step sizes.

[0084] S23. Calculate the number of channels in each group. The number of channels in each group is group_dim, that is, the total number of channels C divided by the number of shifting groups shift_groups.

[0085] S24. Pad the input feature map. Use the F.pad() function to pad the input feature map f with the padding size pad.

[0086] S25. Initialize the output feature map, and initialize a zero tensor output with the same shape as the input feature map f.

[0087] S26. For each shift step size in steps, where the step size is in tuple form representing the row and column shift amounts, perform the following operations:

[0088] S261. Obtain the corresponding index idx, row shift amount s_h, and column shift amount s_w for the current shift step size.

[0089] S262. Split the padded feature map f_pad into shift_groups sub-blocks, each with a channel number of group_dim.

[0090] S263. Shift: Move the upper left corner of the sub-block to the lower right corner of the output feature map, and the specific position is determined by (idx + 1) * group_dim;

[0091] Clip: Clip the sub-block to fit its position in the output feature map, with the clipping sizes being (pad + s_h, pad + s_H) and (pad + s_w, pad + W) to ensure the size of the output feature map remains unchanged.

[0092] S27. Return the output feature map, and return the shifted feature map output.

[0093] In this embodiment, given the satellite remote sensing image input feature f, it is first evenly divided into n groups of feature maps along the channel dimension, where n = S, and n thinner tensors f are obtained i , which can be expressed as

[0094]

[0095] Then, each separated feature group is displaced according to the given step size parameter to obtain the displaced feature f shift . Each pixel feature in f shift contains local features along the channel dimension and is filled with a constant 0 by default.

[0096] The shifted convolutional residual module consists of a 1×1 convolutional layer and a spatial displacement operation to form a shifted convolutional layer (SC). Therefore, the shifted convolutional layer extends the ordinary 1×1 convolution, and with fewer parameters and computational amounts, it provides the ability of local feature aggregation. On this basis, a shifted convolutional residual module (SC-ResBlock) is designed by combining the shifted convolutional layer with the residual neural network structure. SC-ResBlock contains an SC layer, ReLU, and a 1×1 convolutional layer. Compared with the residual module based on 3×3 convolution, SC-ResBlock only uses 1×1 convolution, greatly reducing the number of parameters and computational amounts.

[0097] S3. Pair the satellite remote sensing images with low resolution and high resolution, and train the time-delay super-resolution model.

[0098] S31. The construction process of the short-time-delay super-resolution network for satellite remote sensing images is as follows:

[0099] First, use the standard 1×1 convolution as the input layer of the satellite remote sensing image super-resolution network to map the satellite remote sensing image into a potential high-dimensional space. The shallow features f in the potential high-dimensional space head can be described as

[0100]

[0101] where X is the input satellite remote sensing image, and C latent represents the channel dimension of the potential features, and N head represents the input layer of the satellite remote sensing image super-resolution network.

[0102] Secondly, the backbone part N of the satellite remote sensing image super-resolution network backone is stacked by SC-ResBlocks. N backone uses the shallow features f head as the input and outputs the extracted deep features of the satellite remote sensing image, which can be expressed as

[0103] f backbone = N backbone (f head )

[0104] Then, based on the deep features f backbone , use the upsampling module to reconstruct the high-resolution satellite remote sensing image Z. The upsampling module N rec is composed of an SC layer, ReLU, 1×1 convolution, and the PixelShuffle operation. Finally, while using the 1×1 convolution to map the multi-channel features obtained by upsampling into a 3-channel high-resolution satellite remote sensing image Z, the Bilinear interpolation is also introduced to assist the super-resolution process. The whole process can be expressed as

[0105] Z = N rec (f backbone ) + Bilinear(X)

[0106] The overall architecture of the constructed short-time-delay super-resolution model for satellite remote sensing images is as shown in Figure 3 shown.

[0107] Train the super-resolution model. Use a large number of paired satellite remote sensing images of low resolution (LR) and high resolution (HR) as training data. The goal of the training process is to minimize the difference between the super-resolution (SR) images generated by the model and the true high-resolution images. The mean square error is used as the loss function to measure this difference, and the model parameters are adjusted through the backpropagation algorithm to optimize the performance.

[0108] In this embodiment, based on the xBD, the largest natural disaster image dataset released by MIT, a satellite remote sensing image disaster monitoring dataset is constructed to meet the verification requirements of the method. The constructed satellite remote sensing image disaster monitoring dataset contains a total of 7,464 satellite remote sensing images, including 7 different natural disasters and normal satellite remote sensing images. For model development, verification, and testing, the original dataset is divided into a training set, a validation set, and a test set. Among them, the number of image samples in the training set is 4,478 for model training, and the number of image samples in the validation set is 1,120 for adjusting the hyperparameters of the model and evaluating the performance of the model to avoid overfitting or underfitting. First, compress the resolution of each picture, and then Figure 1 crop it into fragments of 256×256 size together and store them pairwise in the final training data set. Set the initial parameters for training the model. In the training configuration, the maximum number of iterations is set to 50, the batch size is set to 32, and the Adam optimizer is used to adjust the model parameters. The learning rate is set to 0.001 to achieve efficient and stable gradient descent. According to the convergence effect of the mean square error loss function, only select the trained model for testing experiments, and select the 45th generation network model for experiments.

[0109] S32. After the model training is completed, use the test dataset to evaluate the generalization ability of the model. The peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are selected as the quality metrics for calculating the super-resolution images.

[0110] The number of image samples in the test set is 1,866 for evaluating the generalization ability of the model on unknown data, that is, to test the performance of the model in actual application scenarios through the test set and evaluate the accuracy and generalization of the model. And the peak signal-to-noise ratio (PSNR) and the structural similarity index metric (SSIM) are used as indicators to comprehensively evaluate the performance of the super-resolution model.

[0111] The formula definition of PSNR is:

[0112]

[0113] where MAX I is the maximum pixel value of the image. The larger the PSNR, the better the effect of the super-resolution model.

[0114] The structural similarity index measure SSIM takes into account the brightness, contrast of the image, and the preservation of the image structure information. Its basic formula can be expressed as:

[0115] SSIM(x,y) = [l(x,y)] α ·[x(x,y)] β ·[s(x,y)] γ

[0116]

[0117] In this embodiment, the average PSNR index is 37.111 dB, the standard deviation is 2.417, the average SSIM index is 0.936, and the standard deviation is 0.024. In order to confirm that the model algorithm is more superior, at the same time, simulation experiments were carried out on the other three interpolation methods: the bilinear interpolation BILINEAR super-resolution method, the bicubic interpolation BICUBIC super-resolution method, and the nearest neighbor interpolation NEAREST super-resolution method, and the objective evaluation index results were recorded, as shown in Table 1. The average PSNR values based on the BICUBIC, BILINEAR, and NEAREST methods are 33.792, 33.491, and 33.310 respectively, and the average SSIM values are 0.746, 0.722, and 0.700 respectively. In contrast, the PSNR value and SSIM value generated by this embodiment are superior to other super-resolution algorithms. The above experimental result analysis shows that the method designed by the present invention can not only preserve the image details and textures, but also can more effectively improve the resolution and clarity of the image.

[0118] Table 1 Statistical results of satellite remote sensing image super-resolution comparison experiments

[0119]

[0120] S4. Process the low-resolution remote sensing image with the trained time-delay super-resolution model to complete the super-resolution conversion and generate a high-resolution image. Use the trained model to perform super-resolution conversion on the new low-resolution satellite remote sensing image to generate a high-resolution image, and perform denormalization processing on the generated super-resolution image.

[0121] Denormalization formula:

[0122] x original = x normalized *(x max - x min ) + x min

[0123] Denormalize each pixel value back to its original range.

[0124] S5. Result feedback and model iteration. The super-resolution image is used for disaster monitoring and assessment, feedback information is collected, and the model is adjusted and optimized as necessary.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution, characterized in that: The specific method is as follows: Acquire low-resolution satellite remote sensing images with a resolution lower than a preset value, and pre-process the low-resolution satellite remote sensing images; The 1×1 convolutional layer is used to implement spatial shift operations, process low-resolution satellite remote sensing images, and complete local feature aggregation; Pair low-resolution and high-resolution satellite remote sensing images to train a time-delay super-resolution model; Use the trained time-delay super-resolution model to process low-resolution remote sensing images, complete super-resolution conversion, and generate high-resolution images; The time-delay super-resolution model is constructed as follows: The output of the 1×1 convolutional layer is used as the input layer of the time-delay super-resolution model to map the low-resolution satellite remote sensing image to a latent high-dimensional space. The shallow features in the latent high-dimensional space Described as in, is the input low-resolution satellite remote sensing image, represents the channel dimension of the latent features, Represents the input layer of the time-delay super-resolution model; The backbone of the time-delay superresolution model Made of SC-ResBlocks stacked together, Shallow features As input, the deep features of the low-resolution satellite remote sensing image extracted are output, which can be expressed as Based on deep features , using the upsampling module to reconstruct high-resolution satellite remote sensing images ; Upsampling module Includes SC layer, ReLU, 1×1 convolution and pixel shuffle operations; Use 1×1 convolution to map the upsampled multi-channel features into a 3-channel high-resolution satellite remote sensing image. , bilinear interpolation is introduced to assist the super-resolution process, and the whole process can be expressed as in, is a low-resolution satellite remote sensing image, Z is a high-resolution satellite remote sensing image, is the deep features of low-resolution satellite remote sensing images. is the upsampling module, is bilinear interpolation.

2. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 1, characterized in that: Preprocess the low-resolution satellite remote sensing images, including normalization. The specific formula is: in is the original pixel value, and are the minimum and maximum pixel values, respectively, and is the normalized pixel value.

3. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 1, characterized in that: Local feature aggregation, the specific method is as follows: Given remote sensing image input features , first divide it evenly along the channel dimension Group feature map, where , and obtain n thinner tensors , expressed as in, represents the i-th thinner tensor, H is the height, W is the width, represents the channel dimension of the latent features, Indicates the number of groups after segmentation; Each separated feature group is shifted according to the given step size parameter to obtain the shifted features , Each pixel feature in contains local features along the channel dimension and is padded with a constant 0 by default.

4. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 3, characterized in that: The displacement is performed according to the given step size parameter. The specific method is as follows: Calculate the number of shift groups. According to the number of channels C of the input feature map and the number of shift steps, calculate the number of shift groups shift_groups. The number of shift groups is equal to the number of channels divided by the number of shift steps. Calculate the number of channels in each group. The number of channels in each group is group_dim, which is the total number of channels C divided by the number of shift groups shift_groups; Fill the input feature map, use the F.pad() function to fill the input feature map f, and the filling size is pad; Initialize the output feature map, initialize a zero tensor output with the same shape as the input feature map f Circular shift operation, for each shift step in steps, get the index idx, row shift amount s_h and column shift amount s_w corresponding to the current shift step; divide the padded feature map f_pad into shift_groups sub-blocks, and the number of channels of each sub-block is group_dim; shift and crop each sub-block; Return the output feature map, return the shifted feature map output; The displacement is to move the upper left corner of the sub-block to the lower right corner of the output feature map, and the specific position is determined by (idx+1)*group_dim; Cropping is to crop the sub-block to adapt it to the position of the output feature map, and the cropping sizes are (pad+s_h, pad+s_H) and (pad+s_w, pad+W), ensuring that the size of the output feature map remains unchanged.

5. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 1, characterized in that: The specific method for training the time-delay super-resolution model is as follows: Divide the original dataset into training set, validation set and test set Each image is compressed and then cropped into 256×256 size fragments together with the original image, and stored in pairs in the final training data set; Initialize the parameters of the training model. In the training configuration, the maximum number of iterations is set to 50, the batch size is set to 32, and the Adam optimizer is used to adjust the model parameters. The learning rate is set to 0.

001. According to the convergence effect of the mean square error loss function.

6. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 5, characterized in that: The method also includes testing and verification, which uses a test set to test the performance of the model in actual application scenarios to evaluate the accuracy and generalization of the model. The specific method is as follows: Peak signal-to-noise ratio and structural similarity index are used as indicators to comprehensively evaluate the performance of the time-delay super-resolution model; The peak signal-to-noise ratio calculation formula is as follows: PSNR in, is the maximum pixel value of the image; The structural similarity index metric calculation formula is as follows: in, and Represent the super-resolution model output image and the real satellite remote sensing super-resolution image, respectively. , , Represent the similarity of brightness, contrast and structure respectively, and , , is a weighting factor, which is usually taken as 1 to ensure equal contribution of all parts, but can be adjusted according to actual application; and The images are and The average brightness, and is an image and The standard deviation of is an image and The covariance of , , is a constant.

7. The short-time-delay super-resolution method for natural disaster satellite remote sensing images based on spatial shift convolution as claimed in claim 1, characterized in that: After generating the high-resolution image, denormalization is performed to normalize each pixel value back to its original range. The specific formula is as follows: in, is the original value, is the normalized value, is the maximum value among the original values, is the minimum of the original values.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method and system based on multi-scale enhancement module

    CN116342389A

  • Satellite cloud picture super-resolution reconstruction method based on multipath aggregation Transformer

    CN117391958A