Multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing image
Through the multi-scale fusion training strategy and the multi-scale super-resolution reconstruction method of Mars remote sensing images with a three-level pyramid structure, the problems of poor large-scale super-resolution reconstruction effect and high computing complexity of Mars remote sensing images in traditional methods are solved, and the ultra-lightweight deployment and high-quality image reconstruction of the network are realized.
Patent Information
- Application Number
- CN202510626387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The current traditional super-resolution reconstruction method based on deep learning has poor large-scale super-resolution reconstruction of Mars remote sensing images, and the algorithm parameters are large and the calculation complexity is high, making it difficult to deploy on embedded devices.
A multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images is proposed. Through multi-scale fusion training strategy and a three-level pyramid structure, a 2-, 4-, and 8-fold super-resolution reconstruction branch is designed, and the parameter sharing and recursive structures are used to reduce the amount of parameters, and the residual learning is focused through global jump connections and non-homologous jump connections.
The network is ultra-lightweight, with only 103K parameters, which can be deployed on embedded devices, improve the network convergence speed, reduce training time, and effectively extract high-frequency details of Mars remote sensing images, generating high-quality super-resolution images.
Smart Images

Figure CN120147141A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Mars remote sensing image processing, and in particular relates to a multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images. Background Art
[0002] At present, deep space exploration has become a strategic commanding height for the world's major space powers to conduct scientific and technological exploration and innovation. Among them, Mars has become a hot spot and growth point for countries to compete in the field of deep space exploration because it is the main target celestial body for extraterrestrial life exploration. The primary task of Mars exploration is to obtain global high-resolution optical images. In addition, the selection of landing sites in future Mars landing missions, Mars rover exploration path planning, and manned Mars programs also require the acquisition of high-resolution images of the Martian surface.
[0003] Affected by imaging conditions, the imaging system usually cannot restore all the information of the original scene. The imaging process will be affected by deformation (translation, rotation, etc.), blur (motion blur, optical blur, defocus blur, etc.), downsampling, noise, etc., resulting in degradation of image quality. Therefore, how to effectively improve the quality of images has become a research hotspot. Based on the number of low-resolution images input during the reconstruction process, super-resolution reconstruction methods can be roughly divided into two categories: multi-frame image super-resolution reconstruction methods and single-frame image super-resolution reconstruction methods (also known as learning-based methods).
[0004] The multi-frame image super-resolution reconstruction method can utilize the information of the sequence of low-resolution images, but it has limitations such as non-unique solutions, strong dependence on initial values, large amount of calculation, slow convergence speed, stability needs to be improved, and easy edge blurring; the single-frame image super-resolution reconstruction method is based on learning to obtain the mapping relationship between high and low resolution images. With the development of deep learning technology, it is divided into shallow learning and deep learning methods. The shallow learning-based method has problems such as high computational complexity, affected reconstruction quality, limited model feature extraction and expression capabilities; although the deep learning-based method has become the mainstream and has advantages in some fields, it has poor reconstruction effect on large-scale super-resolution of Mars remote sensing images with less detailed information, and there are problems such as large number of algorithm parameters and high computational complexity. At the same time, there is little research on super-resolution reconstruction technology in the field of Mars remote sensing image processing, and publicly published literature and related patents are scarce.
[0005] Therefore, based on the current research results, deep learning-based methods have shown obvious advantages in super-resolution reconstruction of faces, texts, natural images, and medical images. However, since Mars remote sensing images contain less detailed information than Earth remote sensing images or natural images, the current traditional deep learning-based super-resolution reconstruction methods have unsatisfactory results for large-scale super-resolution reconstruction of Mars remote sensing images, and the algorithm has a large number of parameters and high computational complexity. Summary of the invention
[0006] In view of this, the present invention aims to provide a multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images, so as to solve the problems that, compared with Earth remote sensing images or natural images, Mars remote sensing images contain less detailed information, and the current traditional deep learning-based super-resolution reconstruction methods have unsatisfactory effects in large-scale super-resolution reconstruction of Mars remote sensing images, and have a large number of algorithm parameters and high computational complexity. Compared with existing networks with 8-fold super-resolution reconstruction capabilities, which often have parameter quantities of 2000K - 6000K or even larger, the parameter quantity of the present invention is only 103K, achieving ultra-lightweight of the network and can be deployed on embedded devices.
[0007] To achieve the above object, the technical solution of the present invention is realized as follows: A multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images specifically includes the following steps: S1: Obtain a pair of high-resolution image and low-resolution image; S2: Input the pair of high-resolution image and low-resolution image into a multi-scale super-resolution reconstruction network, and use a multi-scale fusion training strategy to train the multi-scale super-resolution reconstruction network to obtain a multi-scale super-resolution reconstruction model; The multi-scale super-resolution reconstruction network includes a 2-fold super-resolution reconstruction branch, a 4-fold super-resolution reconstruction branch, and an 8-fold super-resolution reconstruction branch; S3: Input the image to be reconstructed into the multi-scale super-resolution reconstruction model for processing to obtain a super-resolution reconstructed image.
[0008] Furthermore, in the model training stage, when the low-resolution image input into the multi-scale super-resolution reconstruction network is a 2-fold downsampled image, the 2-fold super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network processes the input image and outputs a 2-fold super-resolution reconstructed image; when the low-resolution image input into the multi-scale super-resolution reconstruction network is a 4-fold downsampled image, the 2-fold super-resolution reconstruction branch and the 4-fold super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network process the input image and output a 2-fold super-resolution reconstructed image and a 4-fold super-resolution reconstructed image; when the low-resolution image input into the multi-scale super-resolution reconstruction network is an 8-fold downsampled image, the 2-fold super-resolution reconstruction branch, the 4-fold super-resolution reconstruction branch, and the 8-fold super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network process the input image and output a 2-fold super-resolution reconstructed image, a 4-fold super-resolution reconstructed image, and an 8-fold super-resolution reconstructed image.
[0009] Furthermore, the 2-fold super-resolution reconstruction branch includes a first residual image extraction module; the 4-fold super-resolution reconstruction branch includes a second residual image extraction module; and the 8-fold super-resolution reconstruction branch includes a third residual image extraction module.
[0010] Furthermore, the low-resolution image is upsampled to obtain a 2-fold upsampled image, and the low-resolution image is input into the first residual image extraction module for processing to obtain a 2-fold residual image. The 2-fold upsampled image and the 2-fold residual image are added pixel by pixel to obtain a 2-fold super-resolution reconstructed image. The 2-fold residual image is input into the second residual image extraction module for processing to obtain a 4-fold residual image. The 2-fold super-resolution reconstructed image is upsampled to obtain a 4-fold upsampled image. The 4-fold upsampled image and the 4-fold residual image are added pixel by pixel to obtain a 4-fold super-resolution reconstructed image. The 4-fold residual image is input into the third residual image extraction module for processing to obtain an 8-fold residual image. The 4-fold super-resolution reconstructed image is upsampled to obtain an 8-fold upsampled image. The 8-fold upsampled image and the 8-fold residual image are added pixel by pixel to obtain an 8-fold super-resolution reconstructed image.
[0011] Furthermore, the first residual image extraction module, the second residual image extraction module, and the third residual image extraction module share parameters.
[0012] Furthermore, the first residual image extraction module, the second residual image extraction module, and the third residual image extraction module have the same structure. The first residual image extraction module includes a first 3×3 convolution, a first residual recursion module, a second residual recursion module, a third residual recursion module, a fourth residual recursion module, a fifth residual recursion module, a sixth residual recursion module, a second 3×3 convolution, a first 1×1 convolution, a third 3×3 convolution, a fourth 3×3 convolution, and a pixel recombination module. The low-resolution image input into the first residual image extraction module is processed by the first 3×3 convolution to obtain a feature map A1. After the feature map A1 is processed by the first residual recursion module, a feature map A2 is obtained. The feature map A1 and the feature map A2 are subjected to a non-homologous skip connection to obtain a feature map A3, and the feature map A3 is input into the second residual recursion module. The processing methods of the second residual recursion module, the third residual recursion module, the fourth residual recursion module, the fifth residual recursion module, and the sixth residual recursion module for the image are the same as those of the first residual recursion module. Add the feature map input to the sixth residual recursive module and the feature map output after being processed by the sixth residual recursive module to obtain feature map B1. Feature map B1 is successively processed by a second 3×3 convolution, a first 1×1 convolution, and a third 3×3 convolution to obtain feature map B2. Perform a global skip connection on feature map B2 and feature map A1 to obtain feature map B3. After feature map B3 is successively processed by a fourth 3×3 convolution and a pixel recombination module, a 2× residual image is obtained.
[0013] Further, the network structures of the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module are the same. The first residual recursive module includes a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a window multi-head self-attention module, a first multi-layer perceptron, a second multi-layer perceptron, a shifted window multi-head self-attention module, a fifth 3×3 convolution, a second 1×1 convolution, and a sixth 3×3 convolution. Among them, the feature map C1 input to the first residual recursive module is processed by the first normalization layer and the window multi-head self-attention module to obtain feature map C2. After adding feature map C1 and feature map C2, feature map C3 is obtained. Feature map C3 is processed by the second normalization layer and the first multi-layer perceptron to obtain feature map C4. Add feature map C4 and feature map C3 to obtain feature map C5. Feature map C5 is processed by the third normalization layer and the shifted window multi-head self-attention module to obtain feature map C6. Add feature map C5 and feature map C6 to obtain feature map C7. Feature map C7 is processed by the fourth normalization layer and the second multi-layer perceptron to obtain feature map C8. Add feature map C8 and feature map C7 to obtain feature map C9. Feature map C9 is processed by the fifth 3×3 convolution, the second 1×1 convolution, and the sixth 3×3 convolution to obtain the output feature map of the first residual recursive module.
[0014] Further, the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module share parameters.
[0015] Further, perform 2× downsampling, 4× downsampling, and 8× downsampling on the training images respectively to obtain 2× downsampled images, 4× downsampled images, and 8× downsampled images; Input the 2× downsampled images into the corresponding multi-scale super-resolution reconstruction network for processing to obtain 2× super-resolution reconstructed images, and calculate the 2× reconstruction loss function based on the training images and the 2× super-resolution reconstructed images; Input the 4x downsampled image into the corresponding multi-scale super-resolution reconstruction network for processing to obtain 2x and 4x super-resolution reconstructed images. Calculate the 2x reconstruction loss function based on the training image and the 2x super-resolution reconstructed image, and calculate the 4x loss function based on the training image and the 4x super-resolution reconstructed image; Input the 8x downsampled image into the corresponding multi-scale super-resolution reconstruction network for processing to obtain 2x, 4x, and 8x super-resolution reconstructed images. Calculate the 2x reconstruction loss function based on the training image and the 2x super-resolution reconstructed image, calculate the 4x reconstruction loss function based on the training image and the 4x super-resolution reconstructed image, and calculate the 8x reconstruction loss function based on the training image and the 8x super-resolution reconstructed image; Add all the 2x reconstruction loss functions, all the 4x reconstruction loss functions, and the 8x reconstruction loss functions to obtain the multi-scale fusion loss function. Use the multi-scale fusion loss function to constrain the training of the multi-scale super-resolution reconstruction network to obtain the multi-scale super-resolution reconstruction model.
[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) For the multi-scale super-resolution reconstruction method of the ultra-lightweight Mars remote sensing image described in the present invention, compared with the existing networks with 8x super-resolution reconstruction ability that often have parameter quantities of 2000K - 6000K or even larger, the parameter quantity of the present invention is only 103K, realizing ultra-lightweight of the network and can be deployed on embedded devices.
[0017] (2) For the multi-scale super-resolution reconstruction method of the ultra-lightweight Mars remote sensing image described in the present invention, the proposed cross-network multi-scale fusion training strategy can accelerate the network convergence speed and reduce the training time.
[0018] (3) For the multi-scale super-resolution reconstruction method of the ultra-lightweight Mars remote sensing image described in the present invention, the present invention focuses on residual learning through global skip connections and non-homologous skip connections; enhances the feature extraction ability through the multi-head self-attention layer; and through the moving window, the self-attention mechanism is not only limited to non-overlapping local windows, but also can realize the application of cross-window attention mechanism, improving the utilization efficiency of window edge information. Through the above operations, the present invention can effectively extract the high-frequency detail information of the Mars remote sensing image, thereby realizing high-quality super-resolution image reconstruction. The present invention designs a three-level pyramid structure, and realizes the multi-scale super-resolution reconstruction of the Mars remote sensing image through parameter sharing and recursion, and can simultaneously generate 2x, 4x, and 8x super-resolution reconstructed images. Description of the Drawings
[0019] The accompanying drawings, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a schematic flowchart of the multi-scale super-resolution reconstruction method for the ultra-lightweight Mars remote sensing image described in the embodiment of the present invention; Figure 2 It is a schematic diagram of the network structure of the multi-scale super-resolution reconstruction model described in the embodiment of the present invention; Figure 3 It is a schematic diagram of the network structure of the first residual image extraction module described in the embodiment of the present invention; Figure 4 It is a schematic diagram of the network structure of the first residual recursion module described in the embodiment of the present invention; Figure 5 It is a schematic diagram of the principles of the window multi-head self-attention module and the shifted window multi-head self-attention module described in the embodiment of the present invention; Figure 6 It is a schematic diagram of the training process of the multi-scale super-resolution reconstruction network described in the embodiment of the present invention. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0021] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0022] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0023] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.
[0024] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0025] As Figure 1 shown, the multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images proposed by the present invention specifically includes the following steps: S1: Obtain a pair of high-resolution images and low-resolution images; S2: Input the pair of high-resolution images and low-resolution images into a multi-scale super-resolution reconstruction network, and use a multi-scale fusion training strategy to train the multi-scale super-resolution reconstruction network to obtain a multi-scale super-resolution reconstruction model; the multi-scale super-resolution reconstruction network includes a 2x super-resolution reconstruction branch, a 4x super-resolution reconstruction branch, and an 8x super-resolution reconstruction branch; S3: Input the image to be reconstructed into the multi-scale super-resolution reconstruction model for processing to obtain a super-resolution reconstructed image.
[0026] The present invention designs a three-level pyramid structure (i.e., a 2x super-resolution reconstruction branch, a 4x super-resolution reconstruction branch, and an 8x super-resolution reconstruction branch). Each level of the pyramid is composed of a residual image extraction module. Through parameter sharing of the three-level residual image extraction modules (realizing the parameter reduction to 1 / 3) and recursion, multi-scale super-resolution reconstruction of ultra-lightweight Mars remote sensing images can be achieved, and 2x, 4x, and 8x super-resolution reconstructed images can be generated simultaneously.
[0027] Furthermore, the low-resolution image for training is the original Mars remote sensing image, and the original Mars remote sensing image is the remote sensing image of the Mars surface taken by the High Resolution Imaging Science Experiment (HiRISE) camera carried by the Mars Reconnaissance Orbiter (MRO) of the United States and the High Resolution Camera (HiRIC) carried by the Tianwen-1 Mars probe of China. The obtained original Mars remote sensing image is used as the high-resolution image of the training data set, and the high-resolution image is downsampled to obtain an image as the low-resolution image of the training data set.
[0028] The purpose of super-resolution reconstruction is to obtain a high-resolution image. The input is the currently obtained real image (i.e., the low-resolution image), where the real image refers to the image obtained by taking a photo with a camera. The high-resolution image is obtained through a multi-scale super-resolution reconstruction model.
[0029] When training the multi-scale super-resolution reconstruction model, corresponding high-resolution and low-resolution image pairs are required, and the training process is constrained based on a loss function. Whether the trained multi-scale super-resolution reconstruction model meets the requirements needs to be evaluated using the corresponding high-resolution and low-resolution image pairs. Since there is no real image with a higher resolution, the currently obtained real image is used as the high-resolution image during the training phase, and the corresponding downsampled image is used as the low-resolution image. After training, using the trained multi-scale super-resolution reconstruction model, taking the currently obtained real image as the input image (low-resolution image) can obtain a high-resolution output image.
[0030] In some embodiments, during the model training phase, when the low-resolution image input to the multi-scale super-resolution reconstruction network is a 2-fold downsampled image, the 2-fold super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network processes the input image and outputs a 2-fold super-resolution reconstructed image; when the low-resolution image input to the multi-scale super-resolution reconstruction network is a 4-fold downsampled image, the 2-fold and 4-fold super-resolution reconstruction branches of the multi-scale super-resolution reconstruction network process the input image and output a 2-fold super-resolution reconstructed image and a 4-fold super-resolution reconstructed image; when the low-resolution image input to the multi-scale super-resolution reconstruction network is an 8-fold downsampled image, the 2-fold, 4-fold, and 8-fold super-resolution reconstruction branches of the multi-scale super-resolution reconstruction network process the input image and output a 2-fold super-resolution reconstructed image, a 4-fold super-resolution reconstructed image, and an 8-fold super-resolution reconstructed image.
[0031] In some embodiments, the 2-fold super-resolution reconstruction branch includes a first residual image extraction module; the 4-fold super-resolution reconstruction branch includes a second residual image extraction module; the 8-fold super-resolution reconstruction branch includes a third residual image extraction module.
[0032] In some embodiments, the low-resolution image is upsampled to obtain a 2-fold upsampled image, and the low-resolution image is input to the first residual image extraction module for processing to obtain a 2-fold residual image. The 2-fold upsampled image and the 2-fold residual image are added pixel by pixel to obtain a 2-fold super-resolution reconstructed image; The 2x residual image is input into the second residual image extraction module for processing to obtain a 4x residual image. The 2x super-resolution reconstruction image is upsampled to obtain a 4x upsampled image. The 4x upsampled image and the 4x residual image are added pixel by pixel to obtain a 4x super-resolution reconstruction image; The 4x residual image is input into the third residual image extraction module for processing to obtain an 8x residual image. The 4x super-resolution reconstruction image is upsampled to obtain an 8x upsampled image. The 8x upsampled image and the 8x residual image are added pixel by pixel to obtain an 8x super-resolution reconstruction image.
[0033] Further, as Figure 2 shown, the overall network architecture of the multi-scale super-resolution reconstruction network is a three-level pyramid structure. Each level of the pyramid consists of a residual image extraction module, and the parameters of the three-level residual image extraction modules are shared, achieving a reduction of the parameters to 1 / 3. The original low-resolution image generates a 2x residual image through the first residual image extraction module; the original low-resolution image is upsampled to generate a 2x upsampled image; the 2x residual image and the 2x upsampled image are added pixel by pixel to obtain a 2x super-resolution reconstruction image. The 2x residual image generates a 4x residual image through the second residual image extraction module; the 2x super-resolution reconstruction image is upsampled to generate a 4x upsampled image; the 4x residual image and the 4x upsampled image are added pixel by pixel to obtain a 4x super-resolution reconstruction image. The 4x residual image generates an 8x residual image through the third residual image extraction module; the 4x super-resolution reconstruction image is upsampled to generate an 8x upsampled image; the 8x residual image and the 8x upsampled image are added pixel by pixel to obtain an 8x super-resolution reconstruction image.
[0034] In some embodiments, the first residual image extraction module, the second residual image extraction module, and the third residual image extraction module share parameters.
[0035] In some embodiments, the first residual image extraction module, the second residual image extraction module, and the third residual image extraction module have the same structure. The first residual image extraction module includes a first 3x3 convolution, a first residual recursive module, a second residual recursive module, a third residual recursive module, a fourth residual recursive module, a fifth residual recursive module, a sixth residual recursive module, a second 3x3 convolution, a first 1x1 convolution, a third 3x3 convolution, a fourth 3x3 convolution, and a pixel recombination module. After the low-resolution image input into the first residual image extraction module is processed by the first 3x3 convolution, a feature map A1 is obtained. After the feature map A1 is processed by the first residual recursive module, a feature map A2 is obtained. The feature map A1 and the feature map A2 are subjected to a non-homologous skip connection to obtain a feature map A3, and the feature map A3 is input into the second residual recursive module; The processing methods of the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module for the image are the same as those of the first residual recursive module; Add the feature map input to the sixth residual recursive module and the feature map output after being processed by the sixth residual recursive module to obtain feature map B1. Feature map B1 is successively processed by a second 3×3 convolution, a first 1×1 convolution, and a third 3×3 convolution to obtain feature map B2. Perform a global skip connection between feature map B2 and feature map A1 to obtain feature map B3. After feature map B3 is successively processed by a fourth 3×3 convolution and a pixel recombination module, a 2-fold residual image is obtained.
[0036] As Figure 3 shown, the key to generating a high-quality super-resolution reconstructed image lies in the inference of the residual image, and the residual image should contain the high-frequency information implicit in the low-resolution image. In the present invention, the residual image extraction module is responsible for deeply extracting the input shallow features and inferring the high-frequency information. The residual image extraction module is composed of a 3×3 convolution, six residual recursive modules, three layers of bottleneck convolutions, a 3×3 convolution, and a pixel recombination module connected in series. The six residual recursive modules share parameters, reducing the parameters further to about 1 / 18. The six residual recursive modules focus on residual learning through global skip connections and non-homologous skip connections, reducing the learning difficulty of the residual image and improving the convergence efficiency of the network. The three layers of bottleneck convolutions are composed of a 3×3 convolution, a 1×1 convolution, and a 3×3 convolution connected in series.
[0037] Furthermore, the detailed processing methods of the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module for the image are as follows: Input feature map A3 to the second residual recursive module for processing to obtain feature map A4. Perform a non-homologous skip connection between feature map A3 and feature map A4 to obtain feature map A5. Input feature map A5 to the third residual recursive module for processing to obtain feature map A6. Perform a non-homologous skip connection between feature map A5 and feature map A6 to obtain feature map A7. Input feature map A7 to the fourth residual recursive module for processing to obtain feature map A8. Perform a non-homologous skip connection between feature map A7 and feature map A8 to obtain feature map A9. Input feature map A9 to the fifth residual recursive module for processing to obtain feature map A10. Perform a non-homologous skip connection between feature map A9 and feature map A10 to obtain feature map A11. Input feature map A11 to the sixth residual recursive module for processing to obtain feature map A12. Perform a non-homologous skip connection between feature map A11 and feature map A12 to obtain feature map B1.
[0038] In some embodiments, the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module have the same network structure. The first residual recursive module includes a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a window multi-head self-attention module, a first multi-layer perceptron, a second multi-layer perceptron, a shifted window multi-head self-attention module, a fifth 3×3 convolution, a second 1×1 convolution, and a sixth 3×3 convolution. Among them, after the feature map C1 input to the first residual recursive module is processed by the first normalization layer and the window multi-head self-attention module, the feature map C2 is obtained. After adding the feature map C1 and the feature map C2, the feature map C3 is obtained. The feature map C3 is processed by the second normalization layer and the first multi-layer perceptron to obtain the feature map C4. After adding the feature map C4 and the feature map C3, the feature map C5 is obtained. The feature map C5 is processed by the third normalization layer and the shifted window multi-head self-attention module to obtain the feature map C6. After adding the feature map C5 and the feature map C6, the feature map C7 is obtained. The feature map C7 is processed by the fourth normalization layer and the second multi-layer perceptron to obtain the feature map C8. After adding the feature map C8 and the feature map C7, the feature map C9 is obtained. The feature map C9 is processed by the fifth 3×3 convolution, the second 1×1 convolution, and the sixth 3×3 convolution to obtain the output feature map of the first residual recursive module.
[0039] Further, the window multi-head self-attention module and the shifted window multi-head self-attention module are from the paper "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" by Liu et al.
[0040] It should be noted that, as Figure 4 shown, each residual recursive module is sequentially composed of a first normalization layer, a window multi-head self-attention module, a second normalization layer, a first multi-layer perceptron, a third normalization layer, a shifted window multi-head self-attention module, a fourth normalization layer, a second multi-layer perceptron, and three layers of bottleneck convolution in series. At the same time, skip connections are designed inside the residual recursive module. The window multi-head self-attention module, as Figure 5 shown on the left, divides the input feature map into N×N individual windows and performs multi-head self-attention calculations on each window (containing M×M pixels). To ensure effective information exchange between different windows, a displacement operation is performed on the window multi-head self-attention module to form a shifted window multi-head self-attention module, as Figure 5As shown on the right, move M / 2 pixels to the left and up (or left and down / right and up / right and down) respectively, and perform the multi-head self-attention calculation again. The displacement window multi-head self-attention module realizes the application of the cross-window attention mechanism, improves the utilization efficiency of the window edge information, and thus realizes high-quality super-resolution image reconstruction.
[0041] In some embodiments, the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module share parameters.
[0042] Furthermore, without considering the number of parameters, the part consisting of a normalization layer, a window multi-head self-attention module, a normalization layer, a multi-layer perceptron, a normalization layer, a displacement window multi-head self-attention module, a normalization layer, and a multi-layer perceptron in the residual recursive module can be stacked multiple times.
[0043] In some embodiments, the training image (equivalent to the high-resolution image) is downsampled by 2 times, 4 times, and 8 times respectively to obtain a 2-times downsampled image, a 4-times downsampled image, and an 8-times downsampled image (equivalent to obtaining different low-resolution images); The 2-times downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2-times super-resolution reconstructed image, and the 2-times reconstruction loss function is calculated based on the training image and the 2-times super-resolution reconstructed image; The 4-times downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2-times super-resolution reconstructed image and a 4-times super-resolution reconstructed image, the 2-times reconstruction loss function is calculated based on the training image and the 2-times super-resolution reconstructed image, and the 4-times loss function is calculated based on the training image and the 4-times super-resolution reconstructed image; The 8-times downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2-times super-resolution reconstructed image, a 4-times super-resolution reconstructed image, and an 8-times super-resolution reconstructed image, the 2-times reconstruction loss function is calculated based on the training image and the 2-times super-resolution reconstructed image, the 4-times reconstruction loss function is calculated based on the training image and the 4-times super-resolution reconstructed image, and the 8-times reconstruction loss function is calculated based on the training image and the 8-times super-resolution reconstructed image; All the 2-times reconstruction loss functions, all the 4-times reconstruction loss functions, and the 8-times reconstruction loss functions are added together to obtain a multi-scale fusion loss function, and the multi-scale super-resolution reconstruction network training is constrained by using the multi-scale fusion loss function to obtain a multi-scale super-resolution reconstruction model.
[0044] It should be noted that in the case of streamlining, the above network architecture can obtain a super-resolution reconstruction network corresponding to a two-stage residual image extraction module and a one-stage residual image extraction module. The former can generate 2x and 4x super-resolution reconstruction images simultaneously, while the latter only generates 2x super-resolution reconstruction images.
[0045] Furthermore, the multi-scale super-resolution reconstruction network corresponding to the 2x downsampled image input is called the 2x super-resolution reconstruction network, the multi-scale super-resolution reconstruction network corresponding to the 4x downsampled image input is called the 4x super-resolution reconstruction network, and the multi-scale super-resolution reconstruction network corresponding to the 8x downsampled image input is called the 8x super-resolution reconstruction network.
[0046] As Figure 6 shown, the training image is downsampled by 2x to obtain a 2x downsampled image. The 2x downsampled image is input into the 2x super-resolution reconstruction network for processing to obtain a 2x super-resolution reconstruction image. The 2x reconstruction loss function is calculated based on the training image and the 2x super-resolution reconstruction image; The training image is downsampled by 4x to obtain a 4x downsampled image. The 4x downsampled image is input into the 4x super-resolution reconstruction network for processing to obtain a 2x super-resolution reconstruction image and a 4x super-resolution reconstruction image. The 2x reconstruction loss function is calculated based on the training image and the 2x super-resolution reconstruction image, and the 4x loss function is calculated based on the training image and the 4x super-resolution reconstruction image; The training image is downsampled by 8x to obtain an 8x downsampled image. The 8x downsampled image is input into the 8x super-resolution reconstruction network for processing to obtain a 2x super-resolution reconstruction image, a 4x super-resolution reconstruction image, and an 8x super-resolution reconstruction image. The 2x reconstruction loss function is calculated based on the training image and the 2x super-resolution reconstruction image, the 4x reconstruction loss function is calculated based on the training image and the 4x super-resolution reconstruction image, and the 8x reconstruction loss function is calculated based on the training image and the 8x super-resolution reconstruction image; All the 2x reconstruction loss functions, all the 4x reconstruction loss functions, and the 8x reconstruction loss functions are added together to obtain a multi-scale fusion loss function. The multi-scale fusion loss function is used to simultaneously constrain the training of the 2x super-resolution reconstruction network, the 4x super-resolution reconstruction network, and the 8x super-resolution reconstruction network to obtain a 2x super-resolution reconstruction model, a 4x super-resolution reconstruction model, and an 8x super-resolution reconstruction model.
[0047] It should be noted that during network training, a cross-network multi-scale fusion training strategy is adopted. That is, first, super-resolution reconstruction networks with a magnification factor of 2, 4, and 8 are constructed in the above manner. After downsampling the training images by a factor of 2, 4, and 8 respectively, they are fed into the corresponding networks to generate super-resolution reconstruction images with a magnification factor of 2, 4, and 8. Then, combined with the training images, the reconstruction loss functions with a magnification factor of 2, 4, and 8 are calculated respectively, and added together to obtain the cross-network multi-scale fusion loss function. Through backpropagation, the parameters of the super-resolution reconstruction networks with a magnification factor of 2, 4, and 8 are optimized. By adopting the cross-network multi-scale fusion training strategy, the super-resolution reconstruction networks with a magnification factor of 2, 4, and 8 perform super-resolution image reconstruction at the same scale (magnification factor) under different spatial resolutions, and the networks promote each other to accelerate the convergence speed. The super-resolution reconstruction network with a magnification factor of 8 simultaneously conducts super-resolution reconstruction training at three scales (magnification factors), enabling a single network to have the ability of multi-scale (magnification factor) super-resolution reconstruction.
[0048] The evaluation indicators adopted in the present invention include Peak Signal-to-Noise Ratio (PSNR) and Structural SIMilarity (SSIM), specifically as follows: ; ; where I o is the original high-resolution image, I sr is the super-resolution reconstruction image, MSE is the mean square error, m and n are the image sizes (i.e., the height and width of the image), (x, y) is the pixel coordinate, Io(x, y) is the gray value at the (x, y) position of the original high-resolution image, and Isr(x, y) is the gray value at the (x, y) position of the super-resolution reconstruction image.
[0049] ; where μ o and μ sr respectively represent the gray value means of the original high-resolution image I o and the super-resolution reconstruction image I sr , C 1 =( K 1 L ) 2 where L is the dynamic range of pixel values, K 1 is a constant much smaller than 1 (defined as 0.01 here), σ o andσ sr respectively represent the original high - resolution image I o and the super - resolution reconstructed image I sr of the gray - level variance, where C 2 = ( K 2 L ) 2 , K 2 is a constant much smaller than 1 (defined as 0.03 here). σ osr is the original high - resolution image I o and the super - resolution reconstructed image I sr The covariance between them. The loss function formula adopted in the present invention is as follows: ; ; where S represents the super - resolution multiple (2, 4, 8), IHR is the original high - resolution image, ISR is the super - resolution reconstructed image, is the multi - scale fusion loss function, is the 2 - fold reconstruction loss function, is the 4 - fold reconstruction loss function, is the 8 - fold reconstruction loss function, N is the number of images in each batch of training samples, and L is the number of pyramid levels (1, 2, 3) corresponding to the super - resolution multiples (2, 4, 8).
[0050] It should be understood that the various forms of the process shown above can be reordered, steps can be added or deleted. For example, the steps recorded in the disclosure of the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and no limitation is made herein.
[0051] The above - mentioned specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-scale super-resolution reconstruction method for ultra-lightweight Mars remote sensing images, characterized by: The specific steps include: S1: Acquire high-resolution image and low-resolution image pairs; S2: Input the high-resolution image and the low-resolution image pair into the multi-scale super-resolution reconstruction network, and train the multi-scale super-resolution reconstruction network using the multi-scale fusion training strategy to obtain a multi-scale super-resolution reconstruction model; The multi-scale super-resolution reconstruction network includes a 2-fold super-resolution reconstruction branch, a 4-fold super-resolution reconstruction branch, and an 8-fold super-resolution reconstruction branch; S3: Input the image to be reconstructed into a multi-scale super-resolution reconstruction model for processing to obtain a super-resolution reconstructed image.
2. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 1 is characterized in that: During the model training stage, when the low-resolution image input to the multi-scale super-resolution reconstruction network is a 2x down-sampled image, the 2x super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network processes the input image and outputs a 2x super-resolution reconstructed image; when the low-resolution image input to the multi-scale super-resolution reconstruction network is a 4x down-sampled image, the 2x super-resolution reconstruction branch and the 4x super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network process the input image and output a 2x super-resolution reconstructed image and a 4x super-resolution reconstructed image; when the low-resolution image input to the multi-scale super-resolution reconstruction network is an 8x down-sampled image, the 2x super-resolution reconstruction branch, the 4x super-resolution reconstruction branch and the 8x super-resolution reconstruction branch of the multi-scale super-resolution reconstruction network process the input image and output a 2x super-resolution reconstructed image, a 4x super-resolution reconstructed image and an 8x super-resolution reconstructed image.
3. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 1, characterized in that: The 2x super-resolution reconstruction branch includes a first residual image extraction module; the 4x super-resolution reconstruction branch includes a second residual image extraction module; and the 8x super-resolution reconstruction branch includes a third residual image extraction module.
4. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 3 is characterized by: The low-resolution image is upsampled to obtain a 2-fold upsampled image, and the low-resolution image is input into a first residual image extraction module for processing to obtain a 2-fold residual image, and the 2-fold upsampled image and the 2-fold residual image are added pixel by pixel to obtain a 2-fold super-resolution reconstructed image; The 2x residual image is input into the second residual image extraction module for processing to obtain a 4x residual image, the 2x super-resolution reconstructed image is upsampled to obtain a 4x upsampled image, and the 4x upsampled image and the 4x residual image are added pixel by pixel to obtain a 4x super-resolution reconstructed image; The 4x residual image is input into the third residual image extraction module for processing to obtain an 8x residual image, the 4x super-resolution reconstructed image is upsampled to obtain an 8x upsampled image, and the 8x upsampled image and the 8x residual image are added pixel by pixel to obtain an 8x super-resolution reconstructed image.
5. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 3 or 4, characterized in that: The parameters of the first residual image extraction module, the second residual image extraction module and the third residual image extraction module are shared.
6. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 5, characterized in that: The first residual image extraction module, the second residual image extraction module and the third residual image extraction module have the same structure. The first residual image extraction module includes a first 3×3 convolution, a first residual recursive module, a second residual recursive module, a third residual recursive module, a fourth residual recursive module, a fifth residual recursive module, a sixth residual recursive module, a second 3×3 convolution, a first 1×1 convolution, a third 3×3 convolution, a fourth 3×3 convolution and a pixel reorganization module. The low-resolution image input to the first residual image extraction module is processed by the first 3×3 convolution to obtain a feature map A1. The feature map A1 is input to the first residual recursive module for processing to obtain a feature map A2. The feature map A1 and the feature map A2 are non-homologous jump connected to obtain a feature map A3. The feature map A3 is input to the second residual recursive module. The second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module, and the sixth residual recursive module process the image in the same manner as the first residual recursive module; The feature map input to the sixth residual recursive module and the feature map output after being processed by the sixth residual recursive module are added to obtain a feature map B1. The feature map B1 is processed by the second 3×3 convolution, the first 1×1 convolution and the third 3×3 convolution in sequence to obtain a feature map B2. The feature map B2 is globally jump-connected with the feature map A1 to obtain a feature map B3. The feature map B3 is processed by the fourth 3×3 convolution and the pixel recombination module in sequence to obtain a 2x residual image.
7. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 6, characterized in that: The network structures of the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module and the sixth residual recursive module are the same. The first residual recursive module includes a first normalization layer, a second normalization layer, a third normalization layer, a fourth normalization layer, a window multi-head self-attention module, a first multi-layer perceptron, a second multi-layer perceptron, a displacement window multi-head self-attention module, a fifth 3×3 convolution, a second 1×1 convolution and a sixth 3×3 convolution. The feature map C1 input to the first residual recursive module is processed by the first normalization layer and the window multi-head self-attention module to obtain a feature map C2. After adding the feature map C1 and the feature map C2, The feature map C3 is obtained, and the feature map C3 is processed by the second normalization layer and the first multi-layer perceptron to obtain the feature map C4. The feature map C4 is added to the feature map C3 to obtain the feature map C5. The feature map C5 is processed by the third normalization layer and the displacement window multi-head self-attention module to obtain the feature map C6. The feature map C5 and the feature map C6 are added to obtain the feature map C7. The feature map C7 is processed by the fourth normalization layer and the second multi-layer perceptron to obtain the feature map C8. The feature map C8 and the feature map C7 are added to obtain the feature map C9. The feature map C9 is processed by the fifth 3×3 convolution, the second 1×1 convolution and the sixth 3×3 convolution to obtain the output feature map of the first residual recursive module.
8. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 7, characterized in that: Parameter sharing of the first residual recursive module, the second residual recursive module, the third residual recursive module, the fourth residual recursive module, the fifth residual recursive module and the sixth residual recursive module.
9. The multi-scale super-resolution reconstruction method of ultra-lightweight Mars remote sensing images according to claim 2, characterized in that: The training images are downsampled by 2 times, 4 times and 8 times respectively, and 2 times downsampled images, 4 times downsampled images and 8 times downsampled images are obtained accordingly; The 2x downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2x super-resolution reconstructed image, and a 2x reconstruction loss function is calculated based on the training image and the 2x super-resolution reconstructed image; The 4-fold downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2-fold super-resolution reconstructed image and a 4-fold super-resolution reconstructed image, and a 2-fold reconstruction loss function is calculated based on the training image and the 2-fold super-resolution reconstructed image, and a 4-fold loss function is calculated based on the training image and the 4-fold super-resolution reconstructed image; The 8-fold downsampled image is input into the corresponding multi-scale super-resolution reconstruction network for processing to obtain a 2-fold super-resolution reconstructed image, a 4-fold super-resolution reconstructed image, and an 8-fold super-resolution reconstructed image. A 2-fold reconstruction loss function is calculated based on the training image and the 2-fold super-resolution reconstructed image, a 4-fold reconstruction loss function is calculated based on the training image and the 4-fold super-resolution reconstructed image, and an 8-fold reconstruction loss function is calculated based on the training image and the 8-fold super-resolution reconstructed image. All 2x reconstruction loss functions, all 4x reconstruction loss functions and 8x reconstruction loss functions are added together to obtain the multi-scale fusion loss function. The multi-scale fusion loss function is used to constrain the multi-scale super-resolution reconstruction network training to obtain the multi-scale super-resolution reconstruction model.
Citation Information
Patent Citations
Low-illumination image enhancement method based on recursive interactive attention
CN116309182A
Remote sensing image super-resolution method, system and equipment based on implicit neural representation
CN119399026A