Continuous multiple remote sensing image super-resolution method, device, equipment and medium
By combining discrete wavelet transform and inverse wavelet transform with a three-dimensional attention mechanism, the problems of detail loss and specific magnification limitations in image super-resolution reconstruction are solved, achieving efficient continuous magnification super-resolution image restoration and improved analysis accuracy.
Patent Information
- Application Number
- CN202410578056.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies for image super-resolution reconstruction suffer from problems such as loss of detail and the inability to handle super-resolution tasks at a specific magnification.
Discrete wavelet transform and inverse wavelet transform are used for feature extraction. Combined with a multi-level wavelet feature aggregation module and a three-dimensional attention mechanism, the system achieves peer information interaction through skip connections, learns the relationship between coordinates and RGB values, and optimizes the decoding function to predict RGB values.
It effectively restores details in high-resolution images, is suitable for super-resolution tasks at successive multiples, reduces resources and costs, improves the accuracy of image analysis and recognition, and mitigates information loss caused by increasing network depth.
Smart Images

Figure CN120931486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of pattern recognition technology, and more specifically to methods, apparatus, devices and media for super-resolution of remote sensing images at successive multiples. Background Technology
[0002] Due to inherent optical limitations, airborne and space-based acquisition instruments often produce images with low inherent resolution during oil and petrochemical production. Therefore, obtaining high-resolution images rich in texture detail has become a crucial challenge. Image super-resolution reconstruction, which uses specific algorithms to reconstruct high-resolution images from given low-resolution images, effectively improving spatial resolution and restoring texture detail, has attracted widespread attention. This algorithm-based image super-resolution reconstruction can enhance the spatial or spectral resolution of low-resolution images, providing an effective solution for image processing tasks in the petroleum industry.
[0003] With the development of deep learning, super-resolution algorithms based on convolutional neural networks (CNNs) have shown increasingly impressive performance. Dong et al. first applied CNNs to image super-resolution tasks, proposing the Super-Resolution Convolutional Neural Network (SRCNN), which achieved superior super-resolution results using only three convolutional layers compared to other models. However, due to the limitations of hardware at the time, this computational cost made it difficult to apply to practical tasks. To address this, Dong et al. proposed the Fast Super-Resolution Convolutional Neural Network (FSRCNN), introducing deconvolutional layers to reduce the feature dimension, decrease the kernel size, reduce model parameters, and improve reconstruction speed. With iterative hardware upgrades, researchers have focused more on reconstructed image quality. Lim et al. designed the Enhanced Deep Residual Network (EDSR), reducing the impact of batch normalization layers on super-resolution tasks by designing a deeper and wider network structure. RDN, based on the residual block structure, designed densely connected blocks to enhance the interaction of information before and after the network. RCAN added a channel attention mechanism to the network, improving the network's perception of channel information by modeling channel features. SAN, based on the channel attention mechanism, introduced second-order information to mine higher-order features of channel characteristics. Based on cross-scale self-similarity within images, CS-NL designed a cross-scale nonlocal attention mechanism to model the correlation between local and nonlocal features within images.
[0004] The implementation process of the above image super-resolution tasks can be divided into an image feature extraction stage and an upsampling stage. However, in the feature extraction stage, most works focus on designing deeper and wider networks to improve performance. This causes the network to only consider positional information and ignore the frequency information of the image, resulting in the loss of detail in the super-resolution results. In the reconstruction stage, most algorithms use methods such as bicubic interpolation or pixel shuffle to perform upsampling at a specific factor, which means that the pre-trained model can only handle super-resolution tasks at that specific factor. Summary of the Invention
[0005] In view of this, this application provides a method, apparatus, device and medium for super-resolution of remote sensing images at successive magnifications, in order to overcome the shortcomings of existing technologies in the process of super-resolution image reconstruction, such as loss of detail and the inability to handle super-resolution tasks at specific magnifications.
[0006] In a first aspect, embodiments of this application provide a method for super-resolution of remote sensing images at successive magnifications, including:
[0007] The high-resolution image is downsampled at a random ratio, and the resulting downsampled image is used as the input image of the network. The RGB values of the high-resolution image are used as the true RGB values.
[0008] The input image is used to extract features using discrete wavelet transform and inverse wavelet transform to obtain the extracted features;
[0009] Using the extracted features, the relationship between coordinates and RGB values is learned;
[0010] By utilizing the relationship between coordinates and RGB values, the predicted RGB values are obtained, and the loss between the predicted RGB values and the true RGB values is calculated to optimize the decoding function.
[0011] In one possible implementation, the step of extracting features from the input image using discrete wavelet transform and inverse wavelet transform to obtain extracted features includes: performing three downsampling operations on the input image using discrete wavelet transform to obtain three feature maps at different scales, and then using inverse wavelet transform to extract the features at the smallest scale. Figure 3 Secondary upsampling, in which skip connections are used to enable interaction between information at the same level, to obtain extracted features.
[0012] In one possible implementation, the input image is feature extracted by a multi-level wavelet feature aggregation module, which includes six sequentially connected network layers, each consisting of three fully connected layers and a three-dimensional attention mechanism.
[0013] In one possible implementation, the fully connected layer comprises a convolutional layer with a 3×3 kernel, a batch normalization layer, and a ReLU activation function layer.
[0014] In one possible implementation, the three-dimensional attention mechanism includes an average pooling layer, a first convolutional layer, a ReLU activation function, a second convolutional layer, a first Sigmoid layer, a third convolutional layer, a fourth convolutional layer, and a second Sigmoid layer.
[0015] The average pooling layer, the first convolutional layer, the ReLU activation function, the second convolutional layer, and the first sigmoid layer are connected in sequence and connected in parallel with the third convolutional layer.
[0016] The input feature map F∈R of the three-dimensional attention mechanism C×H×W The channel attention weight matrix T is obtained by sequentially passing the material through an average pooling layer, a first convolutional layer, a ReLU activation function, a second convolutional layer, and a first sigmoid layer. c ∈R C×1×1 The channel attention weight matrix T C The expression is:
[0017] T c =Sigmoid(f 1×1 (ReLU(f 1×1 (Avg(F)))))
[0018] Where Avg represents global average pooling, f 1×1 This indicates a convolution operation with a 1×1 kernel, ReLU represents the ReLU activation function, and Sigmoid represents the Sigmoid function;
[0019] The input feature map F∈R C×H×W The spatial attention weight matrix T is obtained through the third convolutional layer. s ∈R 1×H×W The spatial attention weight matrix T s The expression is:
[0020] T s =Sigmoid(f 1×1 (F))
[0021] Among them, f 1×1 This indicates a convolution operation with a 1×1 kernel, and Sigmoid represents the Sigmoid function.
[0022] The channel attention weight matrix T C With spatial attention weight matrix T s Matrix multiplication is performed, and the result is then passed through a fourth convolutional layer and a second Sigmoid layer to obtain a three-dimensional attention weight matrix. This three-dimensional attention weight matrix is then combined with the input feature map F∈R. C×H×WPerform element-wise multiplication to obtain the weighted feature map. The weighted feature map The expression is:
[0023]
[0024] in, f represents matrix multiplication. 1×1 This indicates a convolution operation with a 1×1 kernel, Sigmoid represents the Sigmoid function, and ⊙ represents element-wise multiplication.
[0025] In one possible implementation, the relationship between coordinates and RGB values is as follows:
[0026]
[0027] Among them, z i (i∈{t1,t2,t3,t4}) represents the target position x. t The latent code that is closest in Euclidean distance in the four directions of up, down, left, and right, x t For the target location, x i For latent encoding z i The coordinates f in the image domain θ S is the decoding function. i It is x t With x i The area of the rectangle formed by the two sides is S = ∑ i S t This is the normalization coefficient.
[0028] In one possible implementation, the formula for calculating the loss between the predicted RGB value and the true RGB value is as follows:
[0029]
[0030] Among them, L i (θ) represents the loss between the predicted RGB values and the true RGB values, where θ is a function parameter, V(x) sr V(x) represents the predicted RGB value. hr ) represents the true RGB value, i represents the number of predicted RGB values, and n represents the total number of predicted RGB values.
[0031] Secondly, embodiments of this application provide a super-resolution apparatus for continuously multiplied remote sensing images, comprising:
[0032] The data preparation module is used to downsample the high-resolution image at a random ratio, use the resulting downsampled image as the input image of the network, and use the RGB values of the high-resolution image as the RGB ground truth values for training the network.
[0033] The feature extraction module is used to extract features from the input image using discrete wavelet transform and inverse wavelet transform to obtain extracted features;
[0034] The relationship building module is used to learn the relationship between coordinates and RGB values using the extracted features;
[0035] The RGB value prediction module is used to obtain the predicted RGB value by utilizing the relationship between coordinates and RGB values, and to calculate the loss between the predicted RGB value and the true RGB value in order to optimize the decoding function.
[0036] Thirdly, embodiments of this application provide an electronic device, including:
[0037] processor;
[0038] Memory;
[0039] And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method described in any one of the first aspects.
[0040] Fourthly, embodiments of this application provide a computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the method described in any one of the first aspects.
[0041] Compared to the traditional U-net network, which uses pooling and deconvolution layers as downsampling and upsampling layers respectively, resulting in the loss of original image feature information during pooling, this embodiment uses multi-level wavelet transform technology. Discrete wavelet transform (DWT) replaces the downsampling layer, and inverse wavelet transform (IDWT) replaces the upsampling layer. By using skip connections to achieve interaction between information at the same level, high-resolution detail information is recovered from low-resolution images, effectively improving the visual quality of the images. This makes the images suitable for applications such as oil exploration, where resolution can be adjusted and optimized as needed. By increasing image resolution, the resources and costs required during data acquisition are reduced. In some cases, this can be achieved by improving existing data rather than re-acquiring high-resolution data. Furthermore, high-resolution images are generally easier to analyze and recognize, especially when fine structural information is required, improving the accuracy of analysis and recognition tasks. Simultaneously, it alleviates the information loss problem caused by increased network depth and enables continuous super-resolution tasks. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A flowchart illustrating the method for super-resolution of continuously multiplied remote sensing images provided in this application embodiment;
[0044] Figure 2 A schematic diagram of the structure of the three-dimensional attention mechanism provided in the embodiments of this application;
[0045] Figure 3 A structural block diagram of the super-resolution device for continuous magnification remote sensing images provided in the embodiments of this application;
[0046] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0047] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0048] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0049] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0050] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0051] See Figure 1 This is a flowchart illustrating the continuous scaling remote sensing image super-resolution method provided in this application embodiment. Figure 1 As shown, it mainly includes the following steps.
[0052] Step S1: Downsample the input high-resolution image at a random ratio, and use the resulting downsampled image as the input image of the network. The position coordinates x of the high-resolution image are then... hr RGB value V(x) hr ) is used as the ground truth in the training network.
[0053] Step S2: Extract features from the input image using discrete wavelet transform and inverse wavelet transform to obtain extracted features. Specifically, the input image is downsampled three times using discrete wavelet transform to obtain feature maps at three different scales, and then the smallest scale feature map is extracted using inverse wavelet transform. Figure 3 Secondary upsampling, in which skip connections are used to enable interaction between information at the same level, to obtain extracted features.
[0054] Step S2 extracts features from the input image using a multi-level wavelet feature aggregation module (MWFA). This module comprises six sequentially connected network layers. The first three layers first perform three downsampling operations on the input image using discrete wavelet transform to obtain feature maps at three different scales. Then, the last three layers sequentially use inverse wavelet transform to refine the smallest scale feature map. Figure 3 Secondary upsampling, during which interaction between sibling information is achieved through skip connections.
[0055] Each layer of the network consists of three fully connected layers and a three-dimensional attention mechanism.
[0056] The fully connected layer includes a 3×3 convolutional layer (Conv), a batch normalization layer (BN), and a ReLU activation function layer. By adding weights to channel information and spatial location, the network's ability to learn and discriminate feature maps is enhanced.
[0057] See Figure 2 This is a schematic diagram of the structure of the three-dimensional attention mechanism provided in an embodiment of this application. Figure 2 As shown, the three-dimensional attention mechanism includes an average pooling layer (AvgPool), a first convolutional layer, a ReLU activation function, a second convolutional layer, a first Sigmoid layer, a third convolutional layer, a fourth convolutional layer, and a second Sigmoid layer.
[0058] The average pooling layer (AvgPool), the first convolutional layer, the ReLU activation function, the second convolutional layer, and the first sigmoid layer are connected in sequence and connected in parallel with the third convolutional layer.
[0059] Input feature map F∈R of the 3D attention mechanism C×H×W The channel attention weight matrix T is obtained by sequentially passing the material through an average pooling layer, a first convolutional layer, a ReLU activation function, a second convolutional layer, and a first sigmoid layer.c ∈R C×1×1 The channel attention weight matrix T C The expression is:
[0060]
[0061] Where Avg represents global average pooling, f 1×1 This indicates a convolution operation with a 1×1 kernel, ReLU represents the ReLU activation function, and Sigmoid represents the Sigmoid function.
[0062] The input feature map F∈R C×H×W The spatial attention weight matrix T is obtained through the third convolutional layer. s ∈R 1×H×W The spatial attention weight matrix T s The expression is:
[0063] T s =Sigmoid(f 1×1 (F)) (2)
[0064] Among them, f 1×1 This indicates a convolution operation with a 1×1 kernel, and Sigmoid represents the Sigmoid function.
[0065] The channel attention weight matrix T C With spatial attention weight matrix T s Perform matrix multiplication, then pass the result through a fourth convolutional layer and a second Sigmoid layer to obtain the three-dimensional attention weight matrix T∈R. C×H×W Then, the three-dimensional attention weight matrix T∈R C×H×W With the input feature map F∈R C×H×W Perform element-wise multiplication to obtain the weighted feature map. The weighted feature map The expression is:
[0066]
[0067] in, f represents matrix multiplication. 1×1 This indicates a convolution operation with a 1×1 kernel, Sigmoid represents the Sigmoid function, and ⊙ represents element-wise multiplication.
[0068] The convolution kernels of the first, second, third, and fourth convolutional layers are all 1×1.
[0069] Step S3: Using the extracted features, learn the relationship between coordinates and RGB values.
[0070] Suppose that each image I(i) can be represented as a two-dimensional feature map. The corresponding decoding function is represented as f θ , where θ is a function parameter. The relationship between the coordinates and the signal can then be expressed as:
[0071]
[0072] Where x represents the coordinate and z represents the vector. When the decoding function f... θ Given that for each vector z, the position space can be mapped to the RGB value space, i.e. Then for any target position x in image I(i) t x can be t The corresponding RGB value V(x) t ) is defined as:
[0073] V(x t )=f θ (z i ,x t -x i (5)
[0074] Among them, z i For two-dimensional feature mapping map M (i) Center and target position x t The latent code with the closest Euclidean distance, x i For latent encoding z i Coordinates in the image domain. For example... Figure 1 As shown in (c) in the figure, z t4 For two-dimensional feature mapping map M (i) Center and position x t The latent code with the closest Euclidean distance, x t4 For latent encoding z t4 The coordinates in the image domain. Thus, in the function f θ Under the influence of this function, a continuous image can be represented as a two-dimensional feature map M. (i) The latent encoding in, each latent encoding z i It represents a portion of a continuous image and can predict the signal values of the nearest set of coordinates.
[0075] Inspired by the bilinear interpolation method, in order to more accurately predict the corresponding RGB values, this embodiment further extends formula (5) to:
[0076]
[0077] Among them, z i (i∈{t1,t2,t3,t4}) represents the target position x. tThe latent code that is closest in Euclidean distance in the four directions (up, down, left, right), x i For latent encoding z i The coordinates f in the image domain θ S is the decoding function. i It is x t With x i The area of the rectangle formed by the two sides is S = ∑ i S t The normalized coefficients are used to achieve merging predictions by using normalized confidence levels.
[0078] Step S4: Utilize the relationship between coordinates and RGB values to obtain the predicted RGB value V(x). sr ), and calculate the loss L between the predicted RGB values and the true RGB values. i (θ), continuously optimize the decoding function f θ The loss L i The formula for calculating (θ) is:
[0079]
[0080] Where θ is the function parameter, V(x) sr V(x) represents the predicted RGB value. hr ) represents the RGB values of the high-resolution image, which are taken as the true values. i represents the number of predicted RGB values, and n represents the total number of predicted RGB values.
[0081] In this embodiment, taking a 48×48 image patch as the input as an example, and assuming B is the batch size, firstly, B = 16 random scales r are sampled in a uniform distribution U(1,4). 1~B Then crop B images of size B from the training images. The image patch. A 48×48 input is used as a downsampled image patch of the corresponding multiple. The data is enhanced and the image patch is expanded by horizontal or vertical flipping and 90° rotation. For the ground truth, in this embodiment, the image is converted into pixel samples, that is, "coordinate-RGB value" data pairs are formed, and 48×48 pixel samples are sampled to ensure that the shape and size are the same as the ground truth during batch processing. Decoder f θ It consists of a 5-layer MLP, including a ReLU activation function and a 256-dimensional hidden layer. The Adam optimizer is used, with an initial learning rate of 1×10⁻⁶. -4 The training process involves 1,000 cycles, with the learning rate decreasing by 0.5 times after every 200 cycles.
[0082] This application conducted experiments using convolutional neural networks. Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) were used as evaluation metrics. PSNR represents the similarity between corresponding pixels in the super-resolution image and the ground truth image; a higher PSNR value indicates a better super-resolution effect. SSIM represents the difference between blocks in the super-resolution image and the ground truth image. The average difference across all blocks is used to calculate the value for the entire image. A higher SSIM value indicates that the super-resolution image is more similar to the ground truth image in terms of brightness, contrast, and structure, and thus a better super-resolution effect.
[0083] Experimental results demonstrate that this invention can improve the performance of convolutional neural networks in image super-resolution tasks. The experiments first compared the technical solution disclosed in this application with integer factor-based algorithms such as EDSR, RDN, RCAN, SAN, CS-NL, and remote sensing super-resolution algorithms CTN, TransENet, and ReFDN on 2x and 4x super-resolution tasks. Compared to suboptimal methods, the technical solution disclosed in this application improves PSNR by 0.15 dB on the RSC11 dataset, 0.01 dB on the UC-Merced dataset, and 0.02 dB on the NWPU45 dataset when the magnification factor is 2; when the magnification factor is 4, it improves PSNR by 0.05 dB on the RSC11 dataset, 0.19 dB on the UC-Merced dataset, and achieves the best result on the NWPU45 dataset. These results indicate that although the technical solution disclosed in this application is designed for continuous magnification super-resolution tasks, it still achieves competitive experimental results on specific magnification super-resolution tasks, proving the effectiveness of the MW-IR method. In the super-resolution task, the comparison model was a deep convolutional neural network, and the experiments used the RSC11 dataset, UC-Merced dataset, and NWPU-RESISC45 dataset. These are remote sensing datasets. The experimental results are shown in Table 1.
[0084] Table 1: Quantitative comparison results (PSNR / SSIM) of magnification factors of 2 and 4 on the RSC11, UC-Merced, and NWPU45 datasets. Bold indicates the best results, and underline indicates the second best results.
[0085]
[0086]
[0087] The results of quantitative analysis of continuous super-resolution of remote sensing images are further presented, and the comparison results are shown in Table 2.
[0088] Table 2: Results of continuous fold quantitative comparison of MW-IR on the NWPU45 dataset, with bold indicating the best results.
[0089]
[0090] Meta-SR employs a meta-learning approach, using coordinate and scaling factor-related vectors as input to predict the weights of the scaling filter. For each pixel location in the super-resolution image to be predicted, the predicted pixel value is predicted by convolving the corresponding mapping features from the low-resolution image with the predicted weights. Liif, based on implicit representations, learns the implicit representation relationships during the upsampling process, enabling the network to achieve arbitrary scaling. Compared to Liif, MW-IR achieves an average PSNR improvement of 0.43 dB. Experiments demonstrate that MW-IR can reconstruct super-resolution images with better evaluation metrics on continuous super-resolution tasks.
[0091] Corresponding to the above embodiments, this application also provides a super-resolution device for continuous multiple remote sensing images.
[0092] See Figure 3 This is a structural block diagram of the super-resolution device for continuous-magnification remote sensing images provided in the embodiments of this application. Figure 3 As shown, it mainly includes the following modules.
[0093] The data preparation module 301 is used to downsample the high-resolution image at a random ratio, use the resulting downsampled image as the input image of the network, and use the RGB values of the high-resolution image as the RGB ground truth values for training the network.
[0094] Feature extraction module 302 is used to extract features from the input image using discrete wavelet transform and inverse wavelet transform to obtain extracted features;
[0095] Relationship building module 303 is used to learn the relationship between coordinates and RGB values using the extracted features;
[0096] The RGB value prediction module 304 is used to obtain the predicted RGB value by utilizing the relationship between coordinates and RGB values, and to calculate the loss between the predicted RGB value and the true RGB value in order to optimize the decoding function.
[0097] The feature extraction module 302 includes a multi-level wavelet feature aggregation module. Each layer of the multi-level wavelet feature aggregation module consists of three fully connected layers and a three-dimensional attention mechanism.
[0098] The fully connected layer includes a 3×3 convolutional layer (Conv), a batch normalization layer (BN), and a ReLU activation function layer.
[0099] In the multi-level wavelet feature aggregation module, Discrete Wavelet Transform (DWT) is used to replace the downsampling layer, and Inverse Wavelet Transform (IDWT) is used to replace the upsampling layer. Interaction between information at the same level is achieved through skip connections.
[0100] The three-dimensional attention mechanism includes an average pooling layer (AvgPool), a first convolutional layer, a ReLU activation function, a second convolutional layer, a first Sigmoid layer, a third convolutional layer, a fourth convolutional layer, and a second Sigmoid layer.
[0101] The average pooling layer (AvgPool), the first convolutional layer, the ReLU activation function, the second convolutional layer, and the first sigmoid layer are sequentially connected and connected in parallel with the third convolutional layer. The output of the first sigmoid layer is multiplied by the output of the third convolutional layer, and the result of the matrix multiplication is then passed sequentially through the fourth convolutional layer and the second sigmoid layer to obtain a three-dimensional attention weight matrix. This three-dimensional attention weight matrix is then multiplied element-wise with the input features to obtain a weighted feature map.
[0102] It should be noted that the specific content involved in the embodiments of this application can be found in the description of the above method embodiments, and will not be repeated here for the sake of brevity.
[0103] Corresponding to the above embodiments, this application also provides an electronic device.
[0104] See Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 400 may include a processor 401, a memory 402, and a communication unit 403. These components communicate via one or more buses. Those skilled in the art will understand that the electronic device structure shown in the figures does not constitute a limitation on the embodiments of this application. It may be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0105] The communication unit 403 is used to establish a communication channel, thereby enabling the electronic device to communicate with other devices.
[0106] The processor 401 serves as the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory to perform various functions and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 401 may consist only of a central processing unit (CPU). In this embodiment, the CPU may have a single processing core or include multiple processing cores.
[0107] Memory 402 is used to store the execution instructions of processor 401. Memory 402 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0108] When the execution instructions in memory 402 are executed by processor 401, the electronic device 400 is able to perform some or all of the steps in the above method embodiments.
[0109] Corresponding to the above embodiments, this application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, wherein when the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. Specifically, the computer-readable storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0110] Corresponding to the above embodiments, this application also provides a computer program product containing executable instructions that, when executed on a computer, cause the computer to perform some or all of the steps in the above method embodiments.
[0111] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0112] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0113] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0114] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0115] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for super-resolution of continuously multiplied remote sensing images, characterized in that, include: The high-resolution image is downsampled at a random ratio, and the resulting downsampled image is used as the input image of the network. The RGB values of the high-resolution image are used as the true RGB values. The input image is used to extract features using discrete wavelet transform and inverse wavelet transform to obtain the extracted features; Using the extracted features, the relationship between coordinates and RGB values is learned; By utilizing the relationship between coordinates and RGB values, the predicted RGB values are obtained, and the loss between the predicted RGB values and the true RGB values is calculated to optimize the decoding function.
2. The method according to claim 1, characterized in that, The step of extracting features from the input image using discrete wavelet transform and inverse wavelet transform includes: performing three downsampling operations on the input image using discrete wavelet transform to obtain three feature maps at different scales; then performing three upsampling operations on the smallest scale feature map using inverse wavelet transform; and achieving interaction between information at the same level through skip connections during the upsampling process to obtain the extracted features.
3. The method according to claim 1, characterized in that, The input image is processed by a multi-level wavelet feature aggregation module, which includes six sequentially connected network layers. Each network layer consists of three fully connected layers and a three-dimensional attention mechanism.
4. The method according to claim 3, characterized in that, The fully connected layer includes a convolutional layer with a 3×3 kernel, a batch normalization layer, and a ReLU activation function layer.
5. The method according to claim 3, characterized in that, The three-dimensional attention mechanism includes an average pooling layer, a first convolutional layer, a ReLU activation function, a second convolutional layer, a first Sigmoid layer, a third convolutional layer, a fourth convolutional layer, and a second Sigmoid layer. The average pooling layer, the first convolutional layer, the ReLU activation function, the second convolutional layer, and the first sigmoid layer are connected in sequence and connected in parallel with the third convolutional layer. The input feature map F∈R of the three-dimensional attention mechanism C×H×W The channel attention weight matrix T is obtained by sequentially passing the material through an average pooling layer, a first convolutional layer, a ReLU activation function, a second convolutional layer, and a first sigmoid layer. c ∈R C×1×1 The channel attention weight matrix T C The expression is: T c =Sigmoid(f 1×1 (ReLU(f 1×1 (Avg(F))))) Where Avg represents global average pooling, f 1×1 This indicates a convolution operation with a 1×1 kernel, ReLU represents the ReLU activation function, and Sigmoid represents the Sigmoid function; The input feature map F∈R C×H×W The spatial attention weight matrix T is obtained through the third convolutional layer. s ∈R 1×H×W The spatial attention weight matrix T s The expression is: T s =Sigmoid(f 1×1 (F)) Among them, f 1×1 This indicates a convolution operation with a 1×1 kernel, and Sigmoid represents the Sigmoid function. The channel attention weight matrix T C With spatial attention weight matrix T s Matrix multiplication is performed, and the result is then passed through a fourth convolutional layer and a second Sigmoid layer to obtain a three-dimensional attention weight matrix. This three-dimensional attention weight matrix is then combined with the input feature map F∈R. C×H×W Perform element-wise multiplication to obtain the weighted feature map. The weighted feature map The expression is: in, f represents matrix multiplication. 1×1 This indicates a convolution operation with a 1×1 kernel, Sigmoid represents the Sigmoid function, and ⊙ represents element-wise multiplication.
6. The method according to claim 1, characterized in that, The relationship between coordinates and RGB values is as follows: Among them, z i (i∈{t1, t2, t3, t4}) represents the target position x. t The latent code that is closest in Euclidean distance in the four directions of up, down, left, and right, x t For the target location, x i For latent encoding z i The coordinates f in the image domain θ S is the decoding function. i It is x t With x i The area of the rectangle formed by the two sides is S = ∑ i S t This is the normalization coefficient.
7. The method according to claim 1, characterized in that, The formula for calculating the loss between the predicted RGB value and the true RGB value is as follows: Among them, L i (θ) represents the loss between the predicted RGB values and the true RGB values, where θ is a function parameter, V(x) sr V(x) represents the predicted RGB value. hr ) represents the true RGB value, i represents the number of predicted RGB values, and n represents the total number of predicted RGB values.
8. A super-resolution device for continuously multiplied remote sensing images, characterized in that, include: The data preparation module is used to downsample the high-resolution image at a random ratio, use the resulting downsampled image as the input image of the network, and use the RGB values of the high-resolution image as the RGB ground truth values for training the network. The feature extraction module is used to extract features from the input image using discrete wavelet transform and inverse wavelet transform to obtain extracted features; The relationship building module is used to learn the relationship between coordinates and RGB values using the extracted features; The RGB value prediction module is used to obtain the predicted RGB value by utilizing the relationship between coordinates and RGB values, and to calculate the loss between the predicted RGB value and the true RGB value in order to optimize the decoding function.
9. An electronic device, characterized in that, include: processor; Memory; And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 7.