Remote sensing landform enhancement algorithm
Through the dual-domain multi-scale attention fusion super-resolution network, combined with the characteristics of the wavelet domain and spatial domain, the limitations and instability problems in the process of improving the resolution of DEM images in the prior art are solved, and high-precision DEM super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202411851756.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems such as restriction of fixed receptive fields, training instability and pattern collapse when improving DEM image resolution, which is difficult to meet the needs of refined applications such as geological disaster prediction and urban drainage system design.
A dual-domain multi-scale attention fusion super-resolution network is proposed. Through the combination of wavelet domain and spatial domain, high- and low-frequency information is extracted using the Hal wavelet transform, and the importance relationship between different depth feature maps is obtained through the multi-scale attention fusion module group, and the loss function at the pixel and edge level is supervised to improve the accuracy of DEM super-resolution reconstruction.
This method can have feature enhancement functions in shallow layers, improve the resolution of DEM images, enhance the performance of topographic and topographic details, and improve the accuracy and robustness of super-resolution reconstruction.
Smart Images

Figure CN119991458A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of remote sensing terrain enhancement, and more specifically, to a remote sensing terrain enhancement algorithm. Background Art
[0002] With the rapid development of geographic information technology (GIS), digital elevation models (DEMs) as digital representations of terrain surface morphology have been widely used in the fields of geology, environment and urban planning. The existing acquisition of DEM images through aerial photography or satellite remote sensing can quickly obtain terrain data over a large range, but its resolution is low and it is difficult to meet the needs of refined applications. The existing acquisition of DEM images through laser radar (LiDAR) can provide high-precision terrain data, but its data acquisition and processing process is complicated and affected by factors such as weather and terrain coverage. Although DEM products around the world provide valuable resources for terrain data research, such as SRTM (Shuttle Radar Topographic Mapping Data), they are often insufficient in applications such as geological disaster prediction and urban drainage system design with a spatial resolution of 30 meters. Therefore, how to use super-resolution technology to improve the resolution of DEM images has become an urgent problem to be solved.
[0003] The existing interpolation methods to improve the resolution of DEM images mainly include: inverse distance weighted interpolation, Kriging interpolation and bilinear interpolation. However, interpolation methods often have problems with smoothing effect and loss of details when processing large-scale or complex terrain data, especially in areas with drastic terrain changes. With the development of deep learning technology, the resolution of DEM images is now improved through convolutional neural network models (CNN) and generative adversarial networks (GAN). CNN can automatically learn the mapping relationship between low-resolution DEMs (LR) and high-resolution DEMs (HR), and GAN generates high-quality HR data through adversarial learning between generators and discriminators.
[0004] However, the above methods still face some challenges in processing DME data, such as the limitation of fixed receptive field, training instability and mode collapse. Therefore, there is an urgent need for a DEM super-resolution reconstruction method to solve the above technical problems. Summary of the invention
[0005] The present invention provides a remote sensing topography enhancement algorithm to solve the technical problems in the above-mentioned background technology.
[0006] The present invention provides a remote sensing terrain enhancement algorithm, comprising the following steps:
[0007] Step S101, preprocessing the existing high-resolution DEM image and low-resolution DEM image to construct a training data set;
[0008] The training data set includes G training samples, of which G1 training samples are used for training and G2 training samples are used for testing;
[0009] A training sample includes a sample data and a sample label, wherein the sample data corresponds to the preprocessed low-resolution DEM image, and the sample label corresponds to the preprocessed high-resolution DEM image, wherein there is a 4-fold mapping ratio between the preprocessed low-resolution DEM image and the preprocessed high-resolution DEM image;
[0010] Step S102, constructing a dual-domain multi-scale attention fusion super-resolution network;
[0011] The dual-domain multi-scale attention fusion super-resolution network includes: a first convolutional layer, a first reconstruction sub-network, a second reconstruction sub-network, a third reconstruction sub-network, a second convolutional layer, an upsampling layer, and a third convolutional layer;
[0012] The first reconstruction subnetwork, the second reconstruction subnetwork and the third reconstruction subnetwork all include: a wavelet guidance module, a fourth convolutional layer, a multi-scale attention fusion module group and a wavelet separation module, wherein the multi-scale attention fusion module group is formed by cascading and stacking 16 multi-scale attention fusion modules;
[0013] Step S103, training the dual-domain multi-scale attention fusion super-resolution network using a training data set;
[0014] The input of the dual-domain multi-scale attention fusion super-resolution network corresponds to the sample data of the training sample, and the output is a super-resolution DEM image;
[0015] The number of training samples G, the number of training samples used for training G1, and the number of training samples used for testing G2 are all custom parameters.
[0016] Furthermore, the high-resolution DEM images come from the DAAC Asian High Mountain Dataset, and the low-resolution DEM images come from the ASTER global digital elevation model. The longitude and latitude range covered by the low-resolution DEM images is the same as that of the high-resolution DEM images.
[0017] Furthermore, the existing high-resolution DEM images and low-resolution DEM images are preprocessed to construct a training dataset, including the following steps:
[0018] Step S201, geospatial registration of the high-resolution DEM image and the low-resolution DEM image;
[0019] Step S202, performing projection transformation processing on the high-resolution DEM image and the low-resolution DEM image;
[0020] Step S203, the high-resolution DEM image is uniformly cropped to a size of 512×512 pixels, and the low-resolution DEM image is uniformly cropped to establish a 4-fold mapping ratio with the high-resolution DEM image.
[0021] Furthermore, the first convolutional layer inputs the low-resolution DEM image and outputs the first convolutional feature map;
[0022] The first reconstruction subnetwork inputs the first convolutional feature map and outputs the first feature map;
[0023] The first convolution feature map output by the first convolution layer and the first feature map output by the first reconstruction subnetwork are fused to obtain a first fused feature map, and then upsampled by bicubic interpolation to obtain a first upsampled map;
[0024] The second reconstruction subnetwork inputs the first up-sampled image and outputs the second feature image;
[0025] The first feature map output by the first reconstruction subnetwork is upsampled by the bicubic interpolation method to obtain a second upsampled map, which is then fused with the first convolutional feature map output by the first convolutional layer to obtain a second fused feature map. The second fused feature map is then upsampled by the bicubic interpolation method to obtain a third upsampled map. The third upsampled map is fused with the second feature map output by the second reconstruction subnetwork to obtain a third fused feature map. The third fused feature map is then upsampled by the bicubic interpolation method to obtain a fourth upsampled map.
[0026] The third reconstruction subnetwork inputs the fourth up-sampled image and outputs the third feature map;
[0027] The second convolution layer inputs the third feature map and outputs the second convolution feature map;
[0028] The upsampling layer performs sub-pixel upsampling on the first convolution feature map output by the first convolution layer to obtain a fifth upsampling map, and then fuses the fifth upsampling map with the second convolution feature map output by the second convolution layer to obtain a fourth fused feature map;
[0029] The third convolutional layer inputs the fourth fused feature map and outputs a super-resolution DEM image.
[0030] Furthermore, the first reconstruction subnetwork I Sub_Network1 The calculation formula is as follows:
[0031] F1=I Sub_Network1 (I Conv1 (DEM LR ));
[0032] Among them I Conv1represents the first convolutional layer, F1 represents the first feature map output by the first reconstruction sub-network, and I Conv1 (DEM LR ) represents the first convolution feature map output by the first convolution layer, DEM LR Represents a low-resolution DEM image;
[0033] Second reconstruction sub-network I Sub_Network2 The calculation formula is as follows:
[0034] F2=I Sub_Network2 (I Bicubic (F1+I Conv1 (DEM LR )));
[0035] Where F2 represents the second feature map output by the second reconstruction sub-network, I Bicubic (F1+I Conv1 (DEM LR )) represents the first upsampling image, F1+I Conv1 (DEM LR ) represents the first fusion feature map, I Bicubic represents bicubic interpolation;
[0036] The third reconstruction sub-network I Sub_Network3 The calculation formula is as follows:
[0037] F3=I Sub_Network3 (I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2));
[0038] Where F3 represents the third feature map output by the third reconstruction sub-network, I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2) represents the fourth upsampling graph, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2 represents the third fusion feature map, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1)) represents the third upsampling graph, I Conv1 (DEM LR )+I Bicubic (F1) represents the second fusion feature map, I Bicubic(F1) represents the second up-sampling graph.
[0039] Furthermore, the calculation formula of the third convolutional layer is as follows:
[0040] DEM SR =I Conv3 (I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR )));
[0041] Among them I Conv2 represents the second convolutional layer, I Conv3 Represents the third convolutional layer, DEM SR represents the super-resolution DEM image output by the third convolutional layer, I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR )) represents the fourth fusion feature map, I Conv2 (F3) represents the second convolutional feature map, I Pixcel (I Conv1 (DEM LR )) represents the fifth upsampling graph, I Pixcel Indicates sub-pixel upsampling.
[0042] Further, the wavelet guided module includes: a wavelet domain processing module, a fifth convolutional layer and a sixth convolutional layer;
[0043] The wavelet domain processing module decomposes the feature map of the input wavelet guidance module into high-frequency components and low-frequency components through discrete wavelet transform, and then performs dimensionality reduction processing on it through the 1×1 convolution kernel of the fifth convolution layer to obtain the third convolution feature map;
[0044] The feature map of the input wavelet guided module is extracted by using the convolution kernel of the sixth convolution layer with a size of 3×3 to obtain the fourth convolution feature map;
[0045] Fusing the third convolution feature map and the fourth convolution feature map to obtain a fourth fused feature map;
[0046] The calculation formula of the wavelet guided module is as follows:
[0047] F Output =Concat[I Conv6 (F Input ), I Conv5 (I DWT (F Input ))];
[0048] Among them I Conv5 represents the fifth convolutional layer, I Conv6represents the sixth convolutional layer, F Output represents the fourth fused feature map output by the wavelet-guided module, F Input Represents the feature map of the input wavelet guided module, I DWT represents the wavelet domain processing module, I Conv4 (I DWT (F Input )) represents the third convolutional feature map, I Conv6 (F Input ) represents the fourth convolutional feature map, and Concat represents the concatenation function.
[0049] Furthermore, the multi-scale attention fusion module group is composed of 16 multi-scale attention fusion modules of three different types, which are cascaded and stacked, namely the first module, the second module and the third module, wherein the first multi-scale attention fusion module is the third module, the second to the seventh multi-scale attention fusion modules are stacked in turn in the order of the first module and the second module, the eighth multi-scale attention fusion module is the first module, and the last eight multi-scale attention fusion modules are stacked in the same order as the first eight;
[0050] The first module is built on the channel attention module;
[0051] After the first convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. After the second convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. This process is repeated to obtain multiple proportion sequences extracted from features at different depths. After weighting, a new proportion sequence is obtained and applied to the original input of the first module as the final output.
[0052] The channel attention mechanism includes two branches. The first branch obtains the first channel vector of the feature map input to the first module through average pooling calculation. The second branch obtains the second channel vector of the feature map input to the first module through maximum pooling calculation. The first channel vector and the second channel vector are both subjected to nonlinear mapping of two fully connected layers to generate a first weight vector and a second weight vector. The first weight vector and the second weight vector are respectively element-wise multiplied with the feature map input to the first module and then concatenated to obtain a first updated feature map, wherein the first channel vector, the second channel vector, the first weight vector and the second weight vector have the same number of dimensions.
[0053] The second module is built on the Inception module based on dilated convolution;
[0054] The Inception module of the dilated convolution includes 3 filters, each of which has a size of 3×3, and the dilation factors are 1, 2, and 3, respectively. The corresponding receptive fields are 3×3, 5×5, and 7×7, respectively. The feature maps input into the second module are extracted through the 3 filters and then concatenated to obtain the second updated feature map.
[0055] The third module is built based on the edge-enhanced diversified branch module;
[0056] The edge enhancement diversification branch module includes 6 parallel branches. The first parallel branch consists of a 1×1 convolution kernel, the second parallel branch consists of a 3×3 convolution kernel, the third parallel branch consists of a 1×1 convolution kernel and a 3×3 convolution kernel, the fourth parallel branch consists of a 1×1 convolution kernel and a horizontal Sobel edge detection operator, the fifth parallel branch consists of a 1×1 convolution kernel and a vertical Sobel edge detection operator, and the sixth parallel branch consists of a 1×1 convolution kernel and a Laplacian operator. The feature maps input into the third module are extracted through the 6 parallel branches and then spliced to obtain the third updated feature map.
[0057] Further, the wavelet separation module includes: a seventh convolution layer, an inverse wavelet transform module, an eighth convolution layer and an interpolation processing layer;
[0058] The feature map of the input wavelet separation module is reconstructed through the seventh convolution layer and the wavelet inverse transform module to obtain the fifth convolution feature map;
[0059] The feature map of the input wavelet separation module is reconstructed through the eighth convolution layer and the interpolation processing layer to obtain the sixth convolution feature map;
[0060] Fusing the fifth convolution feature map and the sixth convolution feature map to obtain a fifth fused feature map;
[0061] The calculation formula of the wavelet separation module is as follows:
[0062] F Output =I ADD (I Bicubic (I Conv8 (F Input )),I Conv7 (I IWT (F Input )));
[0063] Among them I Conv7 represents the seventh convolutional layer, I Conv8 represents the eighth convolutional layer, F Output represents the fifth fusion feature map output by the wavelet separation module, F InputRepresents the feature map of the input wavelet separation module, I Bicubic represents the interpolation processing layer, I IWT represents the inverse wavelet transform module, I ADD Represents the addition function.
[0064] Furthermore, the loss function L of the dual-domain multi-scale attention fusion super-resolution network is MS The calculation formula includes:
[0065] L MS =α×L Pixcel +β×L edge ;
[0066] Where L Pixcel Represents the loss function for reconstructing pixel-level features, L edge represents the edge pixel level feature loss function, α and β represent L Pixcel and L edge The corresponding custom weight coefficient, and the sum of α and β is 1;
[0067] Reconstruct pixel-level feature loss function L Pixcel The calculation formula is as follows:
[0068]
[0070] in and They represent the i-th pixel value of the high-resolution DEM image with 4 times, 2 times and 1 times mapping ratio corresponding to the sample label, respectively. N represents the number of pixels in the high-resolution DEM image. and Respectively represent the i-th pixel value of the super-resolution DEM image with 4 times, 2 times and 1 times mapping ratio predicted by the dual-domain multi-scale attention fusion super-resolution network, || represents the L1 norm;
[0071] Edge pixel level feature loss function L edge The calculation formula is as follows:
[0072]
[0073] in and They represent the i-th pixel value extracted from the high-resolution DEM image corresponding to the sample label using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator, and the Laplace operator, respectively. and They represent the i-th pixel value extracted from the super-resolution DEM image predicted by the dual-domain multi-scale attention fusion super-resolution network using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator and the Laplacian operator respectively. w1, w2 and w3 represent the customized first weight parameter, the second weight parameter and the third weight parameter respectively, and the sum of w1, w2 and w3 is 1.
[0074] The beneficial effects of the present invention are as follows: the present invention comprehensively considers the edge information and spatial continuity characteristics of the DEM image, utilizes the highly sensitive characteristics of the wavelet domain to high and low frequency information, extracts the high frequency and low frequency information of the DEM image through the Haar wavelet transform, and combines the spatial domain feature information obtained by the parallel branch convolution layer, and fuses the two so that the network has the function of feature enhancement at the shallow layer. At the same time, the present invention obtains the importance relationship between feature maps of different depths through a multi-scale attention fusion module group, thereby improving the accuracy of DEM super-resolution reconstruction. In addition, the present invention proposes a multi-scale loss function to supervise the model from the pixel and edge levels, aiming to help the model be more robust and effective. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 It is a flow chart of a remote sensing topography enhancement algorithm of the present invention;
[0076] Figure 2 is a flow chart of the preprocessing of the present invention to construct a training data set;
[0077] Figure 3 is a schematic diagram of a dual-domain multi-scale attention fusion super-resolution network of the present invention;
[0078] Figure 4 is a schematic diagram of a multi-scale attention fusion module group of the present invention;
[0079] Figure 5 It is a comparison diagram of the DEM super-resolution reconstruction results of the present invention;
[0080] Figure 6 This is an enlarged comparison diagram of the DEM super-resolution reconstruction result of the present invention;
[0081] Figure 7 It is a mountain shadow comparison diagram after the DEM super-resolution reconstruction result of the present invention is enlarged;
[0082] Figure 8 It is a comparison diagram of the experimental results of the present invention;
[0083] Fig. 9 It is a comparison diagram of the river network extraction results of the DEM image of the present invention. DETAILED DESCRIPTION
[0084] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that the discussion of these embodiments is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the contents of this specification. Each example may omit, replace or add various processes or components as needed. In addition, the features described relative to some examples may also be combined in other examples.
[0085] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The words "first", "second" and similar words used in one or more embodiments of the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0086] like Figures 1 to 9 As shown, a remote sensing terrain enhancement algorithm includes the following steps:
[0087] Step S101, preprocessing the existing high-resolution DEM image and low-resolution DEM image to construct a training data set;
[0088] The training data set includes G training samples, of which G1 training samples are used for training and G2 training samples are used for testing;
[0089] A training sample includes a sample data and a sample label, wherein the sample data corresponds to the preprocessed low-resolution DEM image, and the sample label corresponds to the preprocessed high-resolution DEM image, wherein there is a 4-fold mapping ratio between the preprocessed low-resolution DEM image and the preprocessed high-resolution DEM image;
[0090] Step S102, constructing a dual-domain multi-scale attention fusion super-resolution network;
[0091] like Figure 3 As shown, the dual-domain multi-scale attention fusion super-resolution network includes: a first convolutional layer, a first reconstruction subnetwork, a second reconstruction subnetwork, a third reconstruction subnetwork, a second convolutional layer, an upsampling layer and a third convolutional layer;
[0092] The first reconstruction sub-network, the second reconstruction sub-network and the third reconstruction sub-network all include: a wavelet guide module (Wavelet Guide Module), a fourth convolutional layer, a multi-scale attention fusion module group (MAFGroup) and a wavelet separation module (Wavelet Separation Module), wherein the multi-scale attention fusion module group is composed of 16 multi-scale attention fusion modules (MAFBlock) cascaded and stacked;
[0093] Step S103, training the dual-domain multi-scale attention fusion super-resolution network using a training data set;
[0094] The input of the dual-domain multi-scale attention fusion super-resolution network corresponds to the sample data of the training samples, and the output is a super-resolution DEM image.
[0095] It should be noted that the output of the DDMAFN model is a super-resolution DEM image, which is infinitely close to the high-resolution DEM image corresponding to the sample label of the training sample.
[0096] In one embodiment of the present invention, the number of training samples G, the number of training samples used for training G1 and the number of training samples used for testing G2 are all custom parameters. The number of training samples G of the present invention is 1055, the number of training samples used for training G1 is 1039, and the number of training samples used for testing G2 is 16.
[0097] In one embodiment of the present invention, the high-resolution DEM image comes from the DAAC Asian High Mountain Dataset (HMA), which is maintained by the National Snow and Ice Data Center (NSIDC) of the United States and is well-known for its reliability and detailed description of the Asian high mountains. The dataset describes in detail the intricate terrain features of the Qinghai-Tibet Plateau and its surrounding areas with a spatial resolution of 8 meters; the low-resolution DEM image comes from the ASTER Global Digital Elevation Model (ASTER GDEM), which is known for its wide coverage and stable data quality. It can provide global elevation information with a spatial resolution of 30 meters. The longitude and latitude range covered by ASTER GDEM is the same as that of HMA, but the resolution is reduced; this combination is intended to simulate the degradation process in the real world and verify the applicability of the model.
[0098] In one embodiment of the present invention, Figure 2 As shown in the figure, the existing high-resolution DEM images and low-resolution DEM images are preprocessed to build a training dataset, including the following steps:
[0099] Step S201, geospatial registration of the high-resolution DEM image and the low-resolution DEM image;
[0100] That is, the high-resolution DEM image and the low-resolution DEM image cover the same geographical area. For example, geographic information system (GIS) software can be used to identify feature points or objects in the image for matching, and adjust the position, rotation and scaling of the image so that the two images can cover the same geographical area and the pixels and geographical coordinates are consistent.
[0101] Step S202, performing projection transformation processing on the high-resolution DEM image and the low-resolution DEM image;
[0102] That is, high-resolution DEM images and low-resolution DEM images are converted to the same map projection system to ensure geometric consistency. The projection system can be UTM, WGS84, etc.;
[0103] Step S203, the high-resolution DEM image is uniformly cropped to a size of 512×512 pixels, and the low-resolution DEM image is uniformly cropped to establish a 4-fold mapping ratio with the high-resolution DEM image.
[0104] That is, the low-resolution DEM images are 512×512 pixels (1 times), 256×256 pixels (2 times), and 128×128 pixels (4 times).
[0105] In one embodiment of the present invention, the first convolution layer inputs a low-resolution DEM image and outputs a first convolution feature map;
[0106] The first reconstruction subnetwork inputs the first convolutional feature map and outputs the first feature map;
[0107] The first convolution feature map output by the first convolution layer and the first feature map output by the first reconstruction subnetwork are fused to obtain a first fused feature map, and then upsampled by bicubic X2 to obtain a first upsampled map;
[0108] The second reconstruction subnetwork inputs the first up-sampled image and outputs the second feature image;
[0109] The first feature map output by the first reconstruction subnetwork is upsampled by the bicubic interpolation method to obtain a second upsampled map, which is then fused with the first convolutional feature map output by the first convolutional layer to obtain a second fused feature map. The second fused feature map is then upsampled by the bicubic interpolation method to obtain a third upsampled map. The third upsampled map is fused with the second feature map output by the second reconstruction subnetwork to obtain a third fused feature map. The third fused feature map is then upsampled by the bicubic interpolation method to obtain a fourth upsampled map.
[0110] The third reconstruction subnetwork inputs the fourth up-sampled image and outputs the third feature map;
[0111] The second convolution layer inputs the third feature map and outputs the second convolution feature map;
[0112] The upsampling layer performs sub-pixel upsampling on the first convolution feature map output by the first convolution layer to obtain a fifth upsampling map, and then fuses the fifth upsampling map with the second convolution feature map output by the second convolution layer to obtain a fourth fused feature map;
[0113] The third convolutional layer inputs the fourth fused feature map and outputs a super-resolution DEM image.
[0114] In one embodiment of the present invention, the first reconstruction subnetwork I Sub_Network1 The calculation formula is as follows:
[0115] F1=I Sub_Network1 (I Conv1 (DEM LR ));
[0116] Among them I Conv1 represents the first convolutional layer, F1 represents the first feature map output by the first reconstruction sub-network, and I Conv1 (DEM LR ) represents the first convolution feature map output by the first convolution layer, DEM LR Represents a low-resolution DEM image;
[0117] Second reconstruction sub-network I Sub_Network2 The calculation formula is as follows:
[0118] F2=I Sub_Network2 (I Bicubic (F1+I Conv1 (DEM LR )));
[0119] Where F2 represents the second feature map output by the second reconstruction sub-network, I Bicubic (F1+I Conv1 (DEM LR )) represents the first upsampling image, F1+I Conv1 (DEM LR ) represents the first fusion feature map, I Bicubic represents bicubic interpolation;
[0120] The third reconstruction sub-network I Sub_Network3 The calculation formula is as follows:
[0121] F3=I Sub_Network3 (I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2));
[0122] Where F3 represents the third feature map output by the third reconstruction sub-network, I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2) represents the fourth upsampling graph, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2 represents the third fusion feature map, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1)) represents the third upsampling graph, I Conv1 (DEM LR )+I Bicubic (F1) represents the second fusion feature map, I Bicubic (F1) represents the second up-sampling graph.
[0123] In one embodiment of the present invention, the calculation formula of the third convolutional layer is as follows:
[0124] DEM SR =I Conv3 (I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR )));
[0125] Among them I Conv2 represents the second convolutional layer, I Conv3 Represents the third convolutional layer, DEM SR represents the super-resolution DEM image output by the third convolutional layer, I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR )) represents the fourth fusion feature map, I Conv2 (F3) represents the second convolutional feature map, I Pixcel (I Conv1 (DEM LR )) represents the fifth upsampling graph, I Pixcel Indicates sub-pixel upsampling.
[0126] It should be noted that in order to avoid the problem of gradient disappearance, the dual-domain multi-scale attention fusion super-resolution network reuses the features extracted from the previous layer and uses global residuals and local residuals to complete the reconstruction task.
[0127] In one embodiment of the present invention, the wavelet guided module includes: a wavelet domain processing module, a fifth convolutional layer and a sixth convolutional layer;
[0128] The wavelet domain processing module decomposes the feature map of the input wavelet guided module into high-frequency components and low-frequency components through discrete wavelet transform (DWT), and then performs dimension reduction processing on it through the 1×1 convolution kernel of the fifth convolution layer to obtain the third convolution feature map;
[0129] The feature map of the input wavelet guided module is extracted by using the convolution kernel of the sixth convolution layer with a size of 3×3 to obtain the fourth convolution feature map;
[0130] The third convolutional feature map and the fourth convolutional feature map are fused to obtain a fourth fused feature map.
[0131] In one embodiment of the present invention, the calculation formula of the wavelet guidance module is as follows:
[0132] F Output =Concat[I Conv6 (F Input ), I Conv5 (I DWT (F Input ))];
[0133] Among them I Conv5 represents the fifth convolutional layer, I Conv6 represents the sixth convolutional layer, F Output represents the fourth fused feature map output by the wavelet-guided module, F Input Represents the feature map of the input wavelet guided module, I DWT represents the wavelet domain processing module, I Conv4 (I DWT (F Input )) represents the third convolutional feature map, I Conv6 (F Input ) represents the fourth convolutional feature map, and Concat represents the concatenation function.
[0134] It should be noted that feature extraction is performed through spatial domain processing (i.e., the convolution kernel of the sixth convolution layer with a size of 3×3), and the feature map is decomposed into high-frequency and low-frequency components through wavelet domain processing (i.e., discrete Hahn wavelet transform), so as to capture the details and structure of the DEM image, wherein the high-frequency component contains the details of the DEM image, and the low-frequency component reflects the overall structure of the DEM image. In addition, in order to reduce the computational complexity and ensure the consistency of the channel size, the wavelet coefficients are subjected to dimensionality reduction processing through the convolution kernel of the fifth convolution layer with a size of 1×1. The present invention integrates the spatial domain and the wavelet domain to enhance the ability of the dual-domain multi-scale attention fusion super-resolution network to represent DEM images.
[0135] It should be noted that a fourth convolution layer is also included between the wavelet guided module and the multi-scale attention fusion module group, which is used to further extract features from the fourth fusion feature map output by the wavelet guided module and then input it into the multi-scale attention fusion module group, which will not be elaborated here.
[0136] In one embodiment of the present invention, Figure 4 As shown in the figure, the multi-scale attention fusion module group is composed of 16 multi-scale attention fusion modules of three different types, which are cascaded and stacked, namely the first module, the second module and the third module, among which the first multi-scale attention fusion module is the third module, and the second to the seventh multi-scale attention fusion modules are stacked in turn in the order of the first module and the second module, and the eighth multi-scale attention fusion module is the first module, and the stacking order of the last eight multi-scale attention fusion modules is the same as the first eight;
[0137] The first module is built based on the channel attention module (CAM);
[0138] After the first convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. After the second convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. This process is repeated to obtain multiple proportion sequences extracted from features at different depths. After weighting, a new proportion sequence is obtained and applied to the original input of the first module as the final output.
[0139] It should be noted that the idea of building this module is to use the importance sequence of different network depth feature maps in the channel direction, jointly weight them to obtain a more reasonable final attention weight sequence and act on the original input as the final output of the current module;
[0140] The channel attention mechanism includes two branches. The first branch obtains the first channel vector of the feature map input to the first module through average pooling calculation. The second branch obtains the second channel vector of the feature map input to the first module through maximum pooling calculation. The first channel vector and the second channel vector are both subjected to nonlinear mapping of two fully connected layers to generate a first weight vector and a second weight vector. The first weight vector and the second weight vector are respectively element-wise multiplied with the feature map input to the first module and then concatenated to obtain a first updated feature map, wherein the first channel vector, the second channel vector, the first weight vector and the second weight vector have the same number of dimensions.
[0141] The second module is built based on the dilated convolutional Inception module (DCIM);
[0142] The Inception module of the dilated convolution includes 3 filters, each of which has a size of 3×3, and the dilation factors are 1, 2, and 3, respectively. The corresponding receptive fields are 3×3, 5×5, and 7×7, respectively. The feature maps input into the second module are extracted through the 3 filters and then concatenated to obtain the second updated feature map.
[0143] It should be noted that the construction idea of the second module is the same as that of the first module. Both use the importance sequence of different network depth feature maps in the channel direction, weight them together as the final attention weight sequence and act on the original input of the current module.
[0144] The third module is built based on the edge-enhanced diverse branch module (EEDB);
[0145] It should be noted that the construction concept of the third module is the same as that of the first module;
[0146] The edge enhancement diversification branch module includes 6 parallel branches. The first parallel branch consists of a 1×1 convolution kernel, the second parallel branch consists of a 3×3 convolution kernel, the third parallel branch consists of a 1×1 convolution kernel and a 3×3 convolution kernel, the fourth parallel branch consists of a 1×1 convolution kernel and a horizontal Sobel edge detection operator (Sobel X), the fifth parallel branch consists of a 1×1 convolution kernel and a vertical Sobel edge detection operator (Sobel Y), and the sixth parallel branch consists of a 1×1 convolution kernel and a Laplacian operator. The feature maps input into the third module are extracted through the 6 parallel branches and then spliced to obtain the third updated feature map.
[0147] It should be noted that, for example, the feature map input to the first module includes C channels, then the dimensions of the first channel vector, the second channel vector, the first weight vector, and the second weight are all C, average pooling means taking the average value of all element values in a single channel feature map, and maximum pooling means taking the maximum value of all element values in a single channel feature map.
[0148] It should be noted that a dilation factor of 1 is actually a 3×3 convolution kernel, and a dilation factor of 2 is actually a 5×5 convolution kernel. However, a zero value is inserted between each element. In fact, there are still only 9 non-zero values, but the receptive field has been expanded to 5×5. Similarly, a dilation factor of 3 is actually a 7×7 convolution kernel, but 2 zero values are inserted between each element. In fact, there are still 9 non-zero values, but the receptive field has been expanded to 7×7. DCIM can capture features and details of different scales, thereby enhancing the ability to extract multi-scale features. The parallel application of these dilated convolutions enables the second module to focus on the features of different perceptual areas of the DEM image, extract multi-scale features in a forward pass, and capture different image details and complexities.
[0149] It should be noted that the 1×1 convolution kernel is used to adjust the number of channels and increase nonlinearity, the 3×3 convolution kernel is used to capture a wider range of spatial information, the Sobel X and Sobel Y operators are used to extract the edge features of the feature map in the horizontal and vertical directions respectively, and the Laplacian operator is used to extract the second-order edge information and emphasize the detail changes. The feature fusion of 6 parallel branches provides richer edge-sensitive features for multi-scale attention fusion, which enables the network to utilize multiple convolution operation scales and edge detection operators to improve the sensitivity and representativeness of edge features, thereby improving the performance of super-resolution tasks.
[0150] In one embodiment of the present invention, the wavelet separation module includes: a seventh convolution layer, an inverse wavelet transform module (Haar wavelet inverse transform function), an eighth convolution layer and an interpolation processing layer (bicubic interpolation method);
[0151] The feature map of the input wavelet separation module is reconstructed through the seventh convolution layer and the wavelet inverse transform module to obtain the fifth convolution feature map;
[0152] The feature map of the input wavelet separation module is reconstructed through the eighth convolution layer and the interpolation processing layer to obtain the sixth convolution feature map;
[0153] The fifth convolutional feature map and the sixth convolutional feature map are fused to obtain a fifth fused feature map.
[0154] In one embodiment of the present invention, the calculation formula of the wavelet separation module is as follows:
[0155] F Output =I ADD (I Bicubic (I Conv8 (F Input )),I Conv7 (I IWT (F Input )));
[0156] Among them I Conv7 represents the seventh convolutional layer, I Conv8 represents the eighth convolutional layer, F Output represents the fifth fusion feature map output by the wavelet separation module, F Input Represents the feature map of the input wavelet separation module, I Bicubic represents the interpolation processing layer, I IWT represents the inverse wavelet transform module, I ADD Represents the addition function.
[0157] It should be noted that the existing DEM image super-resolution reconstruction methods mainly rely on interpolation or a single convolution operation. Although the performance has been improved, these methods are insufficient in capturing the global structure and detail information of the image. Therefore, the wavelet separation module provided by the present invention includes two branches, one of which uses an inverse wavelet transform (IWT) to reconstruct features in the wavelet domain to ensure the accurate restoration of the wavelet features. Because the wavelet transform has excellent multi-scale analysis capabilities, it can capture the detail information of the image at different frequencies, thereby retaining the key structural information of the DEM image during the super-resolution reconstruction process. The other branch reconstructs features in the spatial domain through 3×3 convolution and bicubic upsampling, and maintains the data dimension consistent with the output of the wavelet domain. Finally, the outputs of the two branches are fused together to restore and enhance the feature information of the image from a spatial perspective, thereby improving the quality of DEM image reconstruction.
[0158] In one embodiment of the present invention, the loss function L of the dual-domain multi-scale attention fusion super-resolution network is MS The calculation formula includes:
[0159] L MS =α×L Pixcel +β×L edge ;
[0160] Where L Pixcel Represents the loss function for reconstructing pixel-level features, L edge represents the edge pixel level feature loss function, α and β represent L Pixcel and L edge The corresponding custom weight coefficient, and the sum of α and β is 1;
[0161] Reconstruct pixel-level feature loss function L Pixcel The calculation formula is as follows:
[0162]
[0163] in and They represent the i-th pixel value of the high-resolution DEM image with 4 times, 2 times and 1 times mapping ratio corresponding to the sample label, respectively. N represents the number of pixels in the high-resolution DEM image. and Respectively represent the i-th pixel value of the super-resolution DEM image with 4 times, 2 times and 1 times mapping ratio predicted by the dual-domain multi-scale attention fusion super-resolution network, || represents the L1 norm;
[0164] Edge pixel level feature loss function L edge The calculation formula is as follows:
[0165]
[0166] in and They represent the i-th pixel value extracted from the high-resolution DEM image corresponding to the sample label using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator, and the Laplace operator, respectively. and They represent the i-th pixel value extracted from the super-resolution DEM image predicted by the dual-domain multi-scale attention fusion super-resolution network using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator and the Laplacian operator respectively. w1, w2 and w3 represent the customized first weight parameter, the second weight parameter and the third weight parameter respectively, and the sum of w1, w2 and w3 is 1.
[0167] In one embodiment of the present invention, Figure 8 As shown, the DEM super-resolution reconstruction is performed using a variety of models using the training data set provided by the present invention. The PNSR (peak signal-to-noise ratio), RMSE (peak signal-to-noise ratio), MAE (mean absolute error), and RMSE of the DDMAFN model provided by the present invention are shown in FIG. slope (Root Mean Square Error of Slope) and RMSE aspect The root mean square error of the slope aspect is 49.99, 13.03, 10.36, 4.15° and 111.47° respectively. Compared with other models, the DDMAFN model provided by the present invention has smaller error and higher accuracy in DEM super-resolution reconstruction.
[0168] In one embodiment of the present invention, the low-resolution DEM image is converted into a super-resolution DEM image by the DDMAFN model provided by the present invention, and then the river network is extracted from the generated super-resolution DEM image, such as Fig. 9As shown, when extracting river networks using low-resolution images, there is often a large amount of detailed information lost, while the river network structure extracted by the DDMAFN model provided by the present invention is highly similar to the result obtained from the original high-resolution DEM image and is consistent in detail representation.
[0169] The above describes an embodiment of the present embodiment, but the present embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present embodiment, ordinary technicians in this field can also make many forms, all of which are within the protection of the present embodiment.
Claims
1. A remote sensing terrain enhancement algorithm, characterized in that: The following steps are involved: Step S101, preprocessing the existing high-resolution DEM image and low-resolution DEM image to construct a training data set; The training data set includes G training samples, of which G1 training samples are used for training and G2 training samples are used for testing; A training sample includes a sample data and a sample label, wherein the sample data corresponds to the preprocessed low-resolution DEM image, and the sample label corresponds to the preprocessed high-resolution DEM image, wherein there is a 4-fold mapping ratio between the preprocessed low-resolution DEM image and the preprocessed high-resolution DEM image; Step S102, constructing a dual-domain multi-scale attention fusion super-resolution network; The dual-domain multi-scale attention fusion super-resolution network includes: a first convolutional layer, a first reconstruction sub-network, a second reconstruction sub-network, a third reconstruction sub-network, a second convolutional layer, an upsampling layer, and a third convolutional layer; The first reconstruction subnetwork, the second reconstruction subnetwork and the third reconstruction subnetwork all include: a wavelet guidance module, a fourth convolutional layer, a multi-scale attention fusion module group and a wavelet separation module, wherein the multi-scale attention fusion module group is formed by cascading and stacking 16 multi-scale attention fusion modules; Step S103, training the dual-domain multi-scale attention fusion super-resolution network using a training data set; The input of the dual-domain multi-scale attention fusion super-resolution network corresponds to the sample data of the training sample, and the output is a super-resolution DEM image; The number of training samples G, the number of training samples used for training G1, and the number of training samples used for testing G2 are all custom parameters.
2. A remote sensing terrain enhancement algorithm according to claim 1, characterized in that: The high-resolution DEM images come from the DAAC Asian high mountain dataset, and the low-resolution DEM images come from the ASTER global digital elevation model. The longitude and latitude range covered by the low-resolution DEM images is the same as that of the high-resolution DEM images.
3. The remote sensing terrain enhancement algorithm according to claim 1 is characterized in that: The existing high-resolution DEM images and low-resolution DEM images are preprocessed to construct a training dataset, including the following steps: Step S201, geospatial registration of the high-resolution DEM image and the low-resolution DEM image; Step S202, performing projection transformation processing on the high-resolution DEM image and the low-resolution DEM image; Step S203, the high-resolution DEM image is uniformly cropped to a size of 512×512 pixels, and the low-resolution DEM image is uniformly cropped to establish a 4-fold mapping ratio with the high-resolution DEM image.
4. The remote sensing terrain enhancement algorithm according to claim 1, characterized in that: The first convolution layer inputs a low-resolution DEM image and outputs the first convolution feature map; The first reconstruction subnetwork inputs the first convolutional feature map and outputs the first feature map; The first convolution feature map output by the first convolution layer and the first feature map output by the first reconstruction subnetwork are fused to obtain a first fused feature map, and then upsampled by bicubic interpolation to obtain a first upsampled map; The second reconstruction subnetwork inputs the first up-sampled image and outputs the second feature image; The first feature map output by the first reconstruction subnetwork is upsampled by the bicubic interpolation method to obtain a second upsampled map, which is then fused with the first convolutional feature map output by the first convolutional layer to obtain a second fused feature map. The second fused feature map is then upsampled by the bicubic interpolation method to obtain a third upsampled map. The third upsampled map is fused with the second feature map output by the second reconstruction subnetwork to obtain a third fused feature map. The third fused feature map is then upsampled by the bicubic interpolation method to obtain a fourth upsampled map. The third reconstruction subnetwork inputs the fourth up-sampled image and outputs the third feature map; The second convolution layer inputs the third feature map and outputs the second convolution feature map; The upsampling layer performs sub-pixel upsampling on the first convolution feature map output by the first convolution layer to obtain a fifth upsampling map, and then fuses the fifth upsampling map with the second convolution feature map output by the second convolution layer to obtain a fourth fused feature map; The third convolutional layer inputs the fourth fused feature map and outputs a super-resolution DEM image.
5. A remote sensing topography enhancement algorithm according to claim 4, characterized in that: The first reconstruction sub-network I Sub_Network1 The calculation formula is as follows: F1=I Sub_Network1 (IN Conv1 (THEM LR )); Among them I Conv1 represents the first convolutional layer, F1 represents the first feature map output by the first reconstruction sub-network, and I Conv1 (DEM LR ) represents the first convolution feature map output by the first convolution layer, DEM LR Represents a low-resolution DEM image; Second reconstruction sub-network I Sub_Network2 The calculation formula is as follows: F2=I Sub_Network2 (I Bicubic (F1+I Conv1 (DEM LR ))); Where F2 represents the second feature map output by the second reconstruction sub-network, I Bicubic (F1+I Conv1 (DEM LR )) represents the first upsampling image, F1+I Conv1 (DEM LR ) represents the first fusion feature map, I Bicubic represents bicubic interpolation; The third reconstruction sub-network I Sub_Network3 The calculation formula is as follows: F3=I Sub_Network3 (I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2)); Where F3 represents the third feature map output by the third reconstruction sub-network, I Bicubic (I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2) represents the fourth upsampling graph, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1))+F2 represents the third fusion feature map, I Bicubic (I Conv1 (DEM LR )+I Bicubic (F1)) represents the third upsampling graph, I Conv1 (DEM LR )+I Bicubic (F1) represents the second fusion feature map, I Bicubic (F1) represents the second up-sampling graph.
6. A remote sensing terrain enhancement algorithm according to claim 4, characterized in that: The calculation formula of the third convolutional layer is as follows: DEM SR =I Conv3 (I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR ))); Among them I Conv2 represents the second convolutional layer, I Conv3 Represents the third convolutional layer, DEM SR represents the super-resolution DEM image output by the third convolutional layer, I Conv2 (F3)+I Pixcel (I Conv1 (DEM LR )) represents the fourth fusion feature map, I Conv2 (F3) represents the second convolutional feature map, I Pixcel (I Conv1 (DEM LR )) represents the fifth upsampling graph, I Pixcel Indicates sub-pixel upsampling.
7. The remote sensing terrain enhancement algorithm according to claim 1, characterized in that: The wavelet guided module includes: a wavelet domain processing module, a fifth convolution layer and a sixth convolution layer; The wavelet domain processing module decomposes the feature map of the input wavelet guidance module into high-frequency components and low-frequency components through discrete wavelet transform, and then performs dimensionality reduction processing on it through the 1×1 convolution kernel of the fifth convolution layer to obtain the third convolution feature map; The fourth convolution feature map is obtained by extracting the feature map of the input wavelet guided module through the convolution kernel of the sixth convolution layer with a size of 3×3; Fusing the third convolution feature map and the fourth convolution feature map to obtain a fourth fused feature map; The calculation formula of the wavelet guided module is as follows: F Output =Concat[I Conv6 (F Input ),I Conv5 (I DWT (F Input ))]; Among them I Conv5 represents the fifth convolutional layer, I Conv6 represents the sixth convolutional layer, F Output represents the fourth fused feature map output by the wavelet-guided module, F Input Represents the feature map of the input wavelet guided module, I DWT represents the wavelet domain processing module, I Conv4 (I DWT (F Input )) represents the third convolutional feature map, I Conv6 (F Input ) represents the fourth convolutional feature map, and Concat represents the concatenation function.
8. The remote sensing terrain enhancement algorithm according to claim 1 is characterized in that: The multi-scale attention fusion module group is composed of 16 multi-scale attention fusion modules of three different types, which are cascaded and stacked, namely the first module, the second module and the third module. The first multi-scale attention fusion module is the third module, and the second to seventh multi-scale attention fusion modules are stacked in turn in the order of the first module and the second module. The eighth multi-scale attention fusion module is the first module, and the last eight multi-scale attention fusion modules are stacked in the same order as the first eight. The first module is built based on the channel attention module; After the first convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. After the second convolutional layer extracts the feature map, the channel attention mechanism is used to obtain the proportion sequence of each feature map in the channel direction. This process is repeated to obtain multiple proportion sequences extracted from features at different depths. After weighting, a new proportion sequence is obtained and applied to the original input of the first module as the final output. The channel attention mechanism includes two branches. The first branch obtains the first channel vector of the feature map input to the first module through average pooling calculation. The second branch obtains the second channel vector of the feature map input to the first module through maximum pooling calculation. The first channel vector and the second channel vector are both subjected to nonlinear mapping of two fully connected layers to generate a first weight vector and a second weight vector. The first weight vector and the second weight vector are respectively element-wise multiplied with the feature map input to the first module and then concatenated to obtain a first updated feature map, wherein the first channel vector, the second channel vector, the first weight vector and the second weight vector have the same number of dimensions. The second module is built on the Inception module based on dilated convolution; The Inception module of the dilated convolution includes 3 filters, each of which has a size of 3×3, and the dilation factors are 1, 2, and 3, respectively. The corresponding receptive fields are 3×3, 5×5, and 7×7, respectively. The feature maps input into the second module are extracted through the 3 filters and then concatenated to obtain the second updated feature map. The third module is built based on the edge-enhanced diversified branch module; The edge enhancement diversification branch module includes 6 parallel branches. The first parallel branch consists of a 1×1 convolution kernel, the second parallel branch consists of a 3×3 convolution kernel, the third parallel branch consists of a 1×1 convolution kernel and a 3×3 convolution kernel, the fourth parallel branch consists of a 1×1 convolution kernel and a horizontal Sobel edge detection operator, the fifth parallel branch consists of a 1×1 convolution kernel and a vertical Sobel edge detection operator, and the sixth parallel branch consists of a 1×1 convolution kernel and a Laplacian operator. The feature maps input into the third module are extracted through the 6 parallel branches and then spliced to obtain the third updated feature map.
9. The remote sensing terrain enhancement algorithm according to claim 1, characterized in that: The wavelet separation module includes: the seventh convolution layer, the wavelet inverse transform module, the eighth convolution layer and the interpolation processing layer; The feature map of the input wavelet separation module is reconstructed through the seventh convolution layer and the wavelet inverse transform module to obtain the fifth convolution feature map; The feature map of the input wavelet separation module is reconstructed through the eighth convolution layer and the interpolation processing layer to obtain the sixth convolution feature map; Fusing the fifth convolution feature map and the sixth convolution feature map to obtain a fifth fused feature map; The calculation formula of the wavelet separation module is as follows: F Output =I ADD (I Bicubic (I Conv8 (F Input )),I Conv7 9I IWT (F Input ))); Among them I Conv7 represents the seventh convolutional layer, I Conv8 represents the eighth convolutional layer, F Output represents the fifth fusion feature map output by the wavelet separation module, F Input Represents the feature map of the input wavelet separation module, I Bicubic represents the interpolation processing layer, I IWT represents the inverse wavelet transform module, I ADD Represents the addition function.
10. The remote sensing terrain enhancement algorithm according to claim 8, characterized in that: The loss function L of the dual-domain multi-scale attention fusion super-resolution network MS The calculation formula include: L MS =α×L Pixcel +β×L edge ; Where L Pixcel Represents the loss function for reconstructing pixel-level features, L edge represents the edge pixel level feature loss function, α and β represent L Pixcel and L edge The corresponding custom weight coefficient, and the sum of α and β is 1; Reconstruct pixel-level feature loss function L Pixcel The calculation formula is as follows: in and They represent the i-th pixel value of the high-resolution DEM image with 4 times, 2 times and 1 times mapping ratio corresponding to the sample label, respectively. N represents the number of pixels in the high-resolution DEM image. and Respectively represent the i-th pixel value of the super-resolution DEM image with 4 times, 2 times and 1 times mapping ratio predicted by the dual-domain multi-scale attention fusion super-resolution network, || represents the L1 norm; Edge pixel level feature loss function L edge The calculation formula is as follows: in and They represent the i-th pixel value extracted from the high-resolution DEM image corresponding to the sample label using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator, and the Laplace operator, respectively. and They represent the i-th pixel value extracted from the super-resolution DEM image predicted by the dual-domain multi-scale attention fusion super-resolution network using the horizontal Sobel edge detection operator, the vertical Sobel edge detection operator and the Laplacian operator respectively. w1, w2 and w3 represent the customized first weight parameter, the second weight parameter and the third weight parameter respectively, and the sum of w1, w2 and w3 is 1.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on deep convolutional neural network
CN113222819A
Cited By
Lightweight target detection method and device for remote sensing image, equipment and medium
CN120279260A
Feature matching method based on differential wavelet transform and dynamic multi-scale feature fusion
CN120279288A
Remote sensing image panchromatic sharpening method based on spatial-spectral information enhancement
CN120689243A
Gradual super-resolution reconstruction method and device based on wavelet-space cooperation
CN121724838A
Wavelet-space collaborative based progressive super-resolution reconstruction method and device
CN121724838B