DEM (Digital Elevation Model) super-resolution reconstruction method and system combined with depth map

By combining the DEM super-resolution reconstruction method with depth map, using multimodal data sets and multimodal fusion super-resolution networks, the shortcomings of the DEM super-resolution reconstruction method in the prior art in multi-source information fusion and complex terrain feature recovery are solved, and a higher precision DEM reconstruction is achieved.

CN119991440APending Publication Date: 2025-05-13Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510048357.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing DEM super-resolution reconstruction methods have shortcomings in multi-source information fusion and complex terrain feature recovery, and it is difficult to effectively improve the resolution and reconstruction accuracy of DEM.

Method used

A DEM super-resolution reconstruction method combining depth maps is proposed. By constructing a multimodal data set and designing a multimodal fusion super-resolution network, using the depth map generated by optical images as prior knowledge, guiding the DEM super-resolution reconstruction task and improving reconstruction accuracy.

Benefits of technology

It significantly improves the accuracy and accuracy of DEM reconstruction, especially in the details recovery of complex terrain areas, enhances the recovery ability of details, and provides higher quality super-resolution DEM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991440A_ABST
    Figure CN119991440A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of topographic mapping, in particular to a DEM super-resolution reconstruction method and system combined with a depth map, and the method comprises the steps: constructing a multi-modal data set which comprises an optical image, a depth map, a low-resolution DEM image and a high-resolution DEM image which are paired; designing a multi-modal fusion super-resolution network; designing a pseudo Siamese feature extraction module to extract low-level features and high-level features of the optical image, the depth map and the low-resolution DEM image; fusing the low-layer feature map generated by the pseudo Siamese feature extraction module in a shallow layer through a multi-scale feature fusion module, and fusing the deep-layer feature map in a deep layer; outputting a high-resolution DEM image through a DEM reconstruction module; and training the multi-modal fusion super-resolution network by using the multi-modal data set and the collaborative loss function. According to the method, the depth map generated by the optical image is used as priori knowledge to guide the DEM super-resolution reconstruction task based on the basic depth large model, a complex fine adjustment process is not needed, and the DEM reconstruction precision and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of topographic surveying and mapping, and in particular to a DEM super-resolution reconstruction method and system combined with a depth map. Background Art

[0002] Digital Elevation Model (DEM) not only plays a key role in practical applications such as engineering construction, disaster management and ecological monitoring, but is also a core means of understanding the geographical environment and perceiving the spatial relationship of the earth. High-resolution DEM data can capture the subtle structure of the surface in detail and reveal the complex relationship between terrain changes, geological structures and hydrological processes. High precision makes it possible to conduct detailed spatial analysis in multiple scientific fields. However, NASA DEM, one of the most widely used DEM products in the world, has a low resolution of only 30 meters. This introduces significant uncertainty when applying higher spatial resolution images (Tapete et al., 2021). Therefore, developing reliable technical methods to improve the resolution of DEM is not only the research frontier of geographic information science (GIS) and remote sensing (RS), but also an important means for humans to fully perceive and explore the natural environment of the earth and finely analyze spatial relationships. It has a far-reaching impact on the theoretical research and application expansion of earth science.

[0003] Early DEM super-resolution studies mainly relied on methods based on interpolation (Thvenaz et al., 2000) and filtering (Protter et al., 2009). These methods are simple in structure and may cause problems such as insufficient image sharpness and distortion. To improve this problem, subsequent studies introduced methods based on statistics and machine learning (Ha et al., 2018). The image reconstruction effect is enhanced by artificially extracting features, but it is limited by the empirical and flexibility of feature extraction. With the advancement of computing power and big data, deep learning methods have been widely used in single image super-resolution (SISR) research. Deep learning can achieve more accurate image classification, prediction, and generation through its powerful feature extraction capabilities. Super-resolution methods for natural images can be divided into three main categories: models based on convolutional neural networks (CNNs), such as SRCNN (Chao et al., 2015) and EDSR (Lim et al., 2017); models based on generative adversarial networks (GANs), such as SRGAN (Christian et al., 2017) and ESRGAN (Wang et al., 2018); and Transformer-based models, such as SwinIR (Liang et al., 2021) and SRFormer (Zhou et al., 2023). Most of these models have recently been directly applied to DEM super-resolution, but they often fail to achieve optimal performance.

[0004] To overcome the above limitations, multimodal fusion technology has gradually become a research hotspot. High-resolution optical images have become a powerful supplement to DEM super-resolution tasks due to their rich ground details. In recent years, the Depth Anything Model (DAM) (Yang et al., 2024) has been introduced as a core tool for monocular depth estimation. It generates relative depth maps through unsupervised learning and demonstrates good zero-shot learning and task adaptation capabilities. The DAM with zero-shot depth estimation is used to generate relative depth maps of optical images to provide intuitive elevation information. Although multimodal fusion technology has shown the potential to improve DEMSR, its development still faces challenges. Existing datasets lack systematic multimodal organization and are difficult to support the joint application of optical images and depth maps. In addition, standardized deep learning methods for multimodal fusion have not yet formed a unified framework, which limits the in-depth development of DEM super-resolution research. It can be seen that the existing technology still has shortcomings in efficiently fusing multimodal data and fully capturing complex terrain features, which provides an important research direction for DEM super-resolution reconstruction methods combined with depth maps. Summary of the invention

[0005] The present invention aims to solve the shortcomings of existing DEM super-resolution reconstruction methods in multi-source information fusion and complex terrain feature restoration, and proposes a DEM super-resolution reconstruction method and system combined with a depth map. By using the depth map generated by optical images as prior knowledge, the DEM super-resolution reconstruction task is guided, thereby improving the DEM reconstruction accuracy, especially in the detail restoration of complex terrain.

[0006] To achieve the above purpose, the technical solution adopted is:

[0007] The present invention provides a DEM super-resolution reconstruction method combined with a depth map, comprising:

[0008] Construct a multimodal dataset, which includes paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images;

[0009] Design a multimodal fusion super-resolution network: Design a pseudo-Siamese feature extraction module to extract low-level and high-level features of optical images, depth maps, and low-resolution DEM images; then use a multi-scale feature fusion module to fuse the low-level feature maps generated by the pseudo-Siamese feature extraction module at the shallow layer and the deep-level feature maps at the deep layer; output a high-resolution DEM image through a DEM reconstruction module;

[0010] The multimodal fusion super-resolution network is trained using multimodal datasets and collaborative loss functions.

[0011] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the process of constructing a multimodal data set is:

[0012] Resampling: Optical images were resampled to a ground resolution of 5 m using bilinear interpolation; the acquired 30 m resolution DEM was resampled to 50 m using bilinear interpolation;

[0013] Reprojection: Convert DEM images and optical images into the same WGS-84 coordinate system;

[0014] Registration: The RIFT algorithm is used to geo-register the DEM image with the optical image;

[0015] Clipping: Clip all the grids according to the maximum inscribed rectangle of the overlapping area obtained after processing;

[0016] Dataset segmentation: The high-resolution DEM images, optical images and corresponding depth maps are segmented into blocks of 256×256 pixels, and the low-resolution DEM images are segmented into blocks of 64×64 pixels.

[0017] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the pseudo-Siamese feature extraction module uses independent convolution flows with the same structure to extract features from the optical image, the depth map and the low-resolution DEM image, uses ResNet-101 as the backbone network, obtains the low-level features of optics, depth and DEM, and takes the deepest feature map of the backbone network as the high-level feature.

[0018] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the multi-scale feature fusion module specifically includes: for low-level features, the low-level features of optics, depth, and DEM are spliced ​​in the channel dimension to obtain a low-level fusion feature map; for high-level features, the high-level features of optics, depth, and DEM are spliced ​​in the channel dimension and then passed through a dilated spatial convolution pooling pyramid and up-sampling to obtain a high-level fusion feature map.

[0019] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the atrous spatial convolution pooling pyramid includes a 1×1 convolution and multiple parallel atrous convolutions, and features of different receptive fields are obtained through atrous convolutions with different expansion rates, and global information is obtained through global pooling and 1×1 convolution, and finally these features are fused.

[0020] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the DEM reconstruction module includes: splicing the low-level fusion feature map obtained by the multi-scale feature fusion module with the high-level fusion feature map in the channel dimension, and adding the splicing result to the low-resolution DEM original feature information after passing two layers of convolution, upsampling and two ResGs in sequence to output a high-resolution DEM image.

[0021] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the ResG is composed of multiple residual blocks and short jump connections, and the ResG includes several RCAB modules, and the RCAB module introduces a channel attention mechanism.

[0022] According to the DEM super-resolution reconstruction method combined with the depth map of the present invention, further, the collaborative loss function includes the elevation loss L1 and the slope loss RMSE, and the weight of each loss is adjusted by a dynamic weighting strategy. The total loss function expression is as follows:

[0023]

[0024] in, is the weighted parameter of self-learning, L τ (W) represents various losses.

[0025] Furthermore, the present invention also proposes a DEM super-resolution reconstruction system combined with a depth map, which is used to implement the DEM super-resolution reconstruction method combined with a depth map as described above, comprising:

[0026] The dataset construction module is used to construct a multimodal dataset, which contains paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images;

[0027] Model building module, used to design a multimodal fusion super-resolution network: design a pseudo-Siamese feature extraction module to extract low-level and high-level features of optical images, depth maps and low-resolution DEM images; then use the multi-scale feature fusion module to fuse the low-level feature maps generated by the pseudo-Siamese feature extraction module at the shallow layer and the deep-level feature maps at the deep layer; output a high-resolution DEM image through the DEM reconstruction module;

[0028] The model training module is used to train the multimodal fusion super-resolution network using multimodal datasets and collaborative loss functions.

[0029] The beneficial effects achieved by adopting the above technical solution are:

[0030] 1. Improve the accuracy and precision of DEM reconstruction: The present invention effectively supplements the information loss in the traditional DEM reconstruction method by introducing the depth map generated by optical images as prior knowledge, especially in the recovery of details in complex terrain areas. The depth map can intuitively reflect the height changes of the terrain, provide rich terrain feature information for DEM super-resolution reconstruction, and significantly improve the accuracy of the reconstruction results. By introducing a multimodal fusion strategy, the multimodal fusion super-resolution network (MFSR) can make full use of the terrain information in the optical image and depth map, thereby enhancing the DEM super-resolution model's ability to restore details, especially in areas with complex terrain features such as rivers and valleys, showing significant advantages in detail recovery. Specifically, the MFSR model extracts multimodal features from optical images, depth maps, and DEM images through a multi-branch pseudo-twin architecture. The network adopts a hierarchical fusion strategy to gradually fuse multimodal information to generate a higher quality super-resolution DEM. The multi-branch structure in the network can effectively avoid interference between different modal data and fully explore and utilize the terrain information of each modality.

[0031] 2. Multimodal datasets provide new resources for DEM research: By constructing the DEM-OPT-Depth SR dataset, a new heterogeneous dataset is provided for DEM super-resolution research. The dataset consists of optical images, depth maps, and DEM images, covering three terrain types: the mountainous areas in western Sichuan, the Loess Plateau, and the Hengduan Mountains. It contains rich terrain features and provides a reliable data basis for the research of DEM super-resolution technology. This dataset has broad application prospects and is suitable for DEM super-resolution tasks in different geographical regions and terrain types.

[0032] 3. Strong generalization and adaptability: Through experimental verification of multiple data sets, the MFSR model has demonstrated good generalization ability and can adapt to DEM reconstruction tasks with different terrain features and different resolutions. Whether it is mountainous areas, the Loess Plateau, or complex terrain areas such as the Hengduan Mountains, MFSR can effectively integrate multimodal information and restore high-precision DEMs, which has strong practical application value.

[0033] 4. Simplify the complexity and computational cost of model training: The method of the present invention can directly apply the depth map generated by optical images to the DEM super-resolution reconstruction task without complex fine-tuning process, which simplifies the training and application process of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention, wherein the drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.

[0035] Figure 1 It is a schematic flow chart of a DEM super-resolution reconstruction method combined with a depth map according to an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of generating a relative depth map from an optical image based on DAM according to an embodiment of the present invention;

[0037] Figure 3 is a schematic diagram of a process of constructing a multimodal dataset according to an embodiment of the present invention;

[0038] Figure 4 is a diagram of a multimodal fusion super-resolution network structure according to an embodiment of the present invention;

[0039] Figure 5 is the RMSE evaluation result of different methods in the embodiment of the present invention on the test set;

[0040] Figure 6 is a comparison of error distributions of different DEMSR models of the embodiments of the present invention;

[0041] Figure 7 These are DEM results generated by different methods according to the embodiments of the present invention. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings of specific embodiments of the present invention to clearly and completely describe the exemplary scheme of the embodiment of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should be the common meanings understood by people with ordinary skills in the field.

[0043] This embodiment discloses a DEM super-resolution reconstruction method combined with a depth map, aiming to make full use of three modal data: optical image, relative depth map and DEM, and design a multimodal fusion super-resolution network (MFSR) to achieve high-precision reconstruction of DEM. The implementation process is as follows: Figure 1 As shown, it specifically includes the following contents:

[0044] Step S101: construct a multimodal dataset, which includes paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images.

[0045] This embodiment uses a basic deep model Depth Anything Model (DAM) that does not require additional training to process the input high-resolution optical image to generate a relative depth map, such as Figure 2As shown in Figure 1, the optical image, relative depth map and DEM image are paired to construct a data set containing trimodal information. In order to ensure the consistency and comparability of the data, the acquired data needs to be processed as follows: Figure 3 The preprocessing operations shown.

[0046] ①Resampling

[0047] In order to facilitate the convolution operation, the optical image was resampled to a ground resolution of 5m using bilinear interpolation, and then a relative depth map was generated from the 5m resolution optical image. The DEM super-resolution was adjusted to 4×, and 12.5m and 30m high-resolution DEMs were downloaded. The 30m resolution DEM was sampled to 50m using bilinear interpolation to obtain a low-resolution DEM. The 5m resolution optical image and relative depth map were downsampled to 50m and input into the network model together with the 50m resolution DEM.

[0048] ② Reprojection

[0049] Since the DEM images and remote sensing images used are from different sources, they need to be unified in projection and reference format. Therefore, they are converted to the WGS-84 coordinate system to ensure that the data is processed and analyzed in the same geographic coordinate system.

[0050] ③Registration

[0051] The RIFT algorithm is used to geo-reference the DEM image and the optical image. Through geo-reference, the geometric transformation relationship of the image to be referenced relative to the reference image can be estimated. Finally, pixel-level alignment is achieved through the calculated geometric transformation relationship to ensure their spatial consistency and accuracy.

[0052] ④ Cutting

[0053] All rasters are clipped according to the maximum inscribed rectangle of the overlapping area obtained after processing. This can make the clipped data set more regular and facilitate subsequent pixel segmentation and processing.

[0054] ⑤Dataset segmentation

[0055] The HR DEM (12.5m) and optical image tiff and the corresponding depth map were cut into blocks of 256×256 pixels, and the low-resolution DEM image was cut into blocks of 64×64 pixels, thus obtaining the final paired data. When dividing the training set and the test set, the ratio of 8:2 was used.

[0056] The multimodal dataset proposed in this embodiment, DEM-OPT-Depth SR, pairs optical images, depth maps, and DEMs, becoming the first heterogeneous benchmark dataset for DEM super-resolution tasks (access link https: / / doi.org / 10.6084 / m9.figshare.26409490.v3). The dataset covers three representative terrain types: the mountainous areas in western Sichuan, the Loess Plateau, and the Hengduan Mountains, covering a wealth of terrain features and providing rich training and test data for DEM super-resolution tasks. The dataset has been open for sharing, providing valuable resources for research in the field of geographic information fusion.

[0057] Step S102, design a multimodal fusion super-resolution (MFSR) network: design a pseudo-Siamese feature extraction module to extract low-level and high-level features of optical images, depth maps and low-resolution DEM images; then use a multi-scale feature fusion module to fuse the low-level feature maps generated by the pseudo-Siamese feature extraction module at a shallow layer and fuse the deep-level feature maps at a deep layer; and output a high-resolution DEM image through a DEM reconstruction module.

[0058] In order to solve the problem of multimodal information fusion in the DEM super-resolution reconstruction task, a multimodal fusion super resolution network (MFSR) is proposed. The network extracts multimodal data features through a multi-branch pseudo-twin architecture and uses a hierarchical fusion strategy to generate high-resolution DEM. MFSR includes a pseudo-Siamese feature extraction module, a multi-scale feature fusion module and a DEM reconstruction module:

[0059] ① Design of Pseudo-Siamese feature extraction module

[0060] In order to extract effective features of optical images, depth maps and DEM images, a pseudo-Siamese feature extraction module is designed. This module uses independent convolutional flows with the same structure to extract features from the three-modal data. The network structure is as follows: Figure 4 (a) (taking optical image and DEM image as input). This part contains three independent branches, which extract features from optical image, depth map and DEM image input respectively, and obtain high-level and low-level features of the three types of data for subsequent analysis. This embodiment uses ResNet-101 as the backbone network, takes a shallower layer of feature map as the low-level feature, and the low-level features of optical and DEM are denoted as F low OPT and F low DEM ; Take the deepest feature map of the backbone network as the high-level feature. Correspondingly, the high-level features of the two types of data are recorded as Fhigh OPT and F high DEM .

[0061] ② Multi-scale feature fusion module design

[0062] The shallow feature map generated by the pseudo-Siamese network feature extraction module is fused with the deep feature map respectively, such as Figure 4 As shown in (a), the backbone network features are spliced ​​at two feature levels and the cross-modal features are obtained based on the backbone network features. The two feature maps obtained are the input of the DEM reconstruction module. The specific process is:

[0063] For low-level features, F low OPT and F low DEM The size of is 64x64x256, and F is obtained by the feature fusion module low OPT-DEM The size of is 64x64x48. Specifically, after the two types of data are spliced ​​in the channel dimension at the same low level, the features are enhanced according to the attention mechanism. The Q, K, and V obtained by the cross-modal attention are calculated by 1x1 convolution, and then reduced to 64x64x48 by 1x1 convolution. The feature map after dimensionality reduction is recorded as F low OPT-DEM .

[0064] For high-level features, first perform an F with a size of 32x32x256. high OPT and F high DEM After channel dimension concatenation, the atrous spatial pyramid pooling (ASPP) is used to reduce the dimension to 256, and finally the F with a size of 64x64x256 is obtained by upsampling. high OPT-DEM The ASPP structure contains 1×1 convolution and multiple parallel dilated convolutions. The features of different receptive fields are obtained through dilated convolutions with different expansion rates, and global information is obtained through global pooling and 1×1 convolution. Finally, these features are integrated to enhance the expressiveness of the model.

[0065] ③DEM reconstruction module design

[0066] In order to optimize the fusion and reconstruction of feature information, this scheme designs a strategy based on residual channel attention. The designed DEM reconstruction module network structure is as follows: Figure 4 (b). First, the F obtained by the multi-scale feature fusion module low OPT-DEM and F highOPT-DEM The channel dimension is concatenated to get a size of 64x64x304 Flow-high OPT-DEM . The splicing results are subjected to two layers of convolution and upsampling and then input into two residual group structures connected in series. ResG consists of multiple residual blocks and short jump connections, which are used to focus on the learning of high-frequency information. Short jump connections can effectively alleviate the problem of information loss. ResG includes several RCAB modules. The RCAB module introduces a channel attention mechanism to adaptively weight the channels of fused features, enhance the information representation ability of key features, and suppress the interference of redundant information. The calculation process is Concat->Conv->Upsample->ResG. This mechanism enables the model to focus more accurately on high-dimensional features that are critical to DEM reconstruction, thereby improving the accuracy and reliability of reconstruction. Finally, the original feature information of the low-resolution DEM is used, which enables the network to form residual reasoning, thereby speeding up the fitting speed. The DEM reconstruction module can output a more refined high-resolution DEM, significantly improving the quality of DEM generation and the ability to reconstruct details.

[0067] Step S103: training a multimodal fusion super-resolution network using the multimodal dataset constructed in step S101 and the collaborative loss function.

[0068] In order to ensure the refinement of terrain features in the process of DEM super-resolution reconstruction, a loss module based on collaborative optimization is proposed, including elevation loss L1 and slope loss RMSE, and the weight of each loss is adjusted through a dynamic weighting strategy.

[0069] (1) Height loss (L1)

[0070] Usually, the super-resolution reconstruction task uses the mean square error (MSE) as the loss function. Compared with MSE, the mean absolute error (MAE) has the same degree of penalty for outlier errors, which can avoid blurring details in the DEM super-resolution process. Therefore, the L1 loss is chosen to measure the error between the generated high-resolution DEM and the true high-resolution DEM.

[0071]

[0072] Where m is the number of samples, Y SR (i) Indicates the elevation value for generating a high-resolution DEM image, Y HR (i) Represents the elevation value of a real high-resolution DEM image.

[0073] (2) Slope loss (RMSE): The root mean square error of the DEM slope is calculated and used to optimize local terrain features.

[0074]

[0075] Where N is the number of samples, O i represents the slope derived from the real high-resolution DEM, S i Represents the slope derived from generating a high-resolution DEM.

[0076] (3) Collaborative loss: The resin stability of model training is optimized by dynamically adjusting the weights of each loss based on uncertainty with the same variance.

[0077]

[0078] in, is the weighted parameter of self-learning, L τ (W) represents various losses.

[0079] Corresponding to the above method, this embodiment also proposes a DEM super-resolution reconstruction system combined with a depth map, including:

[0080] The dataset construction module is used to construct a multimodal dataset, which contains paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images.

[0081] A model building module is used to design a multimodal fusion super-resolution network: a pseudo-Siamese feature extraction module is designed to extract low-level and high-level features of optical images, depth maps and low-resolution DEM images; the low-level feature maps generated by the pseudo-Siamese feature extraction module are fused at the shallow layer and the deep-level feature maps are fused at the deep layer through the multi-scale feature fusion module; a high-resolution DEM image is output through the DEM reconstruction module.

[0082] The model training module is used to train the multimodal fusion super-resolution network using multimodal datasets and collaborative loss functions.

[0083] In order to verify the effectiveness and reliability of the present invention, several representative super-resolution algorithms and the MFSR model proposed in the present invention are tested on the test set, which are:

[0084] ①Bicubic: It is a classic algorithm in the field of super-resolution, which reconstructs images by using cubic polynomial interpolation of pixel neighborhoods. It can usually be used as a baseline model for super-resolution tasks.

[0085] ②SRCNN: SRCNN is the pioneering work of deep learning in the field of image super-resolution. It is a lightweight and effective model. It mainly uses a three-layer convolutional network to perform nonlinear mapping to obtain high-resolution DEM.

[0086] ③ESRGAN: ESRGAN uses a generative adversarial network in the field of super-resolution, generating high-resolution images with more realistic textures through adversarial training between the generator and the discriminator.

[0087] ④TfaSR: TfaSR is the first super-resolution network based on terrain features, integrating a deep residual module and a deformable convolution module to extract deep and adaptive terrain features, respectively. The performance on DEM super-resolution has achieved satisfactory results.

[0088] ⑤SRFormer: SRFormer is an image super-resolution method based on the Transformer architecture, which improves the reconstruction effect of low-resolution images to high-resolution through the self-attention mechanism.

[0089] ⑥Real-GDSR: Real-GDSR is an optical image-guided DEM super-resolution method that combines convolutional neural networks and diffusion processes. It has been used to achieve more accurate DSM super-resolution.

[0090] Figure 5 The quantitative results comparing the methods on the test set are listed, including the RMSE values ​​of elevation, slope, and aspect. In general, the proposed MFSR model using multiple modalities (such as DEM+Depth) significantly outperforms all baseline methods in the three main indicators. Specifically, compared with the best benchmark, the MFSR model shows superior performance in the test set, with RMSE-Elevation improved by 24.63%, RMSE-Slope improved by 22.05%, and RMSE-Aspect improved by 11.44%.

[0091] It can be seen that deep learning methods usually have more satisfactory performance and greater stability than traditional interpolation methods (Bicubic methods). The transformer-based SRFormer model can achieve better results for DEMSR tasks than CNN-based SRCNN and GAN-based ESRGAN. It shows that permutation self-attention (PSA), which strikes a proper balance between the channels of self-attention and spatial information, is suitable for DEM images. The RMSE achieved by SRCNN on the test set is not as good as that on the training set, which is due to the overfitting of the model caused by the simple network structure. However, none of the above models considers spatial patterns, while Tfasr performs better in the RMSE-Elevation indicator. Tfasr integrates a deep residual module and a DCN module to extract deep and adaptive terrain features. This shows that combining Terrain feature-aware can play a positive role in DEMSR tasks. Unlike designing a terrain perception module, the study directly solves the fundamental problem from the input data. Multimodal data can directly provide auxiliary information for DEM, thereby guiding DEMSR. However, Real-GDSR is also a joint multimodal (DEM + optical image), but the quantitative results obtained are not ideal. On the one hand, the reason is that Real-GDSR is designed to solve the DSM optimization problem with obvious ground objects through a two-step method of local refinement and edge enhancement diffusion. However, it fails for the undulating DEM in the dataset of the present invention. On the other hand, the terrain auxiliary information provided by optical images is not as intuitive as the terrain information represented by the relative depth map generated by DAM for the network to extract useful detail features. Therefore, the proposed MFSR fundamentally provides a more accurate and reliable solution for DEMSR from the perspective of data source.

[0092] In order to better understand the DEMSR results of different methods, box plots of the overall error are drawn, such as Figure 6 As shown. The box plot shows the error distribution and dispersion of each method. It can be observed that, in comparison, the error of Bicubic is relatively large and dispersed, indicating that it is not as good as the learning-based method in the DEMSR task. It is worth mentioning that the box of Real-GDSR in the Elevation error analysis is narrower, indicating that the introduction of optical image guidance SR strategy can make the network more stable. In the three sub-images, the proposed MFSR box is narrower and the median is lower, indicating that the error of this method is more concentrated and the overall performance is better.

[0093] An example of a DEM in the test set is selected to illustrate the meaning of the above evaluation indicators. Figure 7As shown, the MFSR of the present invention is based on trimodal optical images, relative depth maps and low-resolution DEM as input. Real-GDSR uses optical images and low-resolution DEM as input. Other models only input low-resolution DEM. The high-resolution DEM results generated by different methods are visualized in the figure. At first glance, all methods can restore the low-resolution DEM to a higher-resolution version. However, there are many differences between these results. In order to better illustrate the difference between the SR result and the original high-resolution DEM, the error map obtained by subtracting the original high-resolution DEM from the elevation value of SR is visualized. In addition, the statistical analysis of the error distribution is visualized as a line graph. It can be observed that the high-resolution DEM generated by the MFSR of the present invention is closer to the real one than that generated by other methods, and the elevation information can be restored more accurately. It can be seen from the error distribution diagram that most of the errors of MFSR are distributed near 0. This verifies that the introduction of multimodal data can provide more terrain information, which helps to reconstruct and restore details of DEM. It should be pointed out that, except for Real-GDSR, the overall errors of other methods tend to be negative, which means that the super-resolution results are lower than the original high-resolution DEM elevation values. This may be because the selected area is a valley, with high elevation values ​​around and low elevation values ​​in the middle, which makes the model learn to generate a priori lower elevation values. The valley generated by Real-GDSR is wider and the elevation is higher than the original value, which makes the error distribution graph have a peak on the right side of the 0 value.

[0094] Unless otherwise specifically stated, the components, steps, numerical expressions and values ​​set forth in these embodiments do not limit the scope of the present invention.

[0095] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0096] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.

[0097] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.

[0098] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A DEM super-resolution reconstruction method combined with a depth map, characterized in that: include: Construct a multimodal dataset, which includes paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images; Design a multimodal fusion super-resolution network: Design a pseudo-Siamese feature extraction module to extract low-level and high-level features of optical images, depth maps, and low-resolution DEM images; Then, the low-level feature maps generated by the pseudo-Siamese feature extraction module are fused at the shallow level and the deep-level feature maps are fused at the deep level through the multi-scale feature fusion module; the high-resolution DEM image is output through the DEM reconstruction module; The multimodal fusion super-resolution network is trained using multimodal datasets and collaborative loss functions.

2. The DEM super-resolution reconstruction method combined with a depth map according to claim 1, characterized in that: The process of constructing a multimodal dataset is: Resampling: Optical images were resampled to a ground resolution of 5 m using bilinear interpolation; the acquired 30 m resolution DEM was resampled to 50 m using bilinear interpolation; Reprojection: Convert DEM images and optical images into the same WGS-84 coordinate system; Registration: The RIFT algorithm is used to geo-register the DEM image with the optical image; Clipping: Clip all the grids according to the maximum inscribed rectangle of the overlapping area obtained after processing; Dataset segmentation: The high-resolution DEM images, optical images and corresponding depth maps are segmented into blocks of 256×256 pixels, and the low-resolution DEM images are segmented into blocks of 64×64 pixels.

3. The DEM super-resolution reconstruction method combined with a depth map according to claim 1, characterized in that: The pseudo-Siamese feature extraction module uses independent convolutional flows with the same structure to extract features from optical images, depth maps and low-resolution DEM images, uses ResNet-101 as the backbone network to obtain low-level features of optics, depth and DEM, and takes the deepest feature map of the backbone network as the high-level feature.

4. The DEM super-resolution reconstruction method combined with a depth map according to claim 3 is characterized in that: The multi-scale feature fusion module specifically includes: for low-level features, the low-level features of optics, depth, and DEM are spliced ​​in the channel dimension to obtain a low-level fusion feature map; for high-level features, the high-level features of optics, depth, and DEM are spliced ​​in the channel dimension and then passed through a dilated spatial convolution pooling pyramid and up-sampling to obtain a high-level fusion feature map.

5. The DEM super-resolution reconstruction method combined with a depth map according to claim 4, characterized in that: The atrous spatial convolution pooling pyramid includes a 1×1 convolution and multiple parallel atrous convolutions. Features of different receptive fields are obtained through atrous convolutions with different expansion rates, and global information is obtained through global pooling and 1×1 convolution, and finally these features are fused.

6. The DEM super-resolution reconstruction method combined with a depth map according to claim 4, characterized in that: The DEM reconstruction module includes: splicing the low-level fusion feature map obtained by the multi-scale feature fusion module with the high-level fusion feature map in the channel dimension, and adding the splicing result to the low-resolution DEM original feature information after passing two layers of convolution, upsampling and two ResGs in sequence to output a high-resolution DEM image.

7. The DEM super-resolution reconstruction method combined with a depth map according to claim 6, characterized in that: The ResG is composed of multiple residual blocks and short skip connections. ResG includes several RCAB modules, and the RCAB module introduces a channel attention mechanism.

8. The DEM super-resolution reconstruction method combined with a depth map according to claim 1, characterized in that: The collaborative loss function includes the elevation loss L1 and the slope loss RMSE. The weight of each loss is adjusted through a dynamic weighting strategy. The total loss function expression is as follows: in, is the weighted parameter of self-learning, L τ (W) represents various losses.

9. A DEM super-resolution reconstruction system combined with a depth map, characterized in that: The method for implementing the DEM super-resolution reconstruction method combined with a depth map as claimed in any one of claims 1 to 8 comprises: The dataset construction module is used to construct a multimodal dataset, which contains paired optical images, depth maps, low-resolution DEM images, and high-resolution DEM images; Model building module, used to design a multimodal fusion super-resolution network: design a pseudo-Siamese feature extraction module to extract low-level and high-level features of optical images, depth maps and low-resolution DEM images; then use the multi-scale feature fusion module to fuse the low-level feature maps generated by the pseudo-Siamese feature extraction module at the shallow layer and the deep-level feature maps at the deep layer; output a high-resolution DEM image through the DEM reconstruction module; The model training module is used to train the multimodal fusion super-resolution network using multimodal datasets and collaborative loss functions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Super-resolution reconstruction system of digital elevation model based on attention mechanism

    CN120318076A

  • A super-resolution reconstruction system of digital elevation model based on attention mechanism

    CN120318076B

  • Arbitrary-scale super-resolution reconstruction method for multi-source heterogeneous DEM

    CN121437268A

  • Construction method of DEM super-resolution reconstruction model

    CN121599842A

  • A method for constructing a DEM super-resolution reconstruction model

    CN121599842B