Super-resolution reconstruction system of digital elevation model based on attention mechanism
The system addresses the underutilization of high-resolution optical image information in DEM reconstruction by using an attention mechanism to integrate optical and elevation data, resulting in improved DEM resolution and accuracy.
Patent Information
- Application Number
- CN202510786828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing DEM reconstruction technology fails to fully mine the complementary information of high-resolution optical remote sensing images, resulting in the reconstructed DEM images having a lower resolution and poor accuracy.
The super-resolution reconstruction system of digital elevation model based on attention mechanism is adopted, and through the optical remote sensing image processing module, the digital elevation model image processing module and the cross-modal fusion module, the high-resolution optical remote sensing image is used to guide DEM reconstruction. Combined with the attention mechanism in deep learning, it adaptively focuses on key areas and fuses the complementary information of optical remote sensing images and DEM.
Improves the resolution and reconstruction accuracy of DEM, enhances the quality of DEM, and enables more accurate prediction and interpolation of elevation information.
Smart Images

Figure CN120318076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a super-resolution reconstruction system for digital elevation models based on an attention mechanism. Background Art
[0002] Digital Elevation Models (DEMs) can be applied in a variety of fields, including land use planning, environmental monitoring, flood modeling, disaster response, and resource management. High-resolution DEMs can provide more detailed topographic information for these applications, but the direct acquisition of high-resolution DEMs is often costly, time-consuming, and impractical, especially in large-scale or remote areas.
[0003] DEM super-resolution reconstruction technology aims to reconstruct high-resolution DEMs from low-resolution DEMs, thereby restoring the fine details missing in the low-resolution DEMs. However, existing DEM reconstruction technologies have not fully exploited the complementary information of high-resolution optical remote sensing images, and their performance needs to be further improved. Summary of the Invention
[0004] In view of this, the present invention provides a super-resolution reconstruction system for digital elevation models based on an attention mechanism to solve the problems of low resolution and poor accuracy of the reconstructed images.
[0005] A super-resolution reconstruction system for digital elevation models based on an attention mechanism includes an optical remote sensing image processing module, an image processing module for digital elevation models, and a cross-modal fusion module; The optical remote sensing image guidance module takes a high-resolution optical remote sensing image as input, and processes the optical remote sensing image through the convolutional module and the residual group in the optical remote sensing image guidance module to generate the first optical remote sensing feature ; The image processing module for digital elevation models takes the initial image of the digital elevation model as input, and processes the initial image through the convolutional module in the image processing module for digital elevation models to generate the first digital elevation feature ; The cross-modal fusion module uses the attention mechanism to fuse the first optical remote sensing feature and the first digital elevation feature to obtain the second digital elevation feature , and then fuses the second digital elevation feature and the first optical remote sensing feature to obtain the second optical remote sensing feature ; The cross-modal fusion module uses the attention mechanism to process the second optical remote sensing feature and the second digital elevation feature for fusion to obtain the third digital elevation feature . Then, the third digital elevation feature and the second optical remote sensing feature are fused to obtain the third optical remote sensing feature . And so on, an optical remote sensing feature sequence and a digital elevation feature sequence are obtained. represents the total number of residual groups in the optical remote sensing image guidance module; The cross-modal fusion module then performs step-by-step fusion on each digital elevation feature and the optical remote sensing feature , and combines the initial image to obtain the reconstructed image of the digital elevation model .
[0006] According to the super-resolution reconstruction system of the digital elevation model based on the attention mechanism provided by the present invention, a high-resolution optical remote sensing image (HRSI) is used as additional data to guide the reconstruction of the high-resolution DEM, effectively mining the complementary information of the high-resolution optical remote sensing image. The optical remote sensing image can provide rich context information, such as surface coverage, surface texture, and spatial patterns, which can complement the DEM. In addition, by combining the attention mechanism in deep learning, the present invention can learn complex spatial patterns and context information from massive data. The attention mechanism can adaptively focus on the regions in the optical remote sensing image that are most critical for improving the quality of the DEM, thereby achieving more accurate elevation information prediction and interpolation, improving the resolution of the DEM, and enhancing the accuracy and reliability of the reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a structural block diagram of the super-resolution reconstruction system of the digital elevation model based on the attention mechanism provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0008] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the embodiments of the present invention, and should not be construed as limiting the present invention.
[0009] Please refer to Figure 1, an embodiment of the present invention provides a super-resolution reconstruction system for a digital elevation model based on an attention mechanism, including an optical remote sensing image processing module, an image processing module for the digital elevation model, and a cross-modal fusion module.
[0010] The input of this super-resolution reconstruction system is the initial image of the digital elevation model with low resolution and the optical remote sensing image with high resolution , and the output is the reconstructed image of the digital elevation model with high resolution , , , , is the set of real numbers, and are respectively the height and width of, is the scaling factor, is and the product of, is and the product of.
[0011] Specifically, the optical remote sensing image guiding module takes the optical remote sensing image as the input, and processes the optical remote sensing image through the convolutional module and the residual group in the optical remote sensing image guiding module to generate the first optical remote sensing feature . The convolutional module in the optical remote sensing image guiding module is composed of a 3×3 convolutional layer and a Leaky ReLU activation function layer connected in sequence.
[0012] Among them, the optical remote sensing image guiding module satisfies the following formula: ; Among them, represents the residual group operation, represents the convolutional operation; The image processing module for the digital elevation model takes the initial image of the digital elevation model as the input, and processes the initial image through the convolutional module in the image processing module for the digital elevation model to generate the first digital elevation feature . The convolutional module in the image processing module for the digital elevation model is also composed of a 3×3 convolutional layer and a Leaky ReLU activation function layer connected in sequence.
[0013] The image processing module for the digital elevation model satisfies the following formula: .
[0014] The cross-modal fusion module uses the attention mechanism to process the first optical remote sensing feature and the first digital elevation feature for fusion to obtain the second digital elevation feature . Then, the second digital elevation feature and the first optical remote sensing feature are fused through the residual group to obtain the second optical remote sensing feature .
[0015] Among them, the cross-modal fusion module satisfies the following formula: ; Among them, represents the fusion operation, which is used to fuse features of different modalities.
[0016] The cross-modal fusion module plays a key role in adaptively integrating the complementary information of data from two modalities, namely HRSI and DEM. By using the attention mechanism, it can highlight the consistent structures in the two-modal data in a learnable manner while suppressing the inconsistent components in the two-modal data. This ensures that the fused features can more effectively capture the high-frequency spatial details in the image.
[0017] The cross-modal fusion module uses the attention mechanism to fuse the second optical remote sensing feature and the second digital elevation feature to obtain the third digital elevation feature . Then, the third digital elevation feature and the second optical remote sensing feature are fused to obtain the third optical remote sensing feature . Then, the third optical remote sensing feature and the third digital elevation feature are fused to obtain the fourth digital elevation feature . Then, the fourth digital elevation feature and the third optical remote sensing feature are fused to obtain the fourth optical remote sensing feature . And so on, an optical remote sensing feature sequence and a digital elevation feature sequence are obtained. represents the total number of residual groups in the optical remote sensing image guidance module.
[0018] The cross-modal fusion module can fuse features from two modalities, capture spatial and geometric cues, and thus provide richer context information for the reconstruction of DEM. Through an attention-driven mechanism, the cross-modal fusion module selectively integrates high-resolution information from high-resolution optical remote sensing images to enhance the features of DEM.
[0019] Specifically, the cross-modal fusion module includes a feature enhancement sub-module based on a gated unit, a feature enhancement sub-module based on a cross-modal feature similarity matrix, and a cross-modal feature compression and excitation sub-module.
[0020] The feature enhancement sub-module based on the gated unit processes and using a gated unit to obtain and . denotes the optical remote sensing feature, denotes the digital elevation feature, denotes the enhanced optical remote sensing feature, denotes the enhanced digital elevation feature, .
[0021] Among them, the purpose of using a gated unit in the feature enhancement sub-module based on the gated unit to enhance features is to selectively retain and transmit key information by controlling the propagation of features in the network, thereby optimizing the features from high-resolution optical remote sensing images and DEM, and providing higher-quality features for subsequent cross-modal fusion and DEM reconstruction. This process can be expressed by the following formula: ; ; Among them, denotes the parametric rectified linear unit activation function, denotes the Sigmoid function, denotes the convolution operation, , , , are the convolution kernels of the gated unit, , , , are the bias terms of the gated unit, denotes element-wise multiplication.
[0022] The feature enhancement sub-module based on the cross-modal feature similarity matrix calculates and The similarity between them is used to obtain the similarity matrix of cross-modal features. Based on the similarity matrix, Enhance the features to obtain the enhanced features .
[0023] In the feature enhancement sub-module based on the cross-modal feature similarity matrix, first Downsample to the same spatial resolution as Then, along the channel dimension, and are unfolded into feature patches, and are respectively The height and width of, Next, use the normalized inner product to calculate and The similarity of the corresponding feature patches between them. When calculating, in order to better utilize the local context information, for The th feature patch in and th feature patch in , , is odd, such as 3×3 or 5×5. For each element in the neighborhood, different weights are assigned according to its distance from the center. The closer the element is to the center, the greater the weight, which can highlight the importance of the central patch and its nearby areas.
[0024] Distance weight can reflect the association strength between the elements in the neighborhood and the center. The closer the distance the greater, which can be calculated by the Gaussian function or manually setting the distance attenuation rule to make the weight of the central element in the neighborhood higher.
[0025] Specifically, and The similarity between them is calculated as follows:
[0026] ; Among them, represents the distance weight, represents The element in the neighborhood of, represents The element in the neighborhood of, represents the neighborhood range, Represents the coordinates of elements in the neighborhood, Represents the two-norm operation, Represents the size parameter of the neighborhood, Represents the parameter for controlling the attenuation rate.
[0027] Since local context information is considered in calculating similarity, it can measure the similarity between small blocks of different modal features more comprehensively and accurately, thus providing more effective information for subsequent feature fusion and model training.
[0028] Then calculate and The similarity weight matrix between , and the expression is: ; Among them, Represents the folding operation, Represents concatenation along the channel dimension; The calculated similarity weight matrix, as an attention mechanism, is used to selectively enhance the features of the high-resolution image. Small blocks of features with higher similarity scores will be assigned greater weights to ensure their more important contributions to the reconstruction of the final digital elevation model. Using the similarity weight matrix, the enhanced features can be obtained, and the expression is: ; Among them, f us Represents upsampling.
[0029] The cross-modal feature compression and excitation sub-module performs feature compression and excitation on and to obtain , Represents the Digital elevation feature.
[0030] In the cross-modal feature compression and excitation sub-module, feature compression can reduce the dimension while retaining key information and integrating cross-modal data, and feature excitation selectively enhances and optimizes the most informative features.
[0031] During the feature compression process, cross-modal feature fusion is first performed to integrate complementary information from different sources, and then average pooling and variance pooling are applied for feature optimization. The expression is: ; ; Among them, Is the feature generated by cross-modal fusion, Is the feature obtained after using average pooling and variance pooling, and represent the convolution kernel and bias term in cross-modal feature fusion respectively, represents average pooling, represents variance pooling; In feature excitation, is subjected to feature excitation to obtain , and the expression is: ; ; ; Among them, is the first intermediate feature, is the second intermediate feature, and are the first convolution kernel and the first bias term in feature excitation, and are the second convolution kernel and the second bias term in feature excitation.
[0032] The cross-modal fusion module then performs step-by-step fusion on each digital elevation feature and the optical remote sensing feature in , and combines it with the initial image to obtain the reconstructed image of the digital elevation model. The expression of this process is as follows: ; ; ; … ; ; Among them, f bi represents the bilinear interpolation operation, , , , , respectively represent the -th, the -th, the -th, the 2nd, and the 1st cross-modal fusion features, , , respectively represent the -th, the -th, and the 2nd digital elevation features.
[0033] In summary, for the super-resolution reconstruction system of digital elevation model based on the attention mechanism according to the above embodiments, a high-resolution optical remote sensing image (HRSI) is used as additional data to guide the reconstruction of the high-resolution DEM, effectively mining the complementary information of the high-resolution optical remote sensing image. The optical remote sensing image can provide rich context information, such as surface coverage, surface texture, and spatial patterns, which can complement the DEM. In addition, by combining the attention mechanism in deep learning, the present invention can learn complex spatial patterns and context information from massive data. The attention mechanism can adaptively focus on the areas in the optical remote sensing image that are most critical for improving the quality of the DEM, thereby achieving more accurate elevation information prediction and interpolation, improving the resolution of the DEM, and enhancing the accuracy and reliability of the reconstruction.
[0034] The above embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.
Claims
1. A super-resolution reconstruction system for digital elevation models based on an attention mechanism, characterized in that, It includes an optical remote sensing image processing module, an image processing module for digital elevation model, and a cross-modal fusion module; The optical remote sensing image guidance module takes the optical remote sensing image as input, and processes the optical remote sensing image through the convolutional module and residual group in the optical remote sensing image guidance module to generate the first optical remote sensing feature ; The image processing module of the digital elevation model uses the initial image of the digital elevation model as input, and processes the initial image through the convolution module in the image processing module of the digital elevation model to generate the first digital elevation feature ; The cross-modal fusion module uses the attention mechanism to fuse the first optical remote sensing feature and the first digital elevation feature to obtain the second digital elevation feature . Then, the second digital elevation feature and the first optical remote sensing feature are fused to obtain the second optical remote sensing feature ; The cross-modal fusion module uses the attention mechanism to fuse the second optical remote sensing feature and the second digital elevation feature to obtain the third digital elevation feature , and then fuse the third digital elevation feature and the second optical remote sensing feature to obtain the third optical remote sensing feature , and so on, to obtain the optical remote sensing feature sequence and the digital elevation feature sequence , represents the total number of residual groups in the optical remote sensing image guidance module; The cross-modal fusion module then for each digital elevation feature and the optical remote sensing feature are gradually fused, and combined with the initial image to obtain the reconstructed image of the digital elevation model .
2. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 1, wherein The optical remote sensing image guidance module satisfies the following formula: ; Among them, represents a residual group operation, represents a convolution operation; The image processing module for digital elevation model satisfies the following formula: 。 3. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 2, wherein The cross-modal fusion module satisfies the following formula: ; Among them, represents a fusion operation.
4. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 3, wherein The cross-modal fusion module includes a feature enhancement sub-module based on a gated unit, a feature enhancement sub-module based on a cross-modal feature similarity matrix, and a cross-modal feature compression and excitation sub-module; The feature augmentation sub-module based on the gating unit uses the gating unit to process respectively and , obtaining and , denotes the optical remote sensing feature, denotes the digital elevation feature, denotes the enhanced optical remote sensing feature, denotes the enhanced digital elevation feature, ; Calculation of the feature enhancer sub-module based on the cross-modal feature similarity matrix and Calculate the similarity between them to obtain the similarity matrix between cross-modal features, and enhance the features based on the similarity matrix to obtain the enhanced features ; Cross-modal feature compression and excitation sub-module pair and perform feature compression and excitation to obtain , indicating the digital elevation feature.
5. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 4, characterized in that, The feature enhancement sub-module based on a gated unit satisfies the following formula: ; ; Among them, represents the parametric rectified linear unit activation function, represents the Sigmoid function, represents the convolution operation, 、 、 、 are the convolution kernels of the gating unit, 、 、 、 are the bias terms of the gating unit, represents element-wise multiplication.
6. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 5, wherein In the feature enhancement sub-module based on the cross-modal feature similarity matrix, first, is downsampled to the same spatial resolution as, and then, along the channel dimension, and are unfolded into feature patches. and are respectively the height and width of. Next, the normalized inner product is used to calculate the similarity between the corresponding feature patches of and . For the th feature patch in and the th feature patch in , , and , the similarity is calculated as follows: ; Among them, represents the distance weight, represents the elements within the neighborhood of represents the elements within the neighborhood of represents the neighborhood range, represents the coordinates of the elements within the neighborhood, represents the two-norm operation, represents the size parameter of the neighborhood, represents the parameter for controlling the attenuation rate; Then calculate and the similarity weight matrix , and the expression is: ; Among them, represents a folding operation, represents concatenation along the channel dimension; Thereby obtaining the enhanced feature , and the expression is: ; Among them, f us represents upsampling.
7. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 6, wherein In the cross-modal feature compression and excitation sub-module, cross-modal feature fusion is first performed, and then average pooling and variance pooling are applied for feature optimization. The expression is: ; ; Among them, is the feature generated by cross-modal fusion, is the feature obtained after using average pooling and variance pooling, and represent the convolution kernel and bias term in cross-modal feature fusion respectively, represents average pooling, represents variance pooling; Then, perform feature excitation on to obtain , and the expression is: ; ; ; Among them, is the first intermediate feature, is the second intermediate feature, and are the first convolutional kernel and the first bias term in feature excitation, and are the second convolutional kernel and the second bias term in feature excitation.
8. The super-resolution reconstruction system of a digital elevation model based on an attention mechanism according to claim 7, characterized in that, The cross-modal fusion module also satisfies the following formula: ; ; ; … ; ; Among them, f bi represents a bilinear interpolation operation, , , , , respectively represent the th, th, th, 2nd, and 1st cross-modal fusion features, , , respectively represent the th, th, and 2nd digital elevation features.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on self-attention fusion
CN112712488A
Remote sensing image semantic segmentation method based on cross-modal fusion and graph neural network
CN116797787A
Remote sensing image super-resolution reconstruction method based on multi-path fusion and attention
CN117173022A
DEM reconstruction model training method and DEM reconstruction method
CN118037976A
DEM (Digital Elevation Model) super-resolution reconstruction method and system fusing gradient feature Transform
CN119251419A
Cited By
Digital elevation model reconstruction method and system
CN120976468A
Digital elevation model reconstruction method and system
CN120976468B
DEM super-resolution reconstruction system based on RGB optical image guidance
CN122434738A