A super-resolution reconstruction system of digital elevation model based on attention mechanism
Through a digital elevation model super-resolution reconstruction system based on attention mechanism, DEM reconstruction is guided by high-resolution optical remote sensing images, solving the problems of low DEM resolution and poor accuracy in the prior art, and achieving higher precision DEM reconstruction effect.
Patent Information
- Application Number
- CN202510786828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing DEM reconstruction technology fails to fully mine the complementary information of high-resolution optical remote sensing images, resulting in the reconstructed DEM resolution and poor accuracy.
The super-resolution reconstruction system of digital elevation model based on attention mechanism is adopted, and the DEM reconstruction is guided by high-resolution optical remote sensing images through optical remote sensing image processing module, digital elevation model image processing module and cross-modal fusion module. The attention mechanism in deep learning adaptively focuses on key areas and fuses rich context information.
More accurate elevation information prediction and interpolation are achieved, improving the resolution and reconstruction accuracy and reliability of DEM.
Smart Images

Figure CN120318076B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a super-resolution reconstruction system of a digital elevation model based on an attention mechanism. Background Art
[0002] Digital Elevation Models (DEMs) can be used in a variety of fields, including land use planning, environmental monitoring, flood modeling, disaster response, and resource management. High-resolution DEMs can provide more detailed terrain information for these applications, but directly acquiring high-resolution DEMs is often costly, time-consuming, and impractical, especially in large or remote areas.
[0003] DEM super-resolution reconstruction technology aims to reconstruct a high-resolution DEM from a low-resolution DEM, thereby recovering the fine details missing in the low-resolution DEM. However, existing DEM reconstruction techniques fail to fully exploit the complementary information of high-resolution optical remote sensing images, and their performance needs to be further improved. Summary of the Invention
[0004] In view of this, the present invention provides a super-resolution reconstruction system of a digital elevation model based on an attention mechanism to solve the problem of low resolution and poor accuracy of the reconstructed image.
[0005] A super-resolution reconstruction system for digital elevation models based on an attention mechanism, comprising an optical remote sensing image processing module, a digital elevation model image processing module, and a cross-modal fusion module;
[0006] The optical remote sensing image guidance module uses high-resolution optical remote sensing images As input, the optical remote sensing image is processed by the convolution module and the residual group in the optical remote sensing image guidance module to generate the first optical remote sensing feature ;
[0007] The image processing module of the digital elevation model uses the initial image of the digital elevation model As input, the initial image is processed by the convolution module in the image processing module of the digital elevation model to generate the first digital elevation feature ;
[0008] The cross-modal fusion module uses the attention mechanism to focus on the first optical remote sensing feature and the 1st Digital Elevation Feature Fusion is performed to obtain the second digital elevation feature , then the second digital elevation feature and the first optical remote sensing feature Fusion is performed to obtain the second optical remote sensing feature ;
[0009] The cross-modal fusion module uses the attention mechanism to analyze the second optical remote sensing features. and the second digital elevation feature Fusion is performed to obtain the third digital elevation feature , then the third digital elevation feature and the second optical remote sensing feature Fusion is performed to obtain the third optical remote sensing feature , and so on, we get the optical remote sensing feature sequence and digital elevation feature sequences , represents the total number of residual groups in the optical remote sensing image guidance module;
[0010] Cross-modal fusion module Each digital elevation feature and Optical remote sensing characteristics Perform gradual fusion and combine with the initial image , obtain the reconstructed image of the digital elevation model .
[0011] According to the super-resolution reconstruction system of the digital elevation model based on the attention mechanism provided by the present invention, high-resolution optical remote sensing images (HRSI) are used as additional data to guide the reconstruction of high-resolution DEM, effectively mining the complementary information of high-resolution optical remote sensing images. Optical remote sensing images can provide rich contextual information, such as surface cover, surface texture and spatial patterns, which can complement DEM. In addition, by combining the attention mechanism in deep learning, the present invention can learn complex spatial patterns and contextual information from massive data. The attention mechanism can adaptively focus on the areas in the optical remote sensing image that are most critical for improving the quality of DEM, thereby achieving more accurate elevation information prediction and interpolation, improving the resolution of DEM, and enhancing the accuracy and reliability of reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A structural block diagram of a super-resolution reconstruction system for a digital elevation model based on an attention mechanism provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the embodiments of the present invention, and should not be construed as limiting the present invention.
[0014] See also Figure 1 An embodiment of the present invention provides a super-resolution reconstruction system of a digital elevation model based on an attention mechanism, including an optical remote sensing image processing module, an image processing module of a digital elevation model, and a cross-modal fusion module.
[0015] The input of the super-resolution reconstruction system is the initial image of the low-resolution digital elevation model and high-resolution optical remote sensing images , the output is a reconstructed image of a high-resolution digital elevation model , , , , is the set of real numbers, and They are The height and width, is the scaling factor, for and The product of for and The product of .
[0016] Specifically, the optical remote sensing image guidance module uses optical remote sensing images As input, the optical remote sensing image is processed by the convolution module and the residual group in the optical remote sensing image guidance module to generate the first optical remote sensing feature The convolution module in the optical remote sensing image guidance module consists of a 3×3 convolution layer and a Leaky ReLU activation function layer connected in sequence.
[0017] Among them, the optical remote sensing image guidance module satisfies the following formula:
[0018] ;
[0019] in, represents the residual group operation, Represents the convolution operation;
[0020] The image processing module of the digital elevation model uses the initial image of the digital elevation model As input, the initial image is processed by the convolution module in the image processing module of the digital elevation model to generate the first digital elevation feature The convolution module in the image processing module of the digital elevation model is also composed of a 3×3 convolution layer and a Leaky ReLU activation function layer connected in sequence.
[0021] The image processing module of the digital elevation model satisfies the following formula:
[0022] .
[0023] The cross-modal fusion module uses the attention mechanism to focus on the first optical remote sensing feature and the 1st Digital Elevation Feature Fusion is performed to obtain the second digital elevation feature , then the second digital elevation feature and the first optical remote sensing feature The second optical remote sensing feature is obtained by fusing the residual group .
[0024] Among them, the cross-modal fusion module satisfies the following formula:
[0025] ;
[0026] in, Represents a fusion operation, which is used to fuse features of different modalities.
[0027] The cross-modal fusion module plays a key role in adaptively integrating the complementary information from the two modal data, HRSI and DEM. By utilizing the attention mechanism, it can highlight the consistent structures in the two modal data in a learnable way while suppressing the inconsistent components in the two modal data. This ensures that the fused features It can more effectively capture high-frequency spatial details in images.
[0028] The cross-modal fusion module uses the attention mechanism to analyze the second optical remote sensing features. and the second digital elevation feature Fusion is performed to obtain the third digital elevation feature , then the third digital elevation feature and the second optical remote sensing feature Fusion is performed to obtain the third optical remote sensing feature , then the third optical remote sensing feature and 3rd digital elevation features Fusion is performed to obtain the fourth digital elevation feature , then the fourth digital elevation feature and the third optical remote sensing feature Fusion is performed to obtain the fourth optical remote sensing feature , and so on, we get the optical remote sensing feature sequence and digital elevation feature sequences , Represents the total number of residual groups in the optical remote sensing image guidance module.
[0029] The cross-modal fusion module fuses features from both modalities, capturing spatial and geometric cues to provide richer contextual information for DEM reconstruction. Through an attention-driven mechanism, the cross-modal fusion module selectively integrates high-resolution information from high-resolution optical remote sensing images to enhance the features of the DEM.
[0030] Specifically, the cross-modal fusion module includes a feature amplification submodule based on a gating unit, a feature enhancement submodule based on a cross-modal feature similarity matrix, and a cross-modal feature compression and excitation submodule.
[0031] The feature enhancement module based on the gated unit uses the gated unit to process and ,get and , Indicates the Optical remote sensing characteristics, Indicates the Digital elevation features, Indicates the enhanced Optical remote sensing characteristics, Indicates the enhanced Digital elevation features, .
[0032] The purpose of the feature enhancement module based on the gated unit is to enhance the features by controlling the propagation of features in the network, selectively retaining and transmitting key information, thereby optimizing the features from high-resolution optical remote sensing images and DEMs, and providing better features for subsequent cross-modal fusion and DEM reconstruction. This process can be expressed as follows:
[0033] ;
[0034] ;
[0035] in, represents the parameterized rectified linear unit activation function, represents the Sigmoid function, represents the convolution operation, 、 、 、 is the convolution kernel of the gated unit, 、 、 、 is the bias term of the gate unit, Represents element-wise multiplication.
[0036] Feature Enhancer Module Computation Based on Cross-Modal Feature Similarity Matrix and The similarity between them is used to obtain the similarity matrix between cross-modal features. Enhance features and get enhanced features .
[0037] In the feature enhancement submodule based on the cross-modal feature similarity matrix, Downsample to and The same spatial resolution, then, along the channel dimension and Expand to A small feature block, and They are The height and width of , next, use the normalized inner product to calculate and The similarity between the corresponding feature blocks is calculated in order to better utilize the local context information. Middle Feature Block and The Feature Block , , construct a local neighborhood centered on each feature patch. Assume that the neighborhood size is , It is an odd number, such as 3×3 or 5×5. For each element in the neighborhood, different weights are assigned according to its distance from the center. The closer the element is to the center, the greater the weight, which can highlight the importance of the central block and its surrounding areas.
[0038] Distance Weight It can reflect the strength of the association between the elements in the neighborhood and the center. The closer the distance, the stronger the The larger it is, the higher the weight of the neighborhood center element can be calculated using a Gaussian function or manually set distance decay rules.
[0039] Specifically, and The similarity between The calculation formula is:
[0040]
[0041] ;
[0042] in, represents the distance weight, express Elements in the neighborhood of express Elements in the neighborhood of represents the neighborhood range, represents the coordinates of elements in the neighborhood, represents the two-norm operation, represents the size parameter of the neighborhood, Represents the parameter that controls the decay speed.
[0043] Since local context information is taken into account when calculating similarity, the similarity between small blocks of features from different modalities can be measured more comprehensively and accurately, thereby providing more effective information for subsequent feature fusion and model training.
[0044] Then calculate and The similarity weight matrix between , the expression is:
[0045] ;
[0046] in, Represents a folding operation, Indicates splicing along the channel dimension;
[0047] The calculated similarity weight matrix is used as an attention mechanism to selectively enhance the features of the high-resolution image. Feature patches with higher similarity scores are given greater weights to ensure that they make a more important contribution to the final digital elevation model reconstruction. Using the similarity weight matrix, the enhanced features can be obtained. , the expression is:
[0048] ;
[0049] Among them, f us Indicates upsampling.
[0050] Cross-modal feature compression and excitation sub-module pair and Perform feature compression and excitation to obtain , Indicates the Digital elevation feature.
[0051] In the cross-modal feature compression and excitation sub-module, feature compression can reduce dimensionality while retaining key information and integrating cross-modal data, while feature excitation selectively enhances and optimizes the most informative features.
[0052] In the feature compression process, cross-modal feature fusion is first performed to integrate complementary information from different sources, and then average pooling and variance pooling are applied for feature optimization. The expression is:
[0053] ;
[0054] ;
[0055] in, is the feature generated by cross-modal fusion, It is the feature obtained after using average pooling and variance pooling. and Represent the convolution kernel and bias term in cross-modal feature fusion, represents average pooling, represents variance pooling;
[0056] In feature excitation, Perform feature excitation and obtain , the expression is:
[0057] ;
[0058] ;
[0059] ;
[0060] in, is the first intermediate feature, is the second intermediate feature, and is the first convolution kernel and the first bias term in the feature excitation, and It is the second convolution kernel and the second bias term in the feature excitation.
[0061] Cross-modal fusion module Each digital elevation feature and Optical remote sensing characteristics Perform gradual fusion and combine with the initial image , obtain the reconstructed image of the digital elevation model , the expression of this process is as follows:
[0062] ;
[0063] ;
[0064] ;
[0065] …
[0066] ;
[0067] ;
[0068] Among them, f bi represents a bilinear interpolation operation, 、 、 、 、 Respectively represent , No. , No. , the second and first cross-modal fusion features, 、 、 Respectively represent , No. , the second digital elevation feature.
[0069] In summary, according to the super-resolution reconstruction system of the digital elevation model based on the attention mechanism of the above embodiment, high-resolution optical remote sensing images (HRSI) are used as additional data to guide the reconstruction of high-resolution DEM, and the complementary information of high-resolution optical remote sensing images is effectively mined. Optical remote sensing images can provide rich contextual information, such as surface cover, surface texture and spatial patterns, which can complement DEM. In addition, the present invention can learn complex spatial patterns and contextual information from massive data by combining the attention mechanism in deep learning. The attention mechanism can adaptively focus on the areas in the optical remote sensing image that are most critical for improving the quality of DEM, thereby achieving more accurate elevation information prediction and interpolation, which can improve the resolution of DEM and enhance the accuracy and reliability of reconstruction.
[0070] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A super-resolution reconstruction system of a digital elevation model based on an attention mechanism, characterized in that: Including optical remote sensing image processing module, digital elevation model image processing module, and cross-modal fusion module; Optical remote sensing image guidance module uses optical remote sensing image As input, the optical remote sensing image is processed by the convolution module and the residual group in the optical remote sensing image guidance module to generate the first optical remote sensing feature ; The image processing module of the digital elevation model uses the initial image of the digital elevation model As input, the initial image is processed by the convolution module in the image processing module of the digital elevation model to generate the first digital elevation feature ; The cross-modal fusion module uses the attention mechanism to focus on the first optical remote sensing feature and the 1st Digital Elevation Feature Fusion is performed to obtain the second digital elevation feature , then the second digital elevation feature and the first optical remote sensing feature Fusion is performed to obtain the second optical remote sensing feature ; The cross-modal fusion module uses the attention mechanism to analyze the second optical remote sensing features. and the second digital elevation feature Fusion is performed to obtain the third digital elevation feature , then the third digital elevation feature and the second optical remote sensing feature Fusion is performed to obtain the third optical remote sensing feature , and so on, we get the optical remote sensing feature sequence and digital elevation feature sequences , represents the total number of residual groups in the optical remote sensing image guidance module; Cross-modal fusion module Each digital elevation feature and Optical remote sensing characteristics Perform gradual fusion and combine with the initial image , obtain the reconstructed image of the digital elevation model .
2. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 1 is characterized in that: The optical remote sensing image guidance module satisfies the following formula: ; in, represents the residual group operation, Represents the convolution operation; The image processing module of the digital elevation model satisfies the following formula: 。 3. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 2 is characterized in that: The cross-modal fusion module satisfies the following formula: ; in, Represents a fusion operation.
4. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 3 is characterized in that: The cross-modal fusion module includes a feature amplification submodule based on a gated unit, a feature enhancement submodule based on a cross-modal feature similarity matrix, and a cross-modal feature compression and excitation submodule; The feature enhancement module based on the gated unit uses the gated unit to process and ,get and , Indicates the Optical remote sensing characteristics, Indicates the Digital elevation features, Indicates the enhanced Optical remote sensing characteristics, Indicates the enhanced Digital elevation features, ; Feature Enhancer Module Computation Based on Cross-Modal Feature Similarity Matrix and The similarity between them is used to obtain the similarity matrix between cross-modal features. Enhance features and get enhanced features ; Cross-modal feature compression and excitation sub-module pair and Perform feature compression and excitation to obtain , Indicates the Digital elevation feature.
5. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 4 is characterized in that: The feature enhancement module based on the gate control unit satisfies the following formula: ; ; in, represents the parameterized rectified linear unit activation function, represents the Sigmoid function, represents the convolution operation, 、 、 、 is the convolution kernel of the gated unit, 、 、 、 is the bias term of the gate unit, Represents element-wise multiplication.
6. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 5 is characterized in that: In the feature enhancement submodule based on the cross-modal feature similarity matrix, Downsample to The same spatial resolution, then, along the channel dimension and Expand to A small feature block, and They are The height and width of , next, use the normalized inner product to calculate and The similarity between the corresponding feature blocks, for Middle Feature Block and The Feature Block , , and The similarity between The calculation formula is: ; in, represents the distance weight, express Elements in the neighborhood of express Elements in the neighborhood of represents the neighborhood range, represents the coordinates of elements in the neighborhood, represents the two-norm operation, represents the size parameter of the neighborhood, Represents the parameter that controls the decay speed; Then calculate and The similarity weight matrix between , the expression is: ; in, Represents a folding operation, Indicates splicing along the channel dimension; Then we get the enhanced features , the expression is: ; Among them, f us Indicates upsampling.
7. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 6 is characterized in that: In the cross-modal feature compression and excitation submodule, cross-modal feature fusion is performed first, and then average pooling and variance pooling are applied for feature optimization. The expression is: ; ; in, is the feature generated by cross-modal fusion, It is the feature obtained after using average pooling and variance pooling. and Represent the convolution kernel and bias term in cross-modal feature fusion, represents average pooling, represents variance pooling; Then Perform feature excitation and obtain , the expression is: ; ; ; in, is the first intermediate feature, is the second intermediate feature, and is the first convolution kernel and the first bias term in the feature excitation, and It is the second convolution kernel and the second bias term in the feature excitation.
8. The super-resolution reconstruction system of the digital elevation model based on the attention mechanism according to claim 7, characterized in that: The cross-modal fusion module also satisfies the following formula: ; ; ; … ; ; Among them, f bi represents a bilinear interpolation operation, 、 、 、 、 Respectively represent , No. , No. , the second and first cross-modal fusion features, 、 、 Respectively represent , No. , the second digital elevation feature.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on self-attention fusion
CN112712488A
DEM (Digital Elevation Model) super-resolution reconstruction method and system combined with depth map
CN119991440A