A land surface temperature downscaling method based on guided feature enhancement
Patent Information
- Application Number
- CN202611017693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-07-09
AI Technical Summary
然而,现有方法MoCoLSK(Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li, Xiang Li, KangNi, Jianhui Xu, Xiangbo Shu, and Jian Yang. MoCoLSK: Modality-ConditionedHigh-Resolution Downscaling for Land Surface Temperature. IEEE Transactionson Geoscience and Remote Sensing, Volumn 63, Pages 1-17, 2025)更多关注通过大核卷积增强源尺度LST表示,而较少关注各种目标尺度引导中固有的细粒度纹理外观,未能充分释放引导数据在细粒度特征融合和重建中的潜力
(1)本发明提出的混合梯度注意力块(MGAB)通过四种专用差分卷积进行混合梯度提取,配合多级注意力机制(通道、空间、像素注意力),能够精确恢复锐利的热边界,同时忽略无关的周围上下文,有效解决了现有方法过度平滑的问题。
Smart Images

Figure CN122530754B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of remote sensing image processing and land surface temperature reconstruction, and in particular to a land surface temperature downscaling method based on guided feature enhancement. Background Technology
[0002] Land surface temperature (LST) physically represents the thermal radiation emitted by the Earth's surface and is a key geophysical parameter controlling surface-atmosphere energy exchange, urban climate dynamics, and the hydrological cycle. Unlike the relatively stable optical reflectivity, LST exhibits high-frequency spatiotemporal fluctuations, placing stringent demands on satellite observation capabilities. However, due to technological and budgetary constraints, different thermal infrared sensors exhibit a fundamental trade-off between spatial and temporal resolution. For example, the Medium Resolution Imaging Spectroradiometer (MRISD) on the Aqua satellite observes twice daily but has a spatial resolution of only 1 kilometer, while the thermal infrared sensor on the Landsat 8 satellite has a spatial resolution of 100 meters but a revisit period of 16 days. This observational gap significantly limits the potential of satellite LST products in various applications.
[0003] To bridge this gap, the LST downscaling technique has emerged, capable of reconstructing target-scale (high-resolution) thermal fields using only source-scale (low-resolution) observations. Traditional methods typically utilize physical correlations under the scale-invariant assumption (Nurit Agam, William P. Kustas, Martha C. Anderson, Fuqin Li, Christopher M.U. Neale. A vegetation index based technique for spatial sharpening of thermal imagery. Remote sensing of Environment, Volume 107, Issue 4, Pages 545-558, 2007), but this assumption is physically controversial for heterogeneous topography. Due to the inherent nonlinearity of land surface processes, these physical correlation models often fail to capture local thermal anomalies.
[0004] With the development of deep learning, single-image super-resolution (SISR) methods have made significant progress in LST downscaling (BM Nguyen, G. Tian, M.-T. Vo, A. Michel, T. Corpetti and C. Granero-Belinchon. Convolutional Neural Network Modelling for MODIS LandSurface Temperature Super-Resolution. European Signal Processing Conference (EUSIPCO), Pages 1806-1810, 2022). However, directly applying a general SISR model to geophysical data ignores the underlying thermal gradient, inevitably over-smoothing the sharp boundaries needed to resolve thermal anomalies. Recently, guided downscaling methods have achieved significant performance improvements by introducing guided data (composed of multispectral bands) at an auxiliary target scale. However, the existing method MoCoLSK (Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li, Xiang Li, KangNi, Jianhui Xu, Xiangbo Shu, and Jian Yang. MoCoLSK: Modality-Conditioned High-Resolution Downscaling for Land Surface Temperature. IEEE Transactions on Geoscience and Remote Sensing, Vol. 63, Pages 1-17, 2025) focuses more on enhancing the source-scale LST representation through large-kernel convolution, while paying less attention to the fine-grained texture appearance inherent in guides at various target scales, thus failing to fully unleash the potential of guide data in fine-grained feature fusion and reconstruction. Summary of the Invention
[0005] To overcome the shortcomings of existing LST downscaling methods in restoring thermal field details at the target scale, this invention proposes a guided feature enhancement-based land surface temperature downscaling method (GFENet). This method captures enhanced features by hybrid gradient attention blocks, performs fine-grained feature fusion in the spatial and frequency domains using a dual-domain embedding ensemble module, and mitigates upsampling artifacts through a gated guided feature refiner, thereby significantly improving the reconstruction accuracy of land surface temperature downscaling.
[0006] The technical solution of the present invention is as follows: A surface temperature downscaling method based on guided feature enhancement includes the following steps: Step 1: Input Initialization: Given source-scale land surface temperature data and target scale guided data ,in, Indicates the spatial height at the source scale. Indicates the spatial width at the source scale. Indicates the scaling factor. Indicates the number of pilot channels. Represents the real number field.
[0007] Step 2, Input Feature Extraction: Extract the source-scale LST data. and target scale guided data Input two convolutional feature extractors respectively to extract source-scale LST features. and target scale guided features .
[0008] Step 3, Guided Feature Fusion Process: Combine source-scale LST features and target scale guided features The input-guided feature fusion module consists of three fusion stages.
[0009] Step 3.1, First Fusion Stage: Integrating Source-Scale LST Features Input a residual module (ResG) to extract source-scale LST features for the first stage. At the same time, the target scale guides the features. Inputting Hybrid Gradient Attention Block (MGAB) yields the target scale-guided features for the first stage of detail enhancement. The dual-domain embedding ensemble module (DEIM) is used to integrate the source-scale LST features from the first stage. Target-scale guided features with enhanced detail The fusion process is performed to obtain the fusion characteristics of the first stage. ; Step 3.2, Second Fusion Stage: The fusion features from the first stage are then... The input is fed into the next residual module (ResG) to extract the source-scale LST features for the second stage. Simultaneously, the target scale of the first-stage detail enhancement is guided by the feature. Input the next Hybrid Gradient Attention Block (MGAB) to obtain the target scale-guided features for second-stage detail enhancement. Construct a dual-domain embedding integration module (DEIM) to integrate the source-scale LST features from the second stage. Target-scale guided features with enhanced detail The fusion process is performed to obtain the fusion characteristics of the second stage. ; Step 3.3, Third Fusion Stage: Integrating the fusion features from the second stage The input is fed into the next residual module (ResG) to extract the source-scale LST features for the third stage. Simultaneously, the target scale guidance feature for the second stage of detail enhancement is... Input the next Hybrid Gradient Attention Block (MGAB) to obtain the target scale-guided features for third-stage detail enhancement. Construct a dual-domain embedding integration module (DEIM) to integrate the source-scale LST features from the third stage. Target-scale guided features with enhanced detail The fusion process yields the fusion characteristics of the third stage. ; Step 4, Guided LST Reconstruction Process: The guided fusion features are enhanced using a gated guided feature refiner (G2FR), and then a high-resolution LST image at the target scale is output.
[0010] Step 4.1, Initial Enhancement Stage: Integrating the fusion features from the third stage. Upsampling is performed, and the results are used to enhance the target scale-guided features in the second stage of detail enhancement. Simultaneously input the gated guided feature refiner (G2FR) to obtain the initial enhanced features. ; Step 4.2, Secondary Enhancement Stage: The initial enhanced features... Input an MGAB module, its output and the target scale-guided features of the first stage of detail enhancement. Simultaneously input another gated guided feature refiner (G2FR) to obtain secondary enhanced features. ; Step 4.3, Target-scale LST reconstruction: The secondary enhancement features are then reconstructed. Input an MGAB module and add its output to the bicubic upsampling result of the source-scale LST to obtain the final predicted target-scale LST image. ; Step 5: Loss Function Optimization: The GFENet network formed in Steps 1-4 is optimized end-to-end using the Mean Squared Error (MSE) loss function to minimize the predicted target-scale LST image. Compared with real target-scale LST images Pixel-level differences between them; Step 6: Perform surface temperature downscaling inference using the optimized GFENet network: The source-scale LST image to be downscaled and its corresponding target-scale guided data are input into the optimized GFENet network. After three stages—feature extraction, guided feature fusion, and guided LST reconstruction—a high-resolution LST image at the target scale is output.
[0011] Preferably, in step 3.1, the specific operation of the Hybrid Gradient Attention Block (MGAB) is as follows: Input target scale guided features First, four types of differential convolution—horizontal differential convolution, vertical differential convolution, central differential convolution, and diagonal differential convolution—are used to perform mixed gradient extraction to obtain differential convolution features. ; The horizontal and vertical differential convolutions use learnable one-dimensional convolution kernels. To approximate the first-order gradient, central difference convolution approximates the Laplacian operator by imposing a zero-sum constraint on the weights, while diagonal difference convolution approximates the Laplacian operator by applying a zero-sum constraint to the weights. Rotation fusion captures diagonal variations, and four convolutional methods utilize learnable weights. and bias By fusing the features, differential convolution features are obtained. : in : in, These represent the weight matrices for horizontal, vertical, central, and diagonal difference convolutions, respectively. in : in, These represent the biases of the horizontal, vertical, central, and diagonal difference convolutions, respectively.
[0012] Then the differential convolution features The input multi-level attention mechanism operates as follows: Channel attention (CA): ... Global average pooling (GAP) is performed, and then channel-level attention weights are captured through a multilayer perceptron (MLP). ,in, Represents the number of feature channels; Spatial attention (SA): for Two two-dimensional feature maps are obtained by performing average pooling and max pooling along the channel dimension respectively. These two feature maps are then concatenated and processed... Convolution output space attention weights Pixel attention (PA): Attention weights are obtained by combining channel attention and spatial attention. Through interleaving operations and Concatenate, then use grouped convolution Ensure that one feature channel corresponds to one convolution operation to obtain pixel-level attention. : in, This is the Sigmoid activation function.
[0013] Pixel-level attention Sum of differential convolution features Perform multiplication and add the original number. The final output is the target scale-guided feature for the first stage of detail enhancement. : in, This is a multiplication operation.
[0014] The MGAB used in steps 3.2 and 3.3 performs the same calculation process.
[0015] Preferably, in step 3.1, the specific operation of the Dual Domain Embedded Integration Module (DEIM) is as follows: First, the source-scale LST features obtained in step 2 are... The first-stage LST features are obtained after passing through a residual module. Through dense projection Upsampling is performed to obtain upsampled LST features. Then, spatial domain embedding and frequency domain embedding are performed separately, as follows: Spatial domain embedding: for Perform a horizontal flip to obtain mirrored features, then calculate... The residual between the horizontal flip feature and the feature constitutes the symmetric uncertainty plot. This reflects spatially unreliable or inconsistent regions; then, through uncertainty diagrams... Target-scale guided features derived from the first stage of adaptive weighted fusion detail enhancement and upsampled LST features Spatial domain fusion characteristics Represented as: in, This indicates a special splicing operation.
[0016] Frequency domain embedding: upsampling LST features and the first stage of detailed enhancement of target scale guidance features After splicing, the complex spectrum is projected onto the frequency domain using FFT. Decomposed into amplitude spectrum and phase spectrum In polar coordinates.
[0017] Among them, the amplitude spectrum encodes the global thermal baseline and anomalous amplitude, while the phase spectrum controls spatial structure information; these are respectively transmitted via network. and Decouple amplitude and phase nonlinear fusion: in, Indicates the result after FFT and The amplitude spectrum, Indicates the result after FFT and phase spectrum, This indicates the characteristics of the fused amplitude spectrum. This indicates the phase spectrum characteristics after fusion; express Convolution operation, This represents the leaky modified linear unit activation function.
[0018] Frequency domain features are obtained by recombination using Euler's formula. Frequency domain fusion features are obtained through inverse FFT transformation. : in, It represents the imaginary unit.
[0019] The final integration operation of DEIM is as follows: splicing spatial domain fusion features. and frequency domain fusion features Obtain splicing features Attention weights are obtained through channel attention mechanism. : in, For global average pooling, It is a fully connected layer. This is the Sigmoid activation function.
[0020] The first stage of fusion characteristics was finally obtained. : in, This is a convolution operation.
[0021] The DEIM used in steps 3.2 and 3.3 undergoes the same calculation process.
[0022] Preferably, in step 4.1, the gated guided feature refiner (G2FR) operates as follows: Third-stage fusion feature Upsampling is performed to obtain upsampled DEIM features. The target scale guidance features of the second-stage detail enhancement are They are projected into a shared latent space through point convolution, and after aligning the distribution, global spatial integration is performed to obtain global thermodynamic context features. : in, Indicates the first line, number Global thermodynamic context features of columns Upsampled DEIM features after projection Indicates the first line, number Global thermodynamic context features of columns The second-stage detail enhancement of the target scale-guided features after projection. Represents global thermodynamic context features The rows and columns.
[0023] Channel confidence scores are generated using a multilayer perceptron (MLP) and softmax. : The initial enhancement features are: in, This represents the upsampled DEIM features after projection. This represents the target scale-guided feature for the second stage of detail enhancement after projection.
[0024] The secondary feature enhancement in step 4.2 follows the same calculation process.
[0025] Preferably, in step 5, the loss function as follows: in, The number of training samples for the source-scale LST image; For the first The target-scale LST image predicted for each sample. For the first A sample of real target-scale high-resolution LST images. The GFENet network is driven for structure reconstruction and intensity calibration by minimizing this loss function.
[0026] The beneficial effects of this invention are: (1) The Hybrid Gradient Attention Block (MGAB) proposed in this invention extracts mixed gradients through four dedicated differential convolutions and is combined with a multi-level attention mechanism (channel, spatial, pixel attention), which can accurately recover sharp thermal boundaries while ignoring irrelevant surrounding context, effectively solving the problem of over-smoothing in existing methods.
[0027] (2) The dual-domain embedded integration module (DEIM) proposed in this invention performs feature fusion in both the spatial and frequency domains. By decoupling the amplitude (radiation intensity) and phase (spatial structure) through spectral decomposition, it avoids the entanglement between radiation intensity and geometric structure and effectively prevents spectral reconstruction distortion.
[0028] (3) The gated guided feature refiner (G2FR) proposed in this invention effectively alleviates the artifact problem caused by simple upsampling by adaptively balancing thermal semantic preservation and spatial detail injection through a learnable gating mechanism.
[0029] (4) Extensive experiments on the GrokLST benchmark dataset demonstrate that this invention achieves state-of-the-art performance on both ×4 and ×8 downscaling tasks. Compared to the suboptimal method MoCoLSK, the RMSE is reduced by 0.0351 on the ×8 downscaling task. Furthermore, this invention exhibits excellent generalization ability across four heterogeneous landforms: bare land / urban, mountainous, vegetated, and water bodies. Attached Figure Description
[0030] Figure 1 This is a flowchart of the surface temperature downscaling method based on guided feature enhancement provided by the present invention. Figure 2 This is a network model diagram of the surface temperature downscaling method based on guided feature enhancement provided by the present invention; Figure 3 This is a schematic diagram of the hybrid gradient attention (MGAB) structure in this invention; Figure 4 This is a schematic diagram of the structure of the Dual Domain Embedded Integration Module (DEIM) in this invention; Figure 5 This is a schematic diagram of the structure of the gated guided feature refiner (G2FR) in this invention; Figure 6The results are a qualitative comparison and visualization of the present invention with other methods, wherein (a) is a real reference image, (b) is the result obtained by the DSRN method, (c) is the result obtained by the FDKN method, (d) is the result obtained by the DACG method, (e) is the result obtained by the SUFT method, (f) is the result obtained by the MoKoLSK method, and (g) is the result obtained by the GFENet method of the present invention. Detailed Implementation
[0031] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.
[0032] Figure 1 This is a flowchart illustrating the workflow of the surface temperature downscaling method based on guided feature enhancement provided by this invention. Figure 2 As shown, the entire GFENet network of this invention consists of three main stages: feature extraction, guided feature fusion, and guided LST reconstruction, including the following steps: Step 1: Input Initialization. Provide source-scale surface temperature data. and target scale guided data The guiding data consists of multispectral bands and index features at 10 target scales, including spectral bands such as deep blue, blue, green, red, shortwave infrared, and near-infrared, as well as index features such as digital elevation model, normalized difference water index, normalized difference vegetation index, and normalized difference water-vegetation index.
[0033] Step 2, Feature Extraction. Extract the source-scale LST data. and target-scale multimodal guided data Input them into the convolutional feature extractor respectively.
[0034] Step 2.1, LST Data Feature Extraction. Use convolutional layers to extract source-scale LST features, obtaining... .
[0035] Step 2.2: Guided Data Feature Extraction. Convolutional layers are used to extract guided features at the target scale, resulting in... .
[0036] Step 3: Guided Feature Fusion Process. This stage includes two core components: Hybrid Gradient Attention Block (MGAB) and Dual Domain Embedding Integration Module (DEIM). It achieves source-scale LST features through three stages. and target scale guided features Multi-level integration.
[0037] Step 3.1, First Fusion Stage. The source-scale LST features are then... Input a residual module to extract source-scale LST features for the first stage. At the same time, the target scale guides the features. Inputting Hybrid Gradient Attention Block (MGAB) yields the target scale-guided features for the first stage of detail enhancement. ;like Figure 3 As shown, MGAB operates through two coordinated steps: hybrid gradient extraction and multi-level attention mechanism.
[0038] Hybrid Gradient Extraction: The input feature map undergoes hybrid gradient extraction using four specialized differential convolutions: horizontal and vertical differential convolutions approximate the first-order gradient using learnable one-dimensional convolution kernels; the central differential convolution approximates the Laplacian operator through zero-sum constraints; and the diagonal differential convolution captures diagonal variations by fusing the convolution kernel with its 90° rotation. These four convolutions are then integrated into detail-enhancing features using learnable fusion weights. .
[0039] Multi-level attention mechanism: This invention simultaneously generates channel and spatial attention maps, modeling inter-channel relationships and spatial dependencies respectively. Channel attention is generated through global average pooling and a multilayer perceptron (MLP) structure; spatial attention is generated by concatenating channel-dimensional average pooling and max pooling followed by a 7×7 convolution. Finally, through a pixel attention mechanism, the joint channel-spatial attention is interleaved with the feature map and then grouped convolutions to generate pixel-level attention weights, ensuring strict channel-specific refinement.
[0040] The dual-domain embedding ensemble module (DEIM) is used to integrate the source-scale LST features from the first stage. Target-scale guided features with enhanced detail The fusion process is performed to obtain the fusion characteristics of the first stage. ;like Figure 4 As shown, DEIM performs feature fusion simultaneously in the spatial and frequency domains: Spatial domain embedding: using dense projection to... Upsampling Calculate the symmetric uncertainty diagram It comes from adaptive weighted fusion guided features and upsampled LST features.
[0041] Frequency domain embedding: After concatenating the features, project them into the frequency domain using FFT, decompose them into amplitude spectrum (encoding global thermal baseline and anomalous amplitude) and phase spectrum (controlling spatial structure information). Optimize energy conservation and structure fidelity by decoupling nonlinear mapping, and obtain frequency domain fused features by inverse FFT and 1×1 convolution after recombination.
[0042] Finally, the spatial and frequency domain features are fused and stitched together, and then aggregated through channel attention and projection to form the first-stage DEIM output. .
[0043] Step 3.2, Second Fusion Stage. This involves fusing the features from the first stage. The input is fed into the next residual module and the source-scale LST features of the second stage are extracted. Simultaneously, the target scale of the first-stage detail enhancement guides the features. Inputting the hybrid gradient attention block (MGAB) of the second stage yields the target scale-guided features for second-stage detail enhancement. ; Construct the second-stage dual-domain embedded integration module (DEIM) to achieve... and The fusion process, DEIM outputs the fusion features of the second stage. .
[0044] Step 3.3, the third fusion stage. This involves integrating the fusion features from the second stage. The input is fed into the next residual module and the source-scale LST features of the third stage are extracted. Simultaneously, the target scale of the second-stage detail enhancement guides the features. Inputting the third-stage Hybrid Gradient Attention Block (MGAB) yields the third-stage detail-enhanced target-scale guided features. ; Construct the third-stage dual-domain embedded integration module (DEIM) to achieve... and The fusion of DEIM outputs the fusion features of the third stage. .
[0045] Step 4: Guide the LST reconstruction process. A gated guided feature refiner (G2FR) is used to refine the fused features. Enhancement is performed, and then the target-scale LST image is output, including two stages: primary enhancement and secondary enhancement.
[0046] Step 4.1, Initial Enhancement Stage. For example... Figure 5 As shown, G2FR acts as a semantic gate to dynamically filter non-physical texture artifacts. The upsampled... Target-scale guided features and second-stage detail enhancement After projection onto the shared latent space, global thermodynamic context features are calculated. Channel confidence scores α are generated using a multilayer perceptron (MLP) and softmax. Adaptive fusion using a gating mechanism yields the initial enhanced feature representation. .
[0047] Step 4.2, Secondary Enhancement Stage. The features enhanced in the initial stage... Input an MGAB module, its output and target scale-guided features Simultaneously inputting a new gated guided feature refiner (G2FR) yields secondary enhanced features. .
[0048] Step 4.3, Target-scale LST reconstruction. The secondary enhancement features... The image is reconstructed using convolutional layers, and the reconstruction result is added to the bicubic upsampling result of the source-scale LST to obtain the final target-scale LST image. .
[0049] Step 5: Loss Function Optimization. Optimize using the Mean Squared Error (MSE) loss function, training with the Adam optimizer for 200 epochs at a learning rate starting from 1×10⁻⁶. -4 Gradually decay to 1×10 -7 A smooth scheduling strategy was employed. The model was trained on a single NVIDIA A40 GPU.
[0050] Step 6: Use the optimized GFENet network to perform land surface temperature downscaling inference. Input the source scale LST to be downscaled and its corresponding target scale guiding data into the network. After three steps, namely feature extraction, guided feature fusion (three-stage fusion), and guided LST reconstruction (two-stage enhancement), the target scale LST image is reconstructed.
[0051] To verify the effectiveness of the three key components (MGAB, DEIM, and G2FR) in GFENet, ablation experiments were conducted. As shown in Table 1, removing MGAB (Model 1) significantly increased the RMSE error from 0.7680 to 0.8298, confirming the indispensability of explicit gradient modeling for depicting anisotropic thermal boundaries. Removing DEIM (Model 2) caused the most severe performance degradation (RMSE deteriorated to 0.8481), highlighting the crucial role of dual-domain fusion in maintaining spatial and frequency consistency. Replacing G2FR with a static fusion strategy (Model 3) increased the RMSE error to 0.8201, indicating that guided adaptive gating mechanisms are essential for resolving visual artifacts during the reconstruction phase. Model 4 represents the complete GFENet model of this invention.
[0052] This invention compares with 16 state-of-the-art downscaling methods, including those without target-scale guided data (EDSR, RCAN, SRFormer, ACT, DCTLSA, HiT-SIR) and those with target-scale guided data (MSG-Net, SVLRM, P2P, DSRN, FDSR, DKN, FDKN, DAGF, SUFT, MoCoLSK). For a fair comparison, the same feature encoder was used.
[0053] As shown in Tables 2 and 3, GFENet outperforms other competing methods on all evaluation metrics for the 4x and 8x downscaling tasks of the GrokLST dataset. On the 4x downscaling task, the RMSE reaches 0.5560, a decrease of 0.003 compared to the second-best method, MoCoLSK (0.5590). On the 8x downscaling task, the performance improvement is even more significant, with an RMSE of 0.7680, a decrease of 0.0351 compared to MoCoLSK (0.8031), while the CC reaches 0.9564 and the RSD decreases to 0.0397. This trend verifies the robustness of the proposed method at different scales. Figure 6 The reconstruction results of various methods on the images shown in Examples 1-4 are presented. Figure 6 In each row, (a) corresponds to the original image of the four images shown in Examples 1-4, (b) is the image obtained after processing the four images shown in Examples 1-4 using the DSRN method, (c) is the image obtained after processing the four images shown in Examples 1-4 using the FDKN method, (d) is the image obtained after processing the four images shown in Examples 1-4 using the DACG method, (e) is the image obtained after processing the four images shown in Examples 1-4 using the SUFT method, (f) is the image obtained after processing the four images shown in Examples 1-4 using the MoKoLSK method, and (g) is the image obtained after processing the four images shown in Examples 1-4 using the GFENet method of this invention. In terms of qualitative comparison, GFENet demonstrates superior capabilities in texture restoration and structure preservation. Unlike the oversmoothing or checkerboard artifacts commonly found in competing methods, the method of this invention successfully reconstructs sharp thermal boundaries and recovers fine-grained local anomalies, minimizing thermodynamic deviations between heterogeneous topography.
[0054] Table 1. Ablation Experiment Results of Three Key Components in this Invention
[0055] Table 2 Comparison of the results of this invention and other downscaling methods on the GrokLST dataset with 4x downscaling.
[0056] Table 3 Comparison of the results of this invention and other downscaling methods on the GrokLST dataset for 8x downscaling.
Claims
1. A method for downscaling land surface temperature based on guided feature enhancement, characterized in that, Includes the following steps: Step 1: Input Initialization: Given source-scale land surface temperature data and target scale guided data ,in, Indicates the spatial height at the source scale. Indicates the spatial width at the source scale. Indicates the scaling factor. Indicates the number of pilot channels. Represents the real number field; Step 2, Input Feature Extraction: Extract the source-scale LST data. and target scale guided data Input two convolutional feature extractors respectively to extract source-scale LST features. and target scale guided features ; Step 3, Guided Feature Fusion Process: Combine source-scale LST features and target scale guided features The input-guided feature fusion module includes three fusion stages; Step 3.1, First Fusion Stage: Integrating Source-Scale LST Features Input a residual module ResG to extract source-scale LST features for the first stage. At the same time, the target scale guides the features. Inputting the Hybrid Gradient Attention Block (MGAB) yields the target scale-guided features for the first stage of detail enhancement. The dual-domain embedding integration module DEIM is used to integrate the source-scale LST features from the first stage. Target-scale guided features with enhanced detail The fusion process is performed to obtain the fusion characteristics of the first stage. ; Step 3.2, Second Fusion Stage: The fusion features from the first stage are then... The input is fed into the next residual module, ResG, to extract the source-scale LST features for the second stage. Simultaneously, the target scale of the first-stage detail enhancement is guided by the feature. Input the next hybrid gradient attention block (MGA) to obtain the target scale-guided features for second-stage detail enhancement. Construct a dual-domain embedding integration module (DEIM) to integrate the source-scale LST features from the second stage. Target-scale guided features with enhanced detail The fusion process is performed to obtain the fusion characteristics of the second stage. ; Step 3.3, Third Fusion Stage: Integrating the fusion features from the second stage The input is fed into the next residual module, ResG, to extract the source-scale LST features for the third stage. Simultaneously, the target scale guidance feature for the second stage of detail enhancement is... Input the next hybrid gradient attention block MGAB to obtain the target scale-guided features for third-stage detail enhancement. Construct a dual-domain embedding integration module (DEIM) to integrate the source-scale LST features from the third stage. Target-scale guided features with enhanced detail The fusion process yields the fusion characteristics of the third stage. ; Step 4, Guided LST Reconstruction Process: The guided fusion features are enhanced using the gated guided feature refiner G2FR, and then a high-resolution LST image at the target scale is output. Step 4.1, Initial Enhancement Stage: Incorporating the fusion features from the third stage. Upsampling is performed, and the results are used to enhance the target scale-guided features in the second stage of detail enhancement. Simultaneously input the gated guided feature refiner G2FR to obtain the initial enhanced features. ; Step 4.2, Secondary Enhancement Stage: The initial enhanced features... Input an MGAB module, its output and the target scale-guided features of the first stage of detail enhancement. Simultaneously input another gated guided feature refiner G2FR to obtain secondary enhanced features. ; Step 4.3, Target-scale LST reconstruction: The secondary enhancement features are then reconstructed. Input an MGAB module and add its output to the bicubic upsampling result of the source-scale LST to obtain the final predicted target-scale LST image. ; Step 5: Loss Function Optimization: The GFENet network formed in Steps 1-4 is optimized end-to-end using the Mean Squared Error (MSE) loss function to minimize the predicted target-scale LST image. Compared with real target-scale LST images Pixel-level differences between them; Step 6: Perform surface temperature downscaling inference using the optimized GFENet network: The source-scale LST image to be downscaled and its corresponding target-scale guiding data are input into the optimized GFENet network. After three stages—feature extraction, guided feature fusion, and guided LST reconstruction—the target-scale LST image is output.
2. The surface temperature downscaling method based on guided feature enhancement according to claim 1, characterized in that, In step 3.1, the specific operation of the Hybrid Gradient Attention Block (MGAB) is as follows: Input target scale guided features First, four types of differential convolution—horizontal differential convolution, vertical differential convolution, central differential convolution, and diagonal differential convolution—are used to perform mixed gradient extraction to obtain differential convolution features. ; The horizontal and vertical differential convolutions use learnable one-dimensional convolution kernels. To approximate the first-order gradient, central difference convolution approximates the Laplacian operator by imposing a zero-sum constraint on the weights, while diagonal difference convolution approximates the Laplacian operator by applying a zero-sum constraint to the weights. Rotation fusion captures diagonal variations, and four convolutional methods utilize learnable weights. and bias By fusing the features, differential convolution features are obtained. : in : in, These represent the weight matrices for horizontal, vertical, central, and diagonal difference convolutions, respectively. in : in, These represent the biases of the horizontal, vertical, central, and diagonal difference convolutions, respectively. Then the differential convolution features The input multi-level attention mechanism operates as follows: Channel attention (CA): for... Global average pooling (GAP) is performed, and then channel-level attention weights are captured using a multilayer perceptron (MLP). ,in, Represents the number of feature channels; Spatial Attention (SA): for Two two-dimensional feature maps are obtained by performing average pooling and max pooling along the channel dimension respectively. These two feature maps are then concatenated and processed... Convolution output space attention weights Pixel attention PA: Attention weights are obtained by combining channel attention and spatial attention. Through interleaving operations and Concatenate, then use grouped convolution Ensure that one feature channel corresponds to one convolution operation to obtain pixel-level attention. : in, Use the Sigmoid activation function; Pixel-level attention Sum of differential convolution features Perform multiplication and add the original number. The final output is the target scale-guided feature for the first stage of detail enhancement. : in, This is a multiplication operation; The MGAB used in steps 3.2 and 3.3 performs the same calculation process.
3. The surface temperature downscaling method based on guided feature enhancement according to claim 1, characterized in that, In step 3.1, the specific operations of the dual-domain embedded integration module DEIM are as follows: First, the source-scale LST features obtained in step 2 are... The first-stage LST features are obtained after passing through a residual module. Through dense projection Upsampling is performed to obtain upsampled LST features. Then, spatial domain embedding and frequency domain embedding are performed separately, as follows: Spatial domain embedding: for Perform a horizontal flip to obtain mirrored features, then calculate... The residual between the horizontal flip feature and the feature constitutes the symmetric uncertainty plot. Then, through uncertainty diagrams... Target-scale guided features derived from the first stage of adaptive weighted fusion detail enhancement and upsampled LST features Spatial domain fusion characteristics Represented as: in, This indicates a special splicing operation; Frequency domain embedding: upsampling LST features and the first stage of detailed enhancement of target scale guidance features After splicing, the complex spectrum is projected onto the frequency domain using FFT. Decomposed into amplitude spectrum and phase spectrum polar coordinates; Among them, the amplitude spectrum encodes the global thermal baseline and anomalous amplitude, while the phase spectrum controls spatial structure information; these are respectively transmitted via network. and Decouple amplitude and phase nonlinear fusion: in, Indicates the result after FFT and The amplitude spectrum, Indicates the result after FFT and phase spectrum, This indicates the characteristics of the fused amplitude spectrum. This indicates the phase spectrum characteristics after fusion; express Convolution operation, This represents the leaky modified linear unit activation function; Frequency domain features are obtained by recombination using Euler's formula. Frequency domain fusion features are obtained through inverse FFT transformation. : in, Represents the imaginary unit; The final integration operation of DEIM is as follows: splicing spatial domain fusion features. and frequency domain fusion features Obtain splicing features Attention weights are obtained through channel attention mechanism. : in, For global average pooling, It is a fully connected layer. Use the Sigmoid activation function; The first stage of fusion characteristics was finally obtained. : in, This is a convolution operation; The DEIM used in steps 3.2 and 3.3 undergoes the same calculation process.
4. The surface temperature downscaling method based on guided feature enhancement according to claim 1, characterized in that, In step 4.1, the specific operation of the gated guided feature refiner G2FR is as follows: The third stage of fusion feature... Upsampling is performed to obtain upsampled DEIM features. The target scale guidance features of the second-stage detail enhancement are They are projected into a shared latent space through point convolution, and after aligning the distribution, global spatial integration is performed to obtain global thermodynamic context features. : in, Indicates the first line, number Global thermodynamic context features of columns Upsampled DEIM features after projection Indicates the first line, number Global thermodynamic context features of columns The second-stage detail enhancement of the target scale-guided features after projection. Represents global thermodynamic context features Rows and columns; Channel confidence scores are generated using a multilayer perceptron (MLP) and softmax. : The initial enhancement features are: in, This represents the upsampled DEIM features after projection. This represents the target scale guidance feature for the second stage of detail enhancement after projection; The secondary feature enhancement in step 4.2 follows the same calculation process.
5. The surface temperature downscaling method based on guided feature enhancement according to claim 1, characterized in that, In step 5, the loss function as follows: in, The number of training samples for the source-scale LST image; For the first The target-scale LST image predicted for each sample. For the first Each sample contains real target-scale high-resolution LST images; the GFENet network is driven by minimizing the loss function for structure restoration and strength calibration.
Citation Information
Patent Citations
MRI arbitrary super-resolution reconstruction method based on frequency spectrum-space linear attention
CN122222822A
Method for fusing infrared light and visible light images
WO2025103079A1