Progressive soil moisture downscaling method and model based on hybrid attention mechanism

CN122157026BActive Publication Date: 2026-08-07SOUTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHWEST UNIV
Filing Date
2026-03-24
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

若直接将空间分辨率为0.36°的原始SMAP被动微波土壤水分数据直接降尺度至0.01°时,较大的尺度差异容易导致空间结构失真,从而影响降尺度结果的稳定性和精度,限制了其应用

Benefits of technology

[0054] This invention, through the design of a PDHAM downscaling model and a series of operations such as progressive staged modeling, fully utilizes the complementary information of SMAP soil moisture data and various auxiliary image data to achieve a gradual improvement in the spatial resolution of the original SMAP soil moisture data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157026B_ABST
    Figure CN122157026B_ABST
Patent Text Reader

Abstract

The application discloses a progressive soil moisture downscaling method and model based on a hybrid attention mechanism, and relates to the technical fields of remote sensing image processing and artificial intelligence. The application realizes the gradual improvement of the spatial resolution of the original SMAP soil moisture data through the design of the PDHAM downscaling model, and has the following main advantages: 1) the progressive phased modeling is adopted, the instability caused by large-scale cross downscaling is relieved, and the practicability of the model is significantly improved; 2) in the aspect of feature processing, the model improves the fine downscaling ability of the downscaling model from two aspects of “key feature selection” and “detail information strengthening” by integrating multiple attention mechanisms and multi-scale deep convolution modules, so that the accuracy of the soil moisture downscaling result and the stability of the soil moisture downscaling model are improved. The project of the application is a national key research and development plan, the project number is 2024YFF1307700, and the project name is “Research and Demonstration of Ecological Protection and Restoration Technology for Highway Network in Extreme Environmental Conditions”.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of remote sensing image processing and artificial intelligence, specifically to a progressive soil moisture downscaling method and model based on a hybrid attention mechanism. Background Technology

[0002] Soil moisture is a key hydrological variable connecting the atmosphere and the Earth's surface system, playing a crucial role in global climate change, drought prediction, and agricultural irrigation. Traditional soil moisture monitoring methods provide localized, high-precision data, but their limited spatial coverage and high resource consumption make them unsuitable for large-scale and long-term data collection. With the development of remote sensing technology, satellite sensors have become an effective supplement, enabling large-scale, long-term monitoring. Currently, many satellite platforms deploy microwave systems for soil moisture monitoring, such as SMAP and SMOS. Microwave remote sensing technology is divided into active microwave and passive microwave. Active microwave has higher spatial resolution but is affected by vegetation, while passive microwave has higher temporal resolution but lower spatial resolution. Although passive microwave products are widely used globally, their spatiotemporal continuity and spatial resolution at regional scales remain limited. To improve these aspects, researchers have proposed soil moisture downscaling methods, aiming to upscale low-resolution data to high-resolution data, thereby enhancing their application value in agriculture and ecology. Existing downscaling methods can be categorized into three types: physical mechanism-based methods, parametric statistical methods, and learning-based methods.

[0003] Physically based soil moisture downscaling methods rely on surface physical processes to construct quantitative relationships between soil moisture and other surface variables. The core of these methods is utilizing physical models to transform low-resolution soil moisture data into high-resolution data. Commonly used physical models include surface radiative transfer models, thermal infrared radiative transfer models, and surface process models, which can accurately describe the interactions between soil moisture and auxiliary variables. By combining these physical processes, physically based downscaling methods can achieve higher spatial resolution soil moisture retrieval, thus more accurately capturing the distribution of soil moisture at the surface. For example, L-band radiative transfer models and thermal infrared radiative transfer models are often used to establish the relationship between soil moisture and surface radiation, thereby downscaling low-resolution soil moisture products to higher spatial resolution. These methods not only improve spatial resolution but also effectively enhance the spatiotemporal continuity of soil moisture, especially in agricultural and ecological monitoring, providing more accurate soil moisture data. However, due to the complexity of the physical models and their high sensitivity to various input parameters, these methods still face some limitations in practical applications.

[0004] Parametric statistical methods extract detailed information from various high-resolution auxiliary data sources, establish empirical statistical models of the relationship between auxiliary parameters and microwave soil moisture at low resolution, and then apply these models to high-resolution data to obtain high-resolution soil moisture products. Several parametric statistical methods have been developed to improve the spatial resolution of soil moisture products. Typical models include linear regression, second-order polynomial regression, and geographically weighted regression. Knipper et al. proposed a SMAP and SMOS soil moisture downscaling method using second-order polynomial regression, combining it with MODIS data to obtain soil moisture data with a spatial resolution of 1 km. Song Peilin et al. proposed an improved surface soil moisture downscaling framework using geographically weighted regression, integrating AMSR 2 soil moisture data with the MODIS LST / NDVI dataset, achieving higher spatial resolution soil moisture data even under cloudy conditions. Although these methods are simple and efficient, their robustness is often affected in heterogeneous regions due to their sensitivity to the accuracy, continuity, and correlation of auxiliary parameters.

[0005] With the rapid development of machine learning technology, learning-based soil moisture downscaling methods have gradually become a research focus, mainly represented by machine learning and deep learning. These methods have shown significant advantages in handling nonlinear relationships and fusing multi-source remote sensing data, providing new ideas for overcoming the limitations of traditional statistical and physical models under complex, multi-scale surface conditions. Studies have shown that machine learning algorithms such as artificial neural networks (ANN), support vector machines (SVM), and random forests have achieved good results in soil moisture downscaling, especially artificial neural networks, which show superior performance in accuracy and robustness. However, machine learning methods often employ point-to-point modeling, making it difficult to fully utilize the spatial correlation and structural information of soil moisture. In contrast, deep learning methods have shown significant advantages in improving the accuracy of soil moisture downscaling and characterizing spatial features. Studies have shown that deep learning models can better preserve spatial structural features, especially exhibiting stronger stability and accuracy under complex surface conditions. Convolutional neural networks (CNN) and multimodal deep learning networks (MMNet) have been successfully applied to soil moisture downscaling, demonstrating the potential of deep learning in multi-source data fusion and spatial relationship representation. In recent years, deep learning has been widely applied due to its advantages in multi-scale feature extraction and complex nonlinear relationship modeling. However, if the original SMAP passive microwave soil moisture data with a spatial resolution of 0.36° is directly downscaled to 0.01°, the large scale difference can easily lead to spatial structure distortion, thereby affecting the stability and accuracy of the downscaling results and limiting its application.

[0006] Therefore, a new solution is needed to address the above problems. Summary of the Invention

[0007] The purpose of this invention is to provide a progressive soil moisture downscaling method and model based on a hybrid attention mechanism to solve the technical problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a progressive soil moisture downscaling model based on a hybrid attention mechanism, wherein the progressive soil moisture downscaling model based on a hybrid attention mechanism is a PDHAM downscaling model, and the PDHAM downscaling model includes a preliminary spatial downscaling network model and a fine spatial downscaling network model.

[0009] The preliminary spatial downscaling network model outputs medium-resolution soil moisture data with a spatial resolution of 6km, which serves as the input soil moisture data for the fine spatial downscaling network model. The preliminary spatial downscaling network model includes a style multifactor cross attention module (SMCA), a global cross attention module (GCA), a residual channel spatial attention module (RCSA), and a content-guided attention fusion module (CGAF).

[0010] The refined spatial downscaling network model further introduces a residual dense connection Res2Net module based on the preliminary spatial downscaling network model.

[0011] The style multi-factor cross-attention module calibrates the weights of spatial feature information on different auxiliary data one by one using low-resolution soil moisture data;

[0012] The global cross-attention module is used to treat all auxiliary data as a whole and calibrate it with the low-resolution soil moisture data.

[0013] The residual channel spatial attention module is used to perform residual channel spatial attention processing on the feature maps output by SMCA and GCA.

[0014] The content-guided attention fusion module is used to effectively fuse features from different perspectives in the above process;

[0015] The Residual Dense Connections Res2Net module significantly improves network performance through multi-scale feature learning.

[0016] A gradual soil moisture downscaling method based on a hybrid attention mechanism includes at least the following steps:

[0017] S1: Collect satellite remote sensing data from multiple data sources, including at least SMAP soil moisture data and auxiliary data, including CLDAS-V2.0 reanalysis data, MODIS surface reflectance data and SRTM digital elevation model;

[0018] S2: Preprocess the satellite remote sensing data to obtain spatiotemporally continuous low-resolution soil moisture data;

[0019] S3: Construct the PDHAM downscaling model described above, input the preprocessed satellite remote sensing data into the PDHAM downscaling model, use the preliminary spatial downscaling network model to perform preliminary downscaling on the low-resolution soil moisture data, and use the fine spatial downscaling network model to perform multi-scale feature processing on the preliminary downscaling results to obtain high-resolution soil moisture data.

[0020] S4: The PDHAM downscaling model is trained based on the combined loss function. After the loss converges, the trained PDHAM downscaling model is used to predict soil moisture data with a spatial resolution of 1 km.

[0021] The combined loss function includes at least the Charbonnier loss function and the marginal loss function.

[0022] Furthermore, the preprocessing of the satellite remote sensing data includes at least the following steps:

[0023] The obtained satellite remote sensing data is processed by stitching, cropping and reprojection, and outliers are removed by low-pass filtering.

[0024] The data after removing outliers was then resampled to 0.06° and 0.36° and normalized.

[0025] Meanwhile, a deep learning network for SMAP soil moisture data reconstruction was used to reconstruct missing SMAP soil moisture data, thereby obtaining spatiotemporally continuous low-resolution soil moisture data.

[0026] Furthermore, the preliminary spatial downscaling process includes at least the following steps:

[0027] The spatial and channel weights of the input data are calculated using a style multi-factor cross-attention module and a global cross-attention module. In this process, the initial auxiliary data is processed by both holistic and stepwise input methods to obtain high-precision feature maps and recalibrated feature maps, which are then regarded as calibrated preliminary feature maps.

[0028] Deep features in high-dimensional data are extracted by residual channel spatial attention module to explore the intrinsic relationship between multiple factors and obtain enhanced feature maps.

[0029] A content-guided attention fusion module is used to fuse the calibrated preliminary feature maps with the enhanced feature maps. By dynamically fusing low-level and high-level feature maps, the model's ability to represent spatial details is improved, thereby enhancing the network's performance and accuracy.

[0030] Furthermore, the style multi-factor cross-attention module includes at least the following steps:

[0031] Each input auxiliary data and low-resolution soil moisture data is convolved to extract features, resulting in corresponding auxiliary data feature maps and low-resolution soil moisture data feature maps. Channel attention and spatial attention are applied to the auxiliary data feature maps and low-resolution soil moisture data feature maps respectively through style modulation and spatial attention mechanisms. In this process, the low-resolution soil moisture features provide accurate channel information, which is used to calculate channel weights to calibrate the channel information of the high-resolution features, while the high-resolution auxiliary data features provide rich detail information, which is used to calculate spatial weights to calibrate the spatial information of the low-resolution soil moisture features.

[0032] In the interaction between auxiliary factors and soil moisture characteristics, the channel attention branch output and the spatial attention branch output are added element by element, and then the cross attention outputs corresponding to each auxiliary factor are cascaded in the channel dimension to construct a joint representation of multi-source auxiliary information.

[0033] Finally, the convolution module is used to reduce the dimensionality of the cascaded high-dimensional features, ultimately generating a comprehensive high-precision feature map.

[0034] Furthermore, the global cross-attention module treats all auxiliary data as a whole and performs mutual calibration with the low-resolution soil moisture data;

[0035] The global cross-attention module inputs auxiliary data and low-resolution soil moisture data into the convolutional layer for feature extraction, and combines them with the Sigmoid activation function to generate weights. The convolutional layer is used to extract features, while the Sigmoid function constrains the output to the (0,1) interval to calculate weights, thereby guiding the interaction between low-resolution and high-resolution features.

[0036] By combining a cross-attention mechanism, a recalibrated feature map is obtained, thereby enhancing the spatial detail information of the output image.

[0037] Furthermore, the residual channel spatial attention module collaborates with the high-precision feature map output by the style multi-factor cross attention module and the recalibrated feature map output by the global cross attention module. The residual channel spatial attention module uses an embedded attention mechanism to further extract and enhance deep feature information, and outputs an enhanced feature map.

[0038] The residual channel spatial attention module combines dense connections and an attention mechanism, where dense connections can enhance feature reuse and improve model efficiency, and the attention mechanism is used to capture local details.

[0039] The residual channel spatial attention module combines a channel attention layer and a spatial attention layer.

[0040] The channel attention layer uses global average pooling, convolution operations, and the sigmoid activation function to learn channel weights, thereby enhancing the ability to represent channel information.

[0041] The spatial attention layer uses convolution operations and the sigmoid activation function to calculate spatial weights, and reweights the network to focus on important spatial features.

[0042] The residual channel spatial attention module combines multiple channel spatial cross attention blocks through dense connections, which significantly enhances the feature representation capability. Finally, residual learning is used to preserve the original features, stabilize the training process, and prevent network feature degradation.

[0043] Furthermore, the application of the content-guided attention fusion module includes at least the following steps:

[0044] The high-precision feature map output by the style multi-factor cross-attention module and the recalibrated feature image output by the global cross-attention module are used as low-level feature images, and the enhanced feature image output by the residual channel spatial attention module is used as high-level feature images. These are then input into the content-guided attention fusion module for initial fusion.

[0045] At this point, channel attention, spatial attention, and pixel attention mechanisms are employed to optimize the fusion process. By calculating the feature importance of the channel and spatial dimensions, the fusion ratio of low-level and high-level features is dynamically adjusted. Subsequently, 1x1 convolution is applied for dimensionality reduction to obtain the final fused feature map. This process significantly improves the accuracy of soil moisture downscaling.

[0046] Furthermore, the application of the refined spatial downscaling network model includes at least the following steps:

[0047] By utilizing the style multi-factor cross-attention module and the global cross-attention module to calculate spatial and channel weights, and by employing both holistic and stepwise input methods for the initial input auxiliary data, preliminary calibration of the feature map is achieved.

[0048] A residual channel spatial attention module is adopted to obtain enhanced feature maps by fitting the intrinsic correlation between multiple bands. At the same time, a residual dense connection Res2Net module (RDRN) is designed to significantly improve network performance through multi-scale feature learning.

[0049] A content-guided attention fusion module is employed to fuse the initially calibrated feature maps with the enhanced feature maps, thereby improving the model's ability to represent spatial details. The aim is to dynamically fuse low-level and high-level feature maps.

[0050] Overall, the refined spatial downscaling network mainly adds a residual dense connection Res2Net module to the initial spatial downscaling network, which learns spatial detail features more meticulously.

[0051] Furthermore, the application of the residual dense connection Res2Net module includes at least the following steps:

[0052] The input feature map is split into three groups according to the channel dimension. Each group is processed separately using a convolutional layer, and then the input feature information is preserved through residual connections. The processed multi-scale features are concatenated, and channel fusion and feature integration are achieved through 1x1 convolution to finally generate a fused feature map.

[0053] Compared with the prior art, the beneficial effects of the present invention are:

[0054] This invention, through the design of a PDHAM downscaling model and a series of operations such as progressive staged modeling, fully utilizes the complementary information of SMAP soil moisture data and various auxiliary image data to achieve a gradual improvement in the spatial resolution of the original SMAP soil moisture data.

[0055] 1) The progressive, phased modeling approach mitigates the instability caused by large-scale scaling down and significantly improves the model's practicality.

[0056] 2) In terms of feature processing, the model improves the fine-grained downscaling capability of the downscaling model by integrating multiple attention mechanisms and multi-scale deep convolution modules from two aspects: "key feature selection" and "detail information enhancement", thereby improving the accuracy of soil moisture downscaling results and the stability of the soil moisture downscaling model. Attached Figure Description

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a schematic diagram of the overall process of the present invention;

[0059] Figure 2 This is a schematic diagram of the overall model of the present invention. Detailed Implementation

[0060] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0061] Example 1:

[0062] Please see Figure 1 The progressive soil moisture downscaling model based on the hybrid attention mechanism, namely the PDHAM downscaling model, includes a preliminary spatial downscaling network model and a fine spatial downscaling network model.

[0063] A preliminary spatial downscaling network model, used to output medium-resolution soil moisture data with a spatial resolution of 6 km, includes:

[0064] A style multi-factor cross-attention module is used to calibrate the weights of spatial feature information on different auxiliary data;

[0065] The global cross-attention module is used to cross-calibrate all auxiliary data with soil moisture data;

[0066] The residual channel spatial attention module is used to perform residual channel spatial attention processing on the feature maps output by the style multi-factor cross attention module and the global cross attention module.

[0067] The content-guided attention fusion module is used to fuse features from the outputs of the style multi-factor cross attention module, the global cross attention module, and the residual channel spatial attention module.

[0068] The refined spatial downscaling network model further adds a residual dense connection Res2Net module to the output of the initial spatial downscaling network model for multi-scale feature learning.

[0069] Example 2:

[0070] Please see Figure 2 A gradual soil moisture downscaling method based on a hybrid attention mechanism includes at least the following steps:

[0071] S1: Collect satellite remote sensing data from multiple data sources. The satellite remote sensing data includes at least SMAP soil moisture data and auxiliary data, including CLDAS-V2.0 reanalysis data, MODIS surface reflectance data and SRTM digital elevation model.

[0072] The specific data sources for this embodiment are shown in the table below:

[0073] Enter variable name Data source Spatiotemporal resolution SMAP soil moisture data https: / / earthexplorer.usgs.gov / 36 kilometers, 1 day CLDAS-V2.0 reanalysis data https: / / data.cma.cn / data / cdcdetail / dataCode / NAFP_CLDAS2.0_NRT.html 6 kilometers, 1 day MODIS Surface Reflectance https: / / data-starcloud.pcl.ac.cn / iearthdata / 500 meters, 1 day Global seamless daily LST data https: / / www.earthdata.nasa.gov / data / catalog / lpcloud-myd11a1-006 / 1 kilometer, 1 day SRTM Digital Elevation Model https: / / srtm.csi.cgiar.org / srtmdata / 90 meters

[0074] S2: Preprocess satellite remote sensing data to obtain spatiotemporally continuous low-resolution soil moisture data;

[0075] S3: Construct the PDHAM downscaling model proposed in the above embodiments, input the preprocessed satellite remote sensing data into the PDHAM downscaling model, use the preliminary spatial downscaling network model to perform preliminary downscaling on the low-resolution soil moisture data, and use the fine spatial downscaling network model to perform multi-scale feature processing on the preliminary downscaling results to obtain high-resolution soil moisture data.

[0076] S4: The PDHAM downscaling model is trained based on the combined loss function. After the loss converges, the trained PDHAM downscaling model is used to predict soil moisture data with a spatial resolution of 1 km.

[0077] The combined loss function includes at least the Charbonnier loss function and the marginal loss function.

[0078] Preprocessing satellite remote sensing data includes at least the following steps:

[0079] The obtained satellite remote sensing data is processed by stitching, cropping and reprojection, and outliers are removed by low-pass filtering.

[0080] The data after removing outliers was then resampled to 0.06° and 0.36° and normalized.

[0081] Meanwhile, a deep learning network for SMAP soil moisture data reconstruction was used to reconstruct missing SMAP soil moisture data, thereby obtaining spatiotemporally continuous low-resolution soil moisture data.

[0082] Performing an initial spatial downscaling process includes at least the following steps:

[0083] The spatial and channel weights of the input data are calculated using a style multi-factor cross-attention module and a global cross-attention module. In this process, the initial auxiliary data is processed by both holistic and stepwise input methods to obtain high-precision feature maps and recalibrated feature maps, which are then regarded as calibrated preliminary feature maps.

[0084] Deep features in high-dimensional data are extracted by residual channel spatial attention module to explore the intrinsic relationship between multiple factors and obtain enhanced feature maps.

[0085] A content-guided attention fusion module is used to fuse the calibrated preliminary feature maps with the enhanced feature maps. By dynamically fusing low-level and high-level feature maps, the model's ability to represent spatial details is improved, thereby enhancing the network's performance and accuracy.

[0086] The style multifactor cross-attention module includes at least the following steps:

[0087] Each input auxiliary data and low-resolution soil moisture data is convolved to extract features, resulting in corresponding auxiliary data feature maps and low-resolution soil moisture data feature maps. Channel attention and spatial attention are applied to the auxiliary data feature maps and low-resolution soil moisture data feature maps respectively through style modulation and spatial attention mechanisms. In this process, the low-resolution soil moisture features provide accurate channel information, which is used to calculate channel weights to calibrate the channel information of the high-resolution features, while the high-resolution auxiliary data features provide rich detail information, which is used to calculate spatial weights to calibrate the spatial information of the low-resolution soil moisture features.

[0088] In the interaction between auxiliary factors and soil moisture characteristics, the channel attention branch output and the spatial attention branch output are added element by element, and then the cross attention outputs corresponding to each auxiliary factor are cascaded in the channel dimension to construct a joint representation of multi-source auxiliary information.

[0089] Finally, the convolution module is used to reduce the dimensionality of the cascaded high-dimensional features, and a comprehensive high-precision feature map is generated.

[0090] The calculation formula is as follows:

[0091]

[0092] In the formula: This represents the output features generated by the style multi-factor attention module; These are the feature map and spatial attention map from the i-th auxiliary data; It represents the spatial attention weight map generated from the i-th auxiliary feature map; This is a feature map derived from low-resolution soil moisture data; This represents the channel modulation factor generated by the style modulation mechanism; n represents the number of auxiliary data. and This represents element-wise multiplication and addition operations.

[0093] The global cross-attention module treats all auxiliary data as a whole and cross-calibrates it with low-resolution soil moisture data;

[0094] Auxiliary data and low-resolution soil moisture data are input into the convolutional layer for feature extraction, and weights are generated by combining the Sigmoid activation function. The convolutional layer is used to extract features, while the Sigmoid function constrains the output to the (0,1) interval to calculate the weights, thereby guiding the interaction between low-resolution and high-resolution features.

[0095] By combining a cross-attention mechanism, a recalibrated feature map is obtained, thereby enhancing the spatial detail information of the output image.

[0096] The residual channel spatial attention module collaborates with the high-precision feature map output by the style multi-factor cross attention module and the recalibrated feature map output by the global cross attention module. The residual channel spatial attention module uses an embedded attention mechanism to further extract and enhance deep feature information, and outputs an enhanced feature map.

[0097] The residual channel spatial attention module combines dense connections and attention mechanisms. Dense connections can enhance feature reuse and improve model efficiency, while attention mechanisms can capture local details.

[0098] The residual channel spatial attention module combines the channel attention layer and the spatial attention layer.

[0099] The channel attention layer uses global average pooling, convolution operations, and the sigmoid activation function to learn channel weights, thereby enhancing the ability to represent channel information;

[0100] The spatial attention layer uses convolution operations and the sigmoid activation function to calculate spatial weights, and reweights the network to focus on important spatial features.

[0101] The residual channel spatial attention module combines multiple channel spatial cross attention blocks through dense connections, which significantly enhances the feature representation capability. Finally, residual learning is used to preserve the original features, stabilize the training process, and prevent network feature degradation.

[0102] The application of a content-guided attention fusion module includes at least the following steps:

[0103] The high-precision feature map output by the style multi-factor cross-attention module and the recalibrated feature image output by the global cross-attention module are used as low-level feature images, and the enhanced feature image output by the residual channel spatial attention module is used as high-level feature images. These are then input into the content-guided attention fusion module for initial fusion.

[0104] At this point, channel attention, spatial attention, and pixel attention mechanisms are used to optimize the fusion process. By calculating the feature importance of the channel dimension and spatial dimension, the fusion ratio of low-level and high-level features is dynamically adjusted.

[0105] Subsequently, 1x1 convolution is applied for dimensionality reduction to obtain the final fused feature map. This process significantly improves the accuracy of soil moisture downscaling, as shown in the following expression:

[0106]

[0107] In the formula, This represents the low-level features of the network encoder. This represents the high-level features of the decoder section; This represents the spatial weights calculated based on the content-driven attention mechanism; It is a 1x1 convolutional layer used to project fused features into the final output space.

[0108] The application of a fine-grained spatial downscaling network model includes at least the following steps:

[0109] By utilizing the style multi-factor cross-attention module and the global cross-attention module to calculate spatial and channel weights, and by employing both holistic and stepwise input methods for the initial input auxiliary data, preliminary calibration of the feature map is achieved.

[0110] A residual channel spatial attention module is adopted to obtain enhanced feature maps by fitting the intrinsic correlation between multiple bands. At the same time, a residual dense connection Res2Net module (RDRN) is designed to significantly improve network performance through multi-scale feature learning.

[0111] A content-guided attention fusion module is employed to fuse the initially calibrated feature maps with the enhanced feature maps, thereby improving the model's ability to represent spatial details. The aim is to dynamically fuse low-level and high-level feature maps.

[0112] Overall, the refined spatial downscaling network mainly adds a residual dense connection Res2Net module to the initial spatial downscaling network, which learns spatial detail features more meticulously.

[0113] The application of the Res2Net module with residual dense connections includes at least the following steps:

[0114] The input feature map is split into three groups according to the channel dimension. Each group is processed separately using a convolutional layer, and then the input feature information is preserved through residual connections. The processed multi-scale features are concatenated, and channel fusion and feature integration are achieved through 1x1 convolutions to finally generate a fused feature map, as shown in the following expression:

[0115]

[0116] In the formula, It is the final output feature map after applying 1x1 convolution and residual connection; It is a 1x1 convolutional layer. It combines features from feature maps at different scales; This represents the input features.

[0117] In summary, the Residual Dense Connection Res2Net module can capture the subtle differences in the input feature map, thereby extracting and fusing key information from the input data.

[0118] The combined loss function used to constrain the network training process consists of the Charbonnier loss function and the marginal loss function, and its expression is as follows:

[0119]

[0120] In the formula, and For regularization parameters, The Charbonnier loss function and the marginal loss function are respectively determined by the following formulas:

[0121]

[0122]

[0123] In the formula, This indicates the number of sample pairs in the training set. This indicates soil moisture labeling data. This represents the output value of the downscaling network. This is a constant term initially set to 0.001. Combining the two loss functions allows for simultaneous optimization of pixel-level precision and edge information;

[0124] Furthermore, in the early stages of training, the global loss is given a higher weight to guide the model in learning the overall spatial distribution features. As training progresses, the weight of the edge loss gradually increases, prompting the model to pay more attention to spatial details and boundary features. This dynamic weight adjustment strategy allows the model to maintain overall structural consistency while significantly improving detail performance and local accuracy. The core lies in the modulation of two coefficients: and Through extensive testing, the weighting scheme was determined as follows:

[0125]

[0126] In summary:

[0127] This invention, based on deep learning, designs a Progressive Downscaling Framework With Hybrid Attention Mechanisms (PDHAAM) model and method for soil moisture downscaling. This model mitigates the instability caused by large-scale downscaling by progressively reducing spatial resolution differences through staged modeling, thereby achieving a gradual downscaling of SMAP soil moisture data from approximately 0.36° resolution to 0.01° resolution. First, this invention designs a two-stage convolutional neural network structure consisting of preliminary spatial downscaling and refined spatial downscaling. Second, it introduces multiple cross-attention structures to establish intrinsic relationships between bands. Subsequently, through staged modeling, it gradually reduces the differences between different spatial scales, thus achieving refined downscaling of soil moisture data from coarse resolution to high resolution. Furthermore, it utilizes a content-guided attention fusion module to adaptively allocate weights to the overall feature map. Finally, it designs a combined loss function to constrain the network training process, including Charbonnier loss and edge loss. This method can fully utilize the complementary information of SMAP soil moisture imagery and auxiliary factor imagery to generate soil moisture images with high temporal resolution.

[0128] Compared to existing models such as Deep Belief Network (DBN), BackPropagation Neural Network (BPNN), Residual Dense Network (RDN), and Hybrid Attention based Residual Dense Network for Soil Moisture Downscaling (HAND), the PDHAM model can fully utilize the advantages of SMAP soil moisture data and various auxiliary remote sensing images to achieve high spatiotemporal resolution SMAP soil moisture data. Under different surface conditions, the PDHAM method exhibits superior robustness and adaptability compared to traditional methods. Furthermore, this downscaling model effectively expands the potential application scenarios of soil moisture downscaling technology, providing high-quality data support for Earth's environmental monitoring.

[0129] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A progressive soil moisture downscaling model based on a hybrid attention mechanism, characterized by: The progressive soil moisture downscaling model based on the hybrid attention mechanism is the PDHAM downscaling model, which includes a preliminary spatial downscaling network model and a fine spatial downscaling network model. The preliminary spatial downscaling network model is used to output medium-resolution soil moisture data with a spatial resolution of 6km. The preliminary spatial downscaling network model includes a style multi-factor cross-attention module, a global cross-attention module, a residual channel spatial attention module, and a content-guided attention fusion module. Performing an initial spatial downscaling process includes at least the following steps: The spatial and channel weights of the input data are calculated using a style multi-factor cross-attention module and a global cross-attention module. In this process, the initial auxiliary data is processed by both holistic and stepwise input methods to obtain high-precision feature maps and recalibrated feature maps, which are then regarded as calibrated preliminary feature maps. Deep features in high-dimensional data are extracted by residual channel spatial attention module to explore the intrinsic relationship between multiple factors and obtain enhanced feature maps. A content-guided attention fusion module is used to fuse the calibrated preliminary feature map with the enhanced feature map; The style multi-factor cross-attention module includes at least the following steps: Each input auxiliary data and low-resolution soil moisture data is convolved to extract features, resulting in corresponding auxiliary data feature maps and low-resolution soil moisture data feature maps. Through style modulation and spatial attention mechanisms, channel attention and spatial attention are applied to the auxiliary data feature maps and low-resolution soil moisture data feature maps, respectively. This process generates corresponding channel weights and spatial weights, thereby adjusting the high-resolution and low-resolution features and further improving the expressive power of the features. The feature maps generated from each auxiliary data and the feature maps generated from the low-resolution soil moisture data are concatenated by element-wise addition. The convolution module fuses element-wise and spatially weighted features, recalibrating the features of each channel to ultimately generate a comprehensive high-precision feature map. The refined spatial downscaling network model further adds a residual dense connection Res2Net module on the basis of the output of the initial spatial downscaling network model. The refined spatial downscaling network model is used for multi-scale feature learning. The application of the refined spatial downscaling network model includes at least the following steps: By utilizing the style multi-factor cross-attention module and the global cross-attention module to calculate spatial and channel weights, and by employing both holistic and stepwise input methods for the initial input auxiliary data, preliminary calibration of the feature map is achieved. A residual channel spatial attention module is adopted to obtain enhanced feature maps by fitting the intrinsic correlation between multiple bands. At the same time, a residual dense connection Res2Net module is designed to significantly improve network performance through multi-scale feature learning. A content-guided attention fusion module is used to fuse the pre-calibrated feature map with the enhanced feature map, thereby improving the model's ability to represent spatial details.

2. A gradual soil moisture downscaling method based on a hybrid attention mechanism, characterized by: At least the following steps are included: S1: Collect satellite remote sensing data from multiple data sources, including at least SMAP soil moisture data and auxiliary data, including CLDAS-V2.0 reanalysis data, MODIS surface reflectance data and SRTM digital elevation model; S2: Preprocess the satellite remote sensing data to obtain spatiotemporally continuous low-resolution soil moisture data; S3: Construct the PDHAM downscaling model as described in claim 1, input the preprocessed satellite remote sensing data into the PDHAM downscaling model, use the preliminary spatial downscaling network model to perform preliminary downscaling on the low-resolution soil moisture data, and use the fine spatial downscaling network model to perform multi-scale feature processing on the preliminary downscaling results to obtain high-resolution soil moisture data. S4: Train the PDHAM downscaling model based on the combined loss function. After the loss converges, use the trained PDHAM downscaling model to predict soil moisture data with a spatial resolution of 1 km. The combined loss function includes at least the Charbonnier loss function and the marginal loss function.

3. The gradual soil moisture downscaling method based on a hybrid attention mechanism according to claim 2, characterized in that: The preprocessing of the satellite remote sensing data includes at least the following steps: The obtained satellite remote sensing data is processed by stitching, cropping and reprojection, and outliers are removed by low-pass filtering. The data after removing outliers was then resampled to 0.06° and 0.36° and normalized. Meanwhile, a deep learning network for SMAP soil moisture data reconstruction was used to reconstruct missing SMAP soil moisture data, thereby obtaining spatiotemporally continuous low-resolution soil moisture data.

4. The gradual soil moisture downscaling method based on a hybrid attention mechanism according to claim 3, characterized in that: The global cross-attention module treats all auxiliary data as a whole and calibrates it with low-resolution soil moisture data. Convolution operations are applied to the two inputs, auxiliary data and low-resolution soil moisture data, to generate corresponding attention maps. The generated attention maps guide the interaction between low-resolution and high-resolution features. By combining the cross-attention mechanism, a recalibrated feature map is obtained.

5. The gradual soil moisture downscaling method based on a hybrid attention mechanism according to claim 4, characterized in that: The residual channel spatial attention module collaboratively obtains the high-precision feature map output by the style multi-factor cross attention module and the recalibrated feature map output by the global cross attention module. The residual channel spatial attention module uses an embedded attention mechanism to further extract and enhance deep feature information and output an enhanced feature map. The residual channel spatial attention module combines dense connections with an attention mechanism to effectively capture local detail information; The residual channel spatial attention module combines a channel attention layer and a spatial attention layer.

6. The gradual soil moisture downscaling method based on a hybrid attention mechanism according to claim 5, characterized in that: The application of the content-guided attention fusion module includes at least the following steps: The high-precision feature map output by the style multi-factor cross-attention module and the recalibrated feature image output by the global cross-attention module are used as low-level feature images, and the enhanced feature image output by the residual channel spatial attention module is used as high-level feature images. These are then input into the content-guided attention fusion module for initial fusion. At this point, channel attention, spatial attention, and pixel attention mechanisms are used to optimize the fusion process. By calculating the feature importance of the channel dimension and spatial dimension, the fusion ratio of low-level and high-level features is dynamically adjusted. Subsequently, 1x1 convolution is applied to process the fused features, resulting in the final fused feature map. This flexibly combines information from different levels, enhances feature representation, and improves the prediction accuracy in soil moisture data downscaling tasks.

7. The gradual soil moisture downscaling method based on a hybrid attention mechanism according to claim 6, characterized in that: The application of the residual dense connection Res2Net module includes at least the following steps: The input feature map is split into 3 groups according to the channel dimension. Each group is processed separately using convolutional layers with different kernel sizes. Then, the input feature information is preserved through residual connections. The processed multi-scale features are concatenated and integrated through 1x1 convolutions to finally generate a fused feature map.

Citation Information

Patent Citations

  • Total primary productivity high-precision downscaling method based on multi-scale attention mechanism

    CN120894687A

  • Soil moisture downscaling method based on self-attention mechanism

    CN121074670A