Soil moisture inversion construction method integrating deep learning and machine learning

By using multi-source data fusion and a collaborative inversion method combining deep learning and machine learning, the problems of single data source and insufficient model generalization in soil moisture estimation are solved, achieving high-precision and robust soil moisture inversion.

CN121638019APending Publication Date: 2026-03-10INST OF WATER RESOURCES FOR PASTERAL AREA MINIST OF WATER RESOURCES P R C

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing soil moisture estimation methods rely on a single data source, have insufficient model generalization ability, and fail to comprehensively consider the coupling effects of multiple factors, resulting in decreased inversion accuracy and insufficient practicality.

Method used

The method employs multi-source heterogeneous data acquisition, cross-modal attention fusion module and graph neural network layer for data fusion, and combines deep learning and machine learning dual-path collaborative inversion. By adjusting hyperparameters through Bayesian optimization and quantifying uncertainty, high-precision estimation of soil moisture is achieved.

Benefits of technology

It significantly improved the accuracy of soil moisture retrieval by 12–18 percentage points, and enhanced the estimation performance, especially in areas with complex land cover and high soil heterogeneity, where the model’s generalization ability and robustness were significantly enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121638019A_ABST
    Figure CN121638019A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning and machine learning fused soil moisture inversion construction method, and relates to the technical field of measurement of physical properties of materials, and the method comprises the steps: capturing complementary information and spatial context of multi-source data through a multi-source heterogeneous data space-time adaptive fusion step by using a cross-modal attention mechanism and a graph neural network; through a deep learning and machine learning dual-path collaborative inversion step, advantage complementation is realized by combining data-driven nonlinear modeling and a physical constraint interpretable model; according to the method, the defects of single data source, insufficient model generalization ability and incomplete physical mechanism consideration in the prior art are overcome, the inversion precision is improved by 12%-18% under the complex earth surface condition, and the method has the advantages that the method is suitable for large-scale popularization and application. And a high-precision, strong-generalization and reliable technical means is provided for precise monitoring of soil moisture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of materials technology by measuring the chemical or physical properties of materials, and specifically to a method for constructing soil moisture inversion by integrating deep learning and machine learning. Background Technology

[0002] Soil moisture is one of the key physical quantities in land-atmosphere interactions, significantly impacting the climate system and its changes. Accurate estimation of soil moisture is crucial for water resource management, agricultural production, climate prediction, and flood monitoring.

[0003] The Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, and others disclosed a soil moisture estimation method in their invention patent application (application number: 202310266595.4, publication number: CN115980317A, invention title: Soil Moisture Estimation Method Based on Ground-Based GNSS-R Data with Corrected Phase). This method extracts the amplitude and standard phase from the signal-to-noise ratio data of ground-based GNSS-R; then, it uses the extracted amplitude to invert and obtain the vegetation water content over a long period; next, based on the vegetation water content over the long period and combined with pre-acquired vegetation normalized index data, it calculates the phase shift caused by the vegetation canopy; finally, it corrects the standard phase based on the phase shift and, combined with measured soil moisture data, estimates the soil moisture data.

[0004] However, the aforementioned existing technologies have the following shortcomings:

[0005] First, regarding the limitation of data source reliance, this method depends solely on a single ground-based GNSS-R data source for soil moisture estimation. Given the widespread availability of multi-source remote sensing data, it fails to fully utilize the complementary information from various data sources such as optical remote sensing, synthetic aperture radar, and meteorological observations. In practical applications, in areas with complex land cover types, uneven vegetation distribution, and high spatial heterogeneity of soil texture, a single data source cannot comprehensively reflect the land surface condition, leading to a 15-20 percentage point decrease in soil moisture retrieval accuracy, thus failing to meet the needs of refined applications.

[0006] Second, regarding model generalization ability, this method uses empirical equations to establish the relationship between amplitude and vegetation water content, and fits the relationship between vegetation water content and phase shift using a quadratic polynomial. This empirical model, based on data from a specific region, has parameters with strong regional dependence. When applied across regions and time periods, the model parameters need to be recalibrated, limiting its generalization performance. When applied to new regions with significantly different climate conditions, soil types, and vegetation species compared to the training region, the model performance degrades significantly, resulting in insufficient practicality and scalability.

[0007] Third, regarding the physical mechanism, this method primarily addresses the attenuation effect of vegetation cover on GNSS-R signals by retrieving vegetation moisture content through amplitude inversion, and then calculating the phase shift caused by vegetation. However, the remote sensing signal response of soil moisture is a complex multi-factor coupling process. In addition to vegetation cover, factors such as soil texture, soil temperature, surface roughness, and soil organic matter content all significantly affect the characteristics of the reflected signal. This method fails to comprehensively consider the coupling effects of these factors. When these factors change significantly in space or time, the method based solely on vegetation correction cannot accurately compensate for the influence of other factors, leading to increased inversion errors.

[0008] Therefore, there is an urgent need to provide a high-precision soil moisture inversion method that can make full use of multi-source data, has strong generalization ability, and comprehensively considers the coupling effect of multiple factors, so as to solve the problems existing in the above-mentioned technologies. Summary of the Invention

[0009] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide a soil moisture inversion construction method that integrates deep learning and machine learning, thereby solving the technical problems of single data source, insufficient model generalization ability, and incomplete consideration of physical mechanisms in existing technologies.

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] A method for constructing soil moisture inversion by integrating deep learning and machine learning includes:

[0012] The steps for acquiring multi-source heterogeneous data include: acquiring GNSS-R reflection signals, multispectral remote sensing image data, synthetic aperture radar backscattering coefficient data, meteorological parameter data, and topographic data of the samples;

[0013] The spatiotemporal adaptive fusion step of multi-source heterogeneous data includes: based on the multi-source heterogeneous data, calculating the contribution weight of each data source through a cross-modal attention fusion module to obtain adaptive weighted fusion features; constructing spatially adjacent sampling points into a graph structure, performing graph convolution propagation through a graph neural network layer to capture spatial correlations, and obtaining a fusion feature representation containing spatial context information;

[0014] The deep learning and machine learning dual-path collaborative inversion steps include: inputting the fused feature representation into the deep learning branch and the machine learning branch respectively; the deep learning branch extracting high-dimensional nonlinear features through a convolutional-recurrent hybrid network; the machine learning branch establishing a piecewise linear model with physical constraints through gradient boosting decision tree integration; and dynamically adjusting the weights of the two branches according to the characteristics of the input data through an adaptive fusion module to obtain the collaborative inversion features.

[0015] The adaptive hyperparameter adjustment and uncertainty quantification steps of Bayesian optimization include: based on the Bayesian optimization framework, modeling the mapping relationship between hyperparameters and model performance through Gaussian processes, and intelligently searching for the optimal hyperparameter configuration; in the inference stage, performing multiple forward propagations through the Monte Carlo Dropout mechanism to obtain the predicted distribution; and based on the predicted distribution, calculating the estimated value of soil sample moisture and its confidence interval to achieve uncertainty quantification.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] This invention employs a multi-source heterogeneous data spatiotemporal adaptive fusion mechanism, integrating multiple data sources such as GNSS-R, multispectral remote sensing, synthetic aperture radar, meteorological observations, and topographic data, fully utilizing the complementary information from these different sources. The cross-modal attention fusion module adaptively learns the contribution weights of each data source to soil moisture estimation, avoiding information redundancy and suboptimal performance issues associated with simple splicing or fixed-weight fusion. The graph neural network layer constructs a spatial graph structure, capturing spatial correlations and neighborhood context information between sampling points, enabling the model to utilize data from surrounding areas to assist in estimating the current location. Compared to single-data-source methods, this invention improves soil moisture retrieval accuracy by 12–18 percentage points in areas with complex land cover and high soil heterogeneity, significantly enhancing estimation performance. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall process of a soil moisture inversion construction method that integrates deep learning and machine learning according to the present invention.

[0019] Figure 2 This is a detailed flowchart illustrating the spatiotemporal adaptive fusion steps of multi-source heterogeneous data in this invention.

[0020] Figure 3 This is a schematic diagram of the architecture of the dual-path collaborative inversion steps of deep learning and machine learning in this invention;

[0021] Figure 4 This is a flowchart illustrating the adaptive hyperparameter adjustment and uncertainty quantification steps of Bayesian optimization in this invention. Detailed Implementation

[0022] like Figure 1As shown, this invention provides a soil moisture inversion construction method integrating deep learning and machine learning. Starting with data acquisition, the method proceeds through multi-source data fusion and dual-path collaborative inversion, ultimately achieving high-precision estimation and uncertainty quantification of soil moisture. The entire process embodies the design philosophy of combining data-driven approaches with physical constraints, complementing the advantages of deep learning and machine learning, and emphasizing both automated optimization and credibility assessment. In specific implementation, firstly, multiple data sources such as GNSS-R, remote sensing, and meteorological data are collected through a multi-source heterogeneous data acquisition step; then, these data are intelligently fused through a spatiotemporal adaptive fusion step, extracting feature representations containing spatial context; next, an inversion model is constructed through a dual-path collaborative inversion step combining deep learning and machine learning, fully utilizing the advantages of both methods; finally, automatic model optimization and credibility assessment of the results are achieved through Bayesian optimization adaptive hyperparameter adjustment and uncertainty quantification steps. Each step is closely connected through standardized data interfaces and feature transfer mechanisms, forming a complete end-to-end processing chain.

[0023] The present invention provides a method for constructing soil moisture inversion by integrating deep learning and machine learning, which includes the following four main steps, each of which is interconnected and together constitutes a complete technical solution.

[0024] In a specific embodiment of the present invention, the multi-source heterogeneous data acquisition step is responsible for collecting all the input data required for soil moisture retrieval. The data acquired in this step includes, but is not limited to, the following categories:

[0025] First, GNSS-R reflection signal characteristic data is acquired. Specifically, ground-based or spaceborne GNSS-R receivers are used to receive reflection signals from navigation satellites, and features such as signal-to-noise ratio time series, phase information, amplitude information, and delay-Doppler maps are extracted. Preferably, data products from the CYGNSS constellation or ground-based GNSS-R observation network are used. These data products provide pre-processed and calibrated reflection signal parameters that can reflect changes in the dielectric properties of the Earth's surface. The spatial resolution of GNSS-R data is typically in the range of 3-25 kilometers, and the temporal resolution can reach hourly levels or higher frequencies, which can meet the needs of dynamic soil moisture monitoring.

[0026] Secondly, multispectral remote sensing imagery data is acquired. Preferably, multispectral imagery from Sentinel-2 or Landsat series satellites is used to acquire surface reflectance data in the visible light band (red, green, and blue) and near-infrared bands. Based on this data, vegetation parameters such as the Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EVI), as well as water-related indices such as the Normalized Difference Water Index (NDWI), can be calculated. These optical remote sensing parameters reflect surface vegetation cover and vegetation water content, providing important auxiliary information for soil moisture estimation. Multispectral remote sensing data typically has a spatial resolution of 10–30 meters and a temporal resolution of 5–16 days, providing high spatial detail and relatively frequent observations.

[0027] Next, synthetic aperture radar (SAR) backscattering coefficient data is acquired. Preferably, Sentinel-1C band SAR data is used to obtain the VV and VH dual-polarized backscattering coefficients. SAR backscattering coefficients are sensitive to surface roughness and dielectric constant, and soil moisture is the main factor affecting the dielectric constant; therefore, SAR data can directly reflect changes in soil moisture. Compared to optical remote sensing, SAR is not limited by cloud cover and lighting conditions, enabling all-weather, 24 / 7 observation and providing more continuous and stable data support. The spatial resolution of SAR data can reach 10 meters or even higher, and the temporal resolution is 6–12 days.

[0028] In addition, meteorological parameter data is acquired. Specifically, meteorological elements such as temperature, precipitation, evapotranspiration, relative humidity, and wind speed are obtained from meteorological observation stations, reanalysis datasets, or numerical weather prediction products. Preferably, ERA5 reanalysis data or local meteorological station observation data are used, as these data can reflect the driving factors and boundary conditions of soil moisture. Temperature affects soil evaporation and vegetation transpiration, precipitation is the main input to soil moisture, and evapotranspiration is the main output; these meteorological parameters play an important role in controlling the temporal evolution of soil moisture.

[0029] Finally, topographic data is acquired. Specifically, topographic parameters such as elevation, slope, and aspect are extracted using a Digital Elevation Model (DEM). Preferably, global DEM products such as SRTM or ASTERGDEM are used. Topography has a significant impact on the spatial distribution of soil moisture. Elevation affects temperature and precipitation distribution, while slope and aspect affect runoff and redistribution processes. These topographic parameters help the model understand the spatial heterogeneity of soil moisture.

[0030] During data acquisition, spatiotemporal registration of data from different sources is necessary to ensure they correspond to the same time window and spatial location. Preferably, all data is resampled to a uniform spatial grid (e.g., 0.01 degree or 1 km resolution), and observations from the same or similar dates are extracted to construct a dataset for model training and inference. Data preprocessing also includes steps such as missing value imputation, outlier detection and removal, and data standardization to ensure the quality and consistency of the input data.

[0031] like Figure 2 As shown, the spatiotemporal adaptive fusion step of multi-source heterogeneous data is the first core innovation of this invention. This step achieves intelligent fusion of multi-source data and effective capture of spatial context through a cross-modal attention fusion module and a graph neural network layer.

[0032] The cross-modal attention fusion module employs a self-attention mechanism to adaptively learn the contribution weights of different data sources to soil moisture estimation. Specifically, for the features of each data source, a linear transformation is first performed to obtain the query vector, key vector, and value vector. Let the... The features of each data modality are represented as follows: Its dimensions are ,in For the sample size, This represents the feature dimension of the modality. It is achieved through a learnable weight matrix. Calculate the query matrix respectively Key matrix Sum matrix :

[0033] ,

[0034] in: For the first A query matrix with 1 modality and 1 dimension. ; For the first A key matrix of modalities, with dimensions of ; For the first A value matrix of modalities, with dimensions . ; To query the projection matrix, the dimension is ; Let be the key projection matrix, with dimension . ; The value projection matrix has dimensions of . ; The dimensions for queries and keys are typically 64 or 128; The dimension of the value is usually related to same.

[0035] Then, the cross-modal attention weights are calculated. For any two modalities... and Calculate the similarity matrix between them:

[0036] ,

[0037] in: For modality and modality The similarity matrix between them has dimensions of . ; For modality The transpose of the key matrix; This is a scaling factor used to prevent the gradient from vanishing due to an excessively large dot product.

[0038] Next, the similarity matrix is ​​normalized using the softmax function to obtain the attention weights:

[0039] ,

[0040] in: For modality For modes The attention weight matrix, with dimensions of ; The total number of modalities; the softmax function normalizes along the modal dimension to ensure that the sum of the attention weights of all modalities is 1.

[0041] Finally, the value vectors are weighted and summed based on the attention weights to obtain the fused feature representation:

[0042] ,

[0043] in: For the first The feature representation of each modality after attention fusion, with dimensions of . This representation integrates information from all modalities, and the weights are obtained through adaptive learning by the model.

[0044] The fusion features of all modalities are concatenated or averaged to obtain the final cross-modal fusion features. :

[0045] ,

[0046] in: The final cross-modal fusion feature has the following dimensions: This feature contains complementary information from all data sources, and the contribution of each data source is adaptively adjusted based on its correlation with other data sources.

[0047] In soil moisture estimation, the information content and reliability of different data sources vary under different conditions. For example, in densely vegetated areas, the vegetation index from optical remote sensing provides more information, while on bare surfaces, the SAR backscattering coefficient is more reliable. Cross-modal attention mechanisms can automatically learn these conditionally dependent weight relationships, avoiding the subjectivity and limitations of manually designed fusion rules.

[0048] The core idea of ​​the attention mechanism is to model the correlation between features through query-key-value triples. The query vector can be understood as what information I need, and the key vector as what information I can provide. Their dot product measures the degree of matching between information demand and supply. Softmax normalization transforms similarity into a probability distribution of weights, ensuring numerical stability and interpretability. The value vector represents the actual information content being conveyed; weighted summation achieves selective aggregation of information. (Scaling factor) The introduction of is to control the variance of the dot product, when the dimension When the value is large, the dot product becomes larger, causing the softmax function to enter the saturation region, and the gradient approaches zero. Divide by This can keep the variance of the dot product within a reasonable range and avoid the gradient vanishing problem.

[0049] In the soil moisture estimation scenario of this invention, it is assumed that the input includes five data modes: GNSS-R features (signal-to-noise ratio, phase, amplitude, dimension 10), multispectral indices (NDVI, EVI, NDWI, dimension 3), SAR backscattering (VV, VH polarization, dimension 2), meteorological parameters (temperature, precipitation, evapotranspiration, dimension 3), and topographic parameters (elevation, slope, aspect, dimension 3). After linear projection, the features of each mode are mapped to a unified 64-dimensional space (…). Attention weight matrix This reflects the dependencies between different modes. For example, in vegetated areas, multispectral modes may have higher attention weights to GNSS-R modes because vegetation information helps explain variations in the GNSS-R signal. Through adaptive fusion, the model can flexibly adjust the contributions of each mode under different conditions, thereby improving inversion accuracy.

[0050] Compared to simple feature concatenation or fixed-weight averaging, cross-modal attention fusion mechanisms offer several significant advantages. First, adaptability: weights are learned through data-driven processing, automatically adjusting to different surface conditions and data quality. Second, interpretability: the attention weight matrix is ​​visualized, intuitively showing the correlation and contribution between different data sources, enhancing model transparency. Third, robustness: when a data source is noisy or missing, the attention mechanism can reduce its weights, relying on other reliable data sources, thus improving the model's fault tolerance. Fourth, computational efficiency: although matrix multiplication is involved, the computational overhead is relatively controllable due to the use of efficient parallel computing libraries (such as PyTorch and TensorFlow), allowing for real-time processing on GPUs. Experimental results show that the average absolute error of soil moisture estimation is reduced by approximately 0.03-0.05% compared to simple concatenation after adopting cross-modal attention fusion. The correlation coefficient increased by approximately 0.08-0.12.

[0051] After completing cross-modal feature fusion, this invention further captures spatial correlations through graph neural network layers. Soil moisture exhibits significant spatial continuity and correlation; soil moisture in adjacent areas is often influenced by similar climatic and topographical conditions, demonstrating a certain spatial dependence. Graph neural networks can explicitly model this spatial structure, utilizing neighborhood information to assist in estimating the current location.

[0052] First, a spatial graph structure is constructed. Each sampling point is treated as a node in the graph, and the node features are the cross-modal fusion features of that point. The connections between nodes are established using an adjacency matrix. The adjacency matrix is ​​constructed based on geographical distance and environmental similarity. Specifically, for nodes... and nodes If the geographical distance between them is less than a threshold If the sampling points are not adjacent, they are considered neighbors and their corresponding positions in the adjacency matrix are set to 1; otherwise, they are set to 0. Preferably, an adaptive distance threshold is used, which is dynamically adjusted according to the density of sampling points. A smaller threshold is used in densely sampled areas, and a larger threshold is used in sparsely sampled areas to ensure that each node has a reasonable number of neighbors.

[0053] Adjacency Matrix elements Further weighting can be applied, with weights calculated based on environmental similarity:

[0054] ,

[0055] in: For nodes and nodes Edge weights between them; For nodes and nodes The geographical distance between them, in kilometers; This is a distance scale parameter that controls the rate of distance decay, typically ranging from 5 to 20 kilometers. and For nodes and nodes The environmental feature vector includes static attributes such as elevation, slope, and soil type; The Euclidean distance is the environmental feature vector. For environmental similarity scale parameters; This is the distance threshold, typically set to 20-50 kilometers.

[0056] This formula combines geographical distance and environmental similarity; nodes that are closer together and have more similar environments have larger edge weights, indicating a stronger correlation between them. An exponential function is used for smooth decay, avoiding the discontinuities caused by a hard threshold.

[0057] Then, neighborhood information is aggregated through graph convolution operations. The core operation of the Graph Convolutional Network (GCN) is to weight and aggregate the features of each node with the features of its neighboring nodes. Specifically, the... The formula for calculating layer graph convolution is:

[0058] ,

[0059] in: For the first The node feature matrix of the layer has a dimension of ,in For the number of nodes, For the first The feature dimensions of the layer; For the first The node feature matrix of the layer has a dimension of ; To add the adjacency matrix after adding self-loops, For an identity matrix, a self-loop indicates that the node itself also participates in the aggregation; for The degree matrix, diagonal elements ; For the first The learnable weight matrix of the layer has dimensions of ; The ReLU function is typically used as the activation function.

[0060] The physical meaning of this formula is: First, node characteristics Through the weight matrix Perform a linear transformation; then, through the adjacency matrix... Weighted aggregation is performed, and the new feature of each node is the weighted sum of its own feature and the features of its neighbors; finally, the degree matrix is ​​used to... Normalization is performed to ensure that nodes of different degrees are comparable during aggregation; the activation function introduces nonlinearity to enhance the expressive power of the model.

[0061] Through multiple layers of graph convolution (preferably 2-4 layers), the feature representation of each node gradually incorporates a wider range of neighborhood information. The first layer of graph convolution aggregates information from one-hop neighbors, the second layer aggregates information from two-hop neighbors, and so on. By stacking multiple layers, a larger range of spatial context can be captured. Finally, the graph neural network outputs node representations that contain global spatial context. , as input for subsequent branches of deep learning and machine learning.

[0062] The spatial distribution of soil moisture is influenced by a variety of factors, including precipitation distribution, topographic relief, and soil texture. These factors exhibit spatial continuity and correlation, leading to spatial autocorrelation in soil moisture. Traditional point-based estimation methods ignore this spatial dependence, making predictions based solely on single-point observation data, which is susceptible to local noise. Graph neural networks, by explicitly modeling the spatial graph structure, can utilize neighborhood information to smooth noise, improving the stability and accuracy of estimations.

[0063] The mathematical foundation of graph convolutional networks stems from spectral graph theory and signal processing. In the spectral domain, graph convolution can be understood as the eigenvector expansion and filtering operation of the graph Laplacian matrix. In the spatial domain, graph convolution performs message passing and aggregation through neighborhood relations defined by an adjacency matrix. Degree normalization. This ensures the symmetry and numerical stability of message passing, resulting in reasonable weights for nodes of different degrees during aggregation. Multi-level stacking allows information to propagate across the graph in multiple hops, enabling long-distance dependency modeling.

[0064] In the soil moisture estimation scenario of this invention, it is assumed that there are 1000 sampling points in the target area, and the cross-modal fusion feature dimension of each point is 64. When constructing the adjacency matrix, a distance threshold of 30 kilometers is set, and each node connects to an average of about 8 to 12 neighbors. A two-layer graph convolution is used, with the output dimension of the first layer being 128 and the output dimension of the second layer being 64, consistent with the input dimension. After graph convolution propagation, the features of each node not only include its own multi-source data information, but also incorporate the average information of the surrounding 8 to 12 neighboring points, as well as broader contextual information indirectly obtained through two-hop propagation. This introduction of spatial context enables the model to utilize the consistency and continuity of the neighborhood to reduce the estimation error of isolated points. For example, if the GNSS-R observation of a certain point is abnormally disturbed, but the observations of its neighboring points are normal, the graph neural network can use the neighborhood smoothing mechanism to pull the estimation result of that point to the consistency level of its neighbors, thereby reducing the impact of noise.

[0065] The graph neural network layer brings the following significant technical advantages to this invention. First, spatial consistency: through neighborhood aggregation, the soil moisture estimation results output by the model are more spatially smooth and consistent, avoiding unreasonable abrupt changes and isolated outliers, and conforming to the physical continuity of soil moisture. Second, noise robustness: Single-point observations may be affected by measurement errors, instrument malfunctions, or local interference. The graph neural network, through the integration of neighborhood information, can smooth these local noises and improve the stability of the estimation. Third, sample efficiency: in regions with sparse training samples, single-point models may perform poorly due to a lack of sufficient data. The graph neural network can utilize information from neighborhood samples for assistance, indirectly increasing the number of effective training samples. Fourth, scalability: The architecture of the graph neural network has good scalability, capable of handling large-scale graph structures from hundreds to millions of nodes, suitable for soil moisture estimation applications from regional to global scales. Experimental verification shows that after introducing the graph neural network layer, the spatial continuity index of soil moisture estimation (such as Moran's I spatial autocorrelation coefficient) is improved by approximately 0.15-0.20, and the number of outliers is reduced by approximately 30%-40%.

[0066] like Figure 3 As shown, the dual-path collaborative inversion step of deep learning and machine learning is the second core innovation of this invention. This step fully leverages the advantages of both methods through parallel processing and adaptive fusion of the deep learning branch and the machine learning branch, and constructs a high-performance and robust inversion model.

[0067] The deep learning branch employs a convolutional-recurrent hybrid network to extract complex spatial-temporal feature patterns of soil moisture. This network first processes spatial features through convolutional layers, then models temporal dependencies through LSTM layers, and finally outputs a high-dimensional feature representation through fully connected layers.

[0068] The convolutional layer design takes into account the need for multi-scale feature extraction. Specifically, multiple convolutional kernels are used, with sizes of [missing information]. , and This is to capture patterns at different spatial scales. For the input fusion feature tensor... Assuming its dimension is ,in For time steps, and For spatial dimensions (such as the arrangement of sampling points on a two-dimensional grid). This represents the number of feature channels. The formula for calculating the convolution operation is:

[0069] ,

[0070] in: For the first The convolutional kernel at the _th ... Feature maps on each output channel; For the first The weights of each convolutional kernel, with dimension [ ]. ,in and The height and width of the convolution kernel, Number of output channels; Indicates the convolution operation; For bias terms; The activation function is typically ReLU or LeakyReLU.

[0071] Convolutional operations, through local receptive fields and weight sharing, can effectively extract local spatial patterns while maintaining a controllable number of parameters. The combination of multi-scale convolutional kernels enables the network to simultaneously capture fine-grained local features and coarse-grained global patterns, enhancing the richness of feature representation.

[0072] Following the convolutional layers, pooling layers are used to reduce the spatial resolution of the feature maps, thereby reducing computational cost and enhancing the translation invariance of the features. Preferably, max pooling or average pooling is used, with a pooling window size of [size missing]. The step size is 2, which halves the spatial dimension of the feature map.

[0073] Then, the spatial features extracted by the convolutional layers are flattened and fed into the LSTM layer for temporal modeling. The computation of LSTM involves the synergistic effect of forget gates, input gates, and output gates, which can selectively preserve and update hidden states, thereby capturing long-term dependencies. LSTM operates at time steps... The calculation formula is:

[0074] ,

[0075] ,

[0076] ,

[0077] ,

[0078] ,

[0079] ,

[0080] in: The activation value for the forgetting gate, with a dimension equal to the size of the hidden layer. ; The activation value of the input gate, with dimension . ; The activation value of the output gate, with dimension . ; For time step The input dimension is the input feature dimension; For time step The hidden state, with dimension . ; For time step The cell state, dimension is ; For candidate cell states, the dimension is ; For time step cellular state; For time step The hidden state; The weight matrix is ​​the input to the gate; The weight matrix from the hidden state to the gate; It is the bias vector; It is the sigmoid activation function. ; This indicates element-wise multiplication.

[0081] LSTM uses a forget gate to control the degree of forgetting of old information, an input gate to control the degree of receiving new information, and an output gate to control the output of the hidden state. Cell state. As long-term memory, it is passed between time steps, enabling the network to remember important historical information and forget unimportant details. Preferably, the hidden layer dimension of LSTM... Set it to 128 or 256 to balance expressive power and computational efficiency.

[0082] Finally, the hidden state of the last time step of the LSTM is... The data is fed into a fully connected layer, undergoes a nonlinear transformation, and outputs the feature representation of the deep learning branch. :

[0083] ,

[0084] in: The output features of the deep learning branch have a dimension of It is usually set to 64 or 128; Here is the weight matrix of the fully connected layer, with dimension 1. ; It is the bias vector; This is the activation function.

[0085] Soil moisture variation is a complex spatiotemporal process, dynamically influenced by multiple factors such as precipitation, evaporation, infiltration, and runoff. Convolutional neural networks (CNNs) excel at extracting spatial features, identifying patterns in soil moisture distribution across space, such as the impact of vegetation cover and topographic control. Recurrent neural networks (especially LSTMs) are adept at modeling temporal dependencies, capturing the dynamic patterns of soil moisture evolution over time, such as rapid increases after precipitation events and continuous decreases during sunny days. Combining convolutional and recurrent networks allows for the simultaneous utilization of information from both spatial and temporal dimensions, constructing a more comprehensive and accurate feature representation.

[0086] The essence of convolution is to extract spatial patterns by performing a locally weighted average of the input signal through a sliding window and weight sharing. The learnability of the convolution kernel allows the network to adaptively discover useful patterns in the data without the need for manually designed features. The core innovation of LSTM lies in solving the gradient vanishing and gradient exploding problems of traditional RNNs through gating mechanisms and cell states, enabling the network to learn long-term dependencies. The forget gate determines which past information should be discarded, the input gate determines which parts of the current input should be stored, and the output gate determines which parts of the cell state should be output to the hidden layer. This fine-grained information flow control makes LSTM perform exceptionally well in sequence modeling tasks.

[0087] In the soil moisture estimation scenario of this invention, the input fused feature tensor contains daily observation data from the past 30 days, with a spatial resolution of 0.01 degrees (approximately 1 kilometer). The convolutional layer first processes the spatial features of each day to extract local patterns, such as the soil moisture distribution pattern within a small area. Then, the LSTM layer processes the 30-day feature sequence to capture the evolution trend of soil moisture over time, such as the impact of precipitation events and the cumulative effect of evaporation. Finally, the fully connected layer maps the extracted spatiotemporal features to a unified representation space, providing high-quality input for subsequent inversion tasks. Experiments show that the convolutional-recurrent hybrid network reduces the root mean square error of soil moisture estimation by approximately 0.02-0.04 compared to a single convolutional network or a single LSTM network. The correlation coefficient increased by approximately 0.05-0.10, fully validating the effectiveness of spatiotemporal joint modeling.

[0088] The construction of a decision tree is a recursive splitting process. At each node, the algorithm selects a feature and a threshold to divide the data into two subsets that minimize the objective function of the split. For regression tasks, the objective function is typically the mean squared error (MSE). Specifically, the nodes... The formula for calculating the split gain is:

[0089] ,

[0090] in: The split gain represents the reduction in the objective function after the split. The first-order gradient statistic of the left child node. ,in The sample index set for the left child node. For the sample The first derivative of the loss function with respect to the predicted value; The first-order gradient statistic of the right child node. ; Let be the second-order gradient statistic of the left child node. ,in For the sample The second derivative of the loss function with respect to the predicted value; The second-order gradient statistic of the right child node; This is the L2 regularization parameter, used to control the complexity of leaf node weights and prevent overfitting; The minimum gain threshold for splitting is only reached when the gain is greater than 1. It only splits at that time.

[0091] This formula approximates the loss function through Taylor expansion and uses first- and second-order gradient information to quickly calculate the optimal split point and leaf node weights, thus ensuring both the efficiency and accuracy of tree construction.

[0092] The gradient boosting algorithm uses an additive model framework to iteratively add new trees, with each tree fitting the residual (or gradient) of the current model. The training objective of the trees is to minimize the following loss function:

[0093] ,

[0094] in: For the first The loss function of the wheel; For the sample The true label; For the front The cumulative predicted value of each tree; For the first Tree samples The predicted value; The loss function is the mean squared error, which is typically used for regression tasks. The regularization term typically includes the number of leaf nodes in the tree and the L2 norm of the leaf node weights.

[0095] Using the idea of ​​gradient descent, the first Each tree fits the negative gradient of the loss function with respect to the current predicted value, i.e., the residual. Iteration After rounds, the final model prediction is:

[0096] ,

[0097] in: This is the final predicted value; The total number of trees, usually between 100 and 500; each tree Contribution through learning rate Scaling, i.e. Learning rate Typical values ​​for the learning rate range from 0.01 to 0.3. A smaller learning rate combined with a larger number of trees can improve the model's generalization ability.

[0098] In this invention, the machine learning branch introduces physical constraints; specifically, it uses the physical upper and lower bounds of soil moisture as hard constraints on the predicted values. (Soil moisture volume content) Physically limited by soil porosity, its value range is: ,in This represents the saturated water content, typically between 0.3 and 0.6. The range depends on soil texture. After model prediction, a cutoff function is used to ensure the predicted values ​​are within a reasonable range:

[0099] ,

[0100] in: These are the predicted values ​​after applying physical constraints; and The function is used to truncate the predicted value to ensure that it is neither lower than 0 nor higher than the saturated water content.

[0101] Furthermore, monotonicity constraints can be introduced during training to ensure that the predicted soil moisture values ​​maintain a monotonically increasing relationship with certain physical variables (such as precipitation). Frameworks such as LightGBM support imposing monotonicity constraints during tree construction, ensuring that the predicted values ​​change monotonically with the feature by restricting the direction of feature splitting.

[0102] In the soil moisture estimation scenario of this invention, the input features of the machine learning branch include the fused features of the graph neural network output, as well as some additional physical features, such as the cumulative precipitation, average temperature, and cumulative evapotranspiration over the past 7 days. These physical features have a clear causal relationship with soil moisture; for example, increased precipitation leads to increased soil moisture, and increased evapotranspiration leads to decreased soil moisture. Gradient boosting decision trees, through feature selection and splitting, can automatically discover these causal relationships and construct a piecewise linear prediction model. For example, the tree model might learn that when the cumulative precipitation over the past 7 days is greater than 50 mm, the estimated soil moisture value should be higher than 0.25. When the cumulative evapotranspiration exceeds 30 mm, the estimated soil moisture content should be below 0.20. These rules align with the physical processes of soil hydrology, enhancing the model's reliability. Preferably, the number of trees is set to 200, the maximum depth to 6, the learning rate to 0.1, and the L2 regularization parameter to 0.5. Experiments show that the machine learning branch exhibits stable performance and minimal error fluctuations even with sparse samples or simple features, demonstrating its robustness.

[0103] Deep learning and machine learning each have their advantages. Deep learning excels at complex pattern recognition and big data fitting, while machine learning excels at interpretability and few-shot learning. This invention designs an adaptive fusion module that dynamically adjusts the weights of the two branches based on the characteristics of the input data, achieving complementary advantages.

[0104] The core idea of ​​adaptive fusion is to evaluate the complexity and sample density of the input data and adjust the fusion weights accordingly. Specifically, the sample density of the input data is first calculated, defined as the number of samples per unit space:

[0105] ,

[0106] in: This represents the sample density, expressed in samples per square kilometer. The number of samples within the target area; The area of ​​the target region is expressed in square kilometers.

[0107] High sample density means sufficient training data, allowing the data-driven advantages of deep learning to be fully utilized; low sample density means sparse data, making the generalization ability of machine learning and the utilization of physical constraints more important.

[0108] Then, the complexity of the input features is calculated, defined as the degree of non-linear correlation and interaction between features. A simple metric is to calculate the average mutual information between features:

[0109] ,

[0110] in: For feature complexity; The total number of features; Features and The mutual information between them reflects their statistical dependence.

[0111] High feature complexity means that there are complex nonlinear relationships between features, and deep learning has a greater advantage in its strong nonlinear modeling capabilities; low feature complexity means that the feature relationships are relatively simple, and piecewise linear models of machine learning can handle the task.

[0112] Based on sample density and feature complexity, calculate the fusion weights of the deep learning branch and the machine learning branch:

[0113] ,

[0114] ,

[0115] in: These are the weights for the deep learning branches, with values ​​ranging from 0 to 1. Weights for the machine learning branch; For the sigmoid function, This maps linear combinations to the interval between 0 and 1. The parameters are learnable and optimized using performance on the validation set.

[0116] When the sample density is high and the feature complexity is large When the value approaches 1, the model primarily relies on deep learning branches; when the sample density is low or the features are simple, Approaching zero, the model primarily relies on machine learning branches. This dynamic adjustment mechanism enables the model to adaptively select the most suitable path under different conditions, improving overall robustness and generalization ability.

[0117] Finally, the collaborative inversion features are obtained through weighted summation:

[0118] ,

[0119] in: The synergistic inversion features after fusion; This is the output of the deep learning branch; For the output of the machine learning branch, the predicted values ​​of the tree model are mapped to the same feature space as the deep learning branch through a fully connected layer.

[0120] In practical applications, the availability and characteristics of data are often dynamic. In areas with abundant data and complex surface conditions, deep learning can fully extract patterns from the data and leverage its advantages; however, in areas with sparse data and simple features, deep learning may face the risk of overfitting, while simpler machine learning models may be more stable. Traditional fusion methods typically use fixed weights or simple averaging, which cannot be flexibly adjusted according to data characteristics. The adaptive fusion mechanism of this invention introduces quantitative indicators of data characteristics (sample density and feature complexity) to achieve dynamic adjustment of weights, enabling the model to maintain optimal performance under different conditions.

[0121] The mathematical foundation of adaptive fusion lies in conditional probability models and ensemble learning theory. Conditional probability models select different models or weights based on varying input conditions to achieve conditionally dependent predictions. Ensemble learning improves overall performance through the diversity and complementarity of multiple base learners. Adaptive fusion can be viewed as a dynamically weighted ensemble, where weights are not fixed but dynamically calculated based on the characteristics of the input data. The introduction of the Sigmoid function ensures that the weights are between 0 and 1 with a smooth transition, avoiding the discontinuities caused by hard switching. Learnable parameters... Optimization is achieved through end-to-end training, ensuring that the weight calculation strategy aligns with the final task objective, thus enabling automatic learning of the fusion strategy.

[0122] In the soil moisture estimation scenario of this invention, assuming a densely populated plain area with high sampling point density (approximately 5 points per square kilometer) and high feature complexity (multiple crop types, complex irrigation patterns), the adaptive fusion module calculates... , The model primarily relies on deep learning, leveraging its powerful fitting capabilities to capture complex nonlinear relationships. However, in a sparsely observed mountainous area with low sampling point density (approximately 0.5 points per square kilometer) and relatively simple features (mainly controlled by terrain and climate), the adaptive fusion module calculates... , The model primarily relies on machine learning, leveraging its adherence to physical constraints and robustness to small samples to avoid overfitting inherent in deep learning. Experiments show that adaptive fusion, compared to fixed-weight fusion, improves average performance by approximately 5%–8% under different data conditions, and significantly reduces performance fluctuations, demonstrating the effectiveness of the adaptive mechanism.

[0123] In addition to the primary soil moisture estimation task, this invention introduces a multi-task learning framework to simultaneously optimize two auxiliary tasks: confidence prediction and anomaly detection. Multi-task learning, through feature sharing and joint training among tasks, improves the model's generalization ability and overall performance.

[0124] The goal of the confidence prediction task is to estimate the reliability of soil moisture estimates. Specifically, in the output layer of the model, in addition to outputting soil moisture estimates... It also outputs a confidence score. The confidence level indicates the degree of confidence in the estimated value. Training labels for confidence scores are defined through cross-validation or expert knowledge; for example, a label of 1 is assigned when the input data is complete and of high quality, and a label of 0 is assigned when the input data is significantly missing or contains anomalies. The loss function for the confidence prediction task uses binary cross-entropy:

[0125] ,

[0126] in: Predict the loss of the task based on confidence level; For the sample True confidence labels; This represents the confidence score for the model's predictions.

[0127] The goal of an anomaly detection task is to identify anomalous samples in the input data or prediction results. Specifically, during training, some known anomalous samples (such as abnormal observations caused by sensor malfunctions) are labeled, and a binary classifier is trained to determine whether a sample is anomalous. The loss function for the anomaly detection task also uses binary cross-entropy:

[0128] ,

[0129] in: Losses due to anomaly detection tasks; For the sample Abnormal tags, Indicates an anomaly. This indicates that it is normal; This represents the probability of anomalies predicted by the model.

[0130] The total loss function for multi-task learning is a weighted sum of the losses from the three tasks:

[0131] ,

[0132] in: This is the total loss function; The loss from the primary task (soil moisture estimation) is typically expressed as the mean squared error (MSE). and The task weight is used to balance the importance of different tasks, and preferably, the value is between 0.1 and 0.5.

[0133] Through joint training, the model's underlying feature extraction module is optimized by all three tasks. The learned feature representations are beneficial not only to the main task but also to the auxiliary tasks, thereby improving the generalization ability of the features. The introduction of the auxiliary tasks also serves as a regularization mechanism, preventing the model from overfitting on the main task.

[0134] In soil moisture estimation applications, besides focusing on the accuracy of the estimates, it is also necessary to pay attention to the reliability of the results and the identification of anomalies. Confidence prediction can provide users with an assessment of their confidence in the estimation results, assisting in decision-making; anomaly detection can promptly identify anomalies in the data or model, preventing the adoption of erroneous estimation results. Multi-task learning, through collaboration between tasks, can improve the performance of both the main task and auxiliary tasks, achieving multiple benefits in one go.

[0135] The theoretical foundation of multi-task learning lies in transfer learning and regularization theory. Multi-task learning assumes that different tasks share a common underlying structure or features. Through joint training, the model can utilize information from auxiliary tasks to improve the performance of the main task. From an optimization perspective, multi-task learning introduces additional constraints and objectives, subjecting the model's solution space to multiple limitations, thereby avoiding overfitting on a single task. Task weights and The choice of task weights affects the balance between tasks; a larger weight means the auxiliary task has a greater impact on the model, while a smaller weight means the auxiliary task acts only as a slight regularization term. In practice, task weights can be adjusted using grid search or Bayesian optimization to achieve optimal overall performance.

[0136] In the soil moisture estimation scenario of this invention, the introduction of a multi-task learning framework brings significant performance improvements. The confidence prediction task enables the model to identify samples with poor data quality or high model uncertainty, assigning these samples lower confidence scores to remind users to use them with caution. The anomaly detection task enables the model to identify abnormal observations caused by sensor malfunctions, data transmission errors, or extreme weather events, preventing these anomalous data from contaminating the model's training or inference. Experiments show that after introducing multi-task learning, the root mean square error of the main task (soil moisture estimation) is reduced by approximately 0.01-0.02. Meanwhile, the confidence prediction accuracy reached over 85%, and the F1 score for anomaly detection reached 0.78, fully validating the effectiveness of multi-task learning.

[0137] like Figure 4As shown, the adaptive hyperparameter adjustment and uncertainty quantification step of Bayesian optimization is the third core innovation of this invention. This step realizes the automated intelligent search of model hyperparameters through Bayesian optimization and realizes the uncertainty quantification of prediction results through Monte Carlo Dropout, providing the ability to automatically optimize and assess the credibility of soil moisture estimation.

[0138] Hyperparameter tuning is a crucial step in machine learning model development, directly impacting model performance. Traditional hyperparameter tuning methods, such as grid search and random search, are computationally expensive and inefficient, often failing to find optimal configurations in high-dimensional hyperparameter spaces. Bayesian optimization, by establishing a probabilistic model between hyperparameters and model performance, uses previous evaluation results to guide subsequent searches, enabling the finding of near-optimal hyperparameter configurations with fewer evaluation iterations.

[0139] The core of Bayesian optimization is Gaussian process (GP) modeling and the acquisition function (ACF). A Gaussian process is a nonparametric probabilistic model used to model unknown functions. The prior distribution. Assume the hyperparameters are configured as follows. The model's performance on the validation set is The Gaussian process assumes that for any finite number of hyperparameter configurations Corresponding performance Follows a multivariate normal distribution:

[0140] ,

[0141] in: The observed performance vector; It is a mean function, usually set to zero; The covariance matrix, also known as the kernel matrix, consists of elements... By kernel function definition.

[0142] Kernel functions reflect the similarity between different hyperparameter configurations. Commonly used kernel functions include the Radial Basis Function (RBF) kernel and the Matérn kernel. The expression for the RBF kernel is:

[0143] ,

[0144] in: Configuration of hyperparameters and Covariance between them; The signal variance controls the overall variation of the function value. The length scale parameter controls the smoothness of the function; a smaller value indicates a smoother performance. This indicates that the function changes rapidly and is relatively large. This indicates that the function changes slowly; The Euclidean distance between the two hyperparameter configurations.

[0145] The RBF kernel assumption assumes that similar hyperparameter configurations will produce similar performance, with closer proximity resulting in larger covariance and stronger correlation. This smoothness assumption allows Gaussian processes to predict performance in unknown regions using known evaluation results.

[0146] After evaluation Configure hyperparameters and observe performance. In the case of a new hyperparameter configuration Gaussian processes can predict the posterior distribution of their performance:

[0147] ,

[0148] in: For new configuration Performance; The posterior mean is used to provide a point estimate of the performance; The posterior variance reflects the uncertainty of the prediction.

[0149] The formulas for calculating the posterior mean and variance are:

[0150] ,

[0151] ,

[0152] in: This is the covariance vector between the new configuration and the previously evaluated configuration. , dimension ; The kernel matrix between the evaluated configurations has dimensions of . ; The inverse of the kernel matrix; The covariance between the new configuration and itself is equal to the signal variance. .

[0153] In Gaussian process-based predictions, the acquisition function is used to determine the next hyperparameter configuration to be evaluated. The acquisition function balances exploration and exploitation: exploration means selecting regions with high uncertainty to discover potentially better configurations; exploitation means selecting regions with high predictive performance to quickly find good configurations. Commonly used acquisition functions include expected improvement (EI), probabilistic improvement (PI), and confidence upper bound (UCB).

[0154] The definition of expected improvement in EI is:

[0155] ,

[0156] in: For configuration Expected improvements; This represents the best performance observed so far. To improve the quantity, when Superior It is positive when it is positive and zero otherwise; the expectation is the integral of the posterior distribution.

[0157] Under the Gaussian distribution assumption, the expected improvement has a closed-form solution:

[0158] ,

[0159] in: This refers to the improvement amount after standardization; The cumulative distribution function of the standard normal distribution; It is the probability density function of the standard normal distribution.

[0160] The goal is to improve the function to achieve a larger value in regions with high prediction mean (utilization) or high prediction uncertainty (exploration), thus balancing exploration and utilization. The iterative process of Bayesian optimization is as follows: First, several initial configurations are randomly selected in the hyperparameter space for evaluation; then, a Gaussian process model is fitted based on the evaluated configurations; next, the next configuration to be evaluated is selected by maximizing the acquisition function; the performance of this configuration is evaluated, and the Gaussian process model is updated; the above steps are repeated until a predetermined number of evaluations is reached or the performance improvement is no longer significant.

[0161] The soil moisture inversion model of this invention involves multiple hyperparameters, including the learning rate of the deep learning branch, the dimension of the LSTM hidden layer, the number of convolutional kernels, the Dropout ratio, the number of trees, the maximum depth, the learning rate of the machine learning branch, and the weight calculation parameters of the fusion module. The combination space of these hyperparameters is high-dimensional and complex, and different hyperparameters interact with each other. Traditional grid search requires traversing all possible combinations, and the computational cost increases exponentially with dimensionality, often making it infeasible. While random search is simple, it lacks intelligence and may require a large number of evaluations to find a good configuration. Bayesian optimization, by utilizing existing evaluation information, constructs a surrogate model between hyperparameters and performance, which can intelligently guide the search direction and find an approximately optimal hyperparameter configuration within a fewer number of evaluations, significantly reducing the time and computational cost of hyperparameter tuning.

[0162] The mathematical foundation of Gaussian processes lies in stochastic process theory and Bayesian inference. A Gaussian process can be viewed as an infinite-dimensional Gaussian distribution. By defining the covariance structure between different inputs through a kernel function, nonparametric modeling of the unknown function is achieved. The core of Bayesian inference is to update the posterior distribution using Bayes' theorem, utilizing the prior distribution and observed data. In a Gaussian process, the prior is a set of data related to the function... The Gaussian process hypothesis states that the posterior is based on the observed data. Then, the function at the new point The updated distribution of values ​​at each point. Due to the conjugate property of the Gaussian distribution, the posterior distribution is still Gaussian, and the mean and variance have closed-form solutions, making the Gaussian process computationally efficient and interpretable. The design of the acquisition function embodies the sequential decision-making concept in decision theory. By quantifying the value of each candidate configuration (expected improvement, probabilistic improvement, or confidence upper bound), the optimal decision is selected, achieving optimization of multi-step decision-making.

[0163] Specific application of this invention: In the soil moisture estimation scenario of this invention, the hyperparameter search space of Bayesian optimization is defined as follows: the learning rate of the deep learning branch is... The LSTM hidden layer dimension is chosen in the logarithmic space. Choose from an integer space, the number of convolution kernels is in Choose from an integer space, the Dropout ratio is in The number of branches in a continuous space is selected; the number of branches in a machine learning tree is... Choose from an integer space, with the maximum depth in The learning rate is chosen from an integer space. The model is selected from a continuous space. Gaussian process regression is used as a surrogate model, and the Matérn5 / 2 kernel is chosen as the kernel function, which achieves a good balance between smoothness and flexibility. The acquisition function is the expected improvement EI, as it performs well in balancing exploration and exploitation. The number of iterations for Bayesian optimization is set to 50. Each evaluation uses 5-fold cross-validation to calculate the model's average performance on the validation set (e.g., the negative value of the root mean square error, since Bayesian optimization defaults to maximizing the objective function). Experiments show that Bayesian optimization converges to a better hyperparameter configuration after about 30 iterations, which is comparable to or even better than 500 random searches, saving about 90% of computation time. The finally found optimal hyperparameter configuration reduces the model's root mean square error by about 0.015-0.025 compared to the default parameters. This fully verifies the effectiveness and efficiency of Bayesian optimization.

[0164] Bayesian optimization brings the following significant technical advantages to this invention. First, automation: The hyperparameter tuning process is fully automated, requiring no manual intervention or domain knowledge, thus lowering the technical threshold and time cost of model development. Second, efficiency: Through intelligent search strategies, it finds near-optimal configurations in fewer evaluation iterations, saving approximately 80%–90% of computational resources and time compared to traditional methods. Third, robustness: Bayesian optimization can handle noisy observations and the non-smoothness of the objective function, achieving stable convergence even with a degree of randomness during evaluation. Fourth, scalability: The Bayesian optimization framework can be easily extended to more hyperparameters or different model architectures, exhibiting good versatility. Furthermore, the intermediate results of Bayesian optimization (Gaussian process models and collected function values) can be visualized, helping developers understand the performance distribution and search trajectory in the hyperparameter space, enhancing the transparency and controllability of model development.

[0165] After the model training is completed and the optimal hyperparameters are determined, this invention further quantifies the uncertainty of the prediction results through the Monte Carlo Dropout mechanism. Uncertainty quantification can provide confidence intervals for soil moisture estimation results, helping users assess the reliability of the results, which is of great value in risk-sensitive application scenarios.

[0166] Dropout is a commonly used regularization technique in deep learning that randomly deactivates a portion of neurons during training to prevent overfitting. Monte Carlo Dropout applies the idea of ​​Dropout to the inference stage. By maintaining the activation of the Dropout layer during inference and performing multiple forward propagations, it obtains the statistical properties of the prediction distribution, thereby quantifying the cognitive uncertainty of the model.

[0167] Specifically, in the inference phase, for an input sample Maintain the Dropout layer with probability Randomly inactivated neurons, Each forward propagation yields a predicted value. ,in Because the number of neurons inactivated varies in each propagation, the predicted value... This will vary, forming a predicted distribution. The predicted mean will be used as the final soil moisture estimate.

[0168] ,

[0169] in: The predicted mean of Monte Carlo Dropout is used as an estimate of soil moisture. The number of forward propagations is usually between 50 and 200. The more propagations, the more accurate the prediction of the distribution, but the greater the computational cost.

[0170] Prediction variance reflects the uncertainty of the model:

[0171] ,

[0172] in: The variance is the prediction variance, which reflects the uncertainty of the model's prediction for the sample. The larger the variance, the more uncertain the model's prediction for the sample, which may be due to poor input data quality, incomplete features, or the sample being located outside the training data distribution.

[0173] Confidence intervals can be constructed based on the predicted mean and variance. Assuming the predicted distribution follows a normal distribution, the 95% confidence interval is:

[0174] ,

[0175] in: The 95th percentile of the standard normal distribution is given; the confidence interval provides the upper and lower bounds of the estimated soil moisture value, indicating that the true value has a 95% probability of falling within the interval.

[0176] In practical applications of soil moisture estimation, besides focusing on the accuracy of the estimates, reliability is also crucial. Because input data may contain missing data, noise, or anomalies, the model's predictions for some samples may be unreliable. Traditional point estimation models only provide a deterministic prediction, failing to inform users of the prediction's confidence level. Uncertainty quantification, by providing the prediction distribution and confidence intervals, allows users to understand the degree of uncertainty in the prediction, thus enabling more informed decision-making. For example, in drought early warning applications, if the model's estimate of soil moisture for a certain area has high uncertainty, decision-makers can collect more data or adopt a more conservative response strategy; conversely, if the uncertainty is low, they can more confidently formulate irrigation plans or issue early warnings based on the estimates.

[0177] The mathematical foundation of Monte Carlo Dropout lies in Bayesian deep learning and the Monte Carlo method. From a Bayesian perspective, Dropout can be interpreted as an approximate inference of the posterior distribution of network weights. During training, Dropout implicitly samples different subsets of weights, each corresponding to a different network configuration. During inference, Monte Carlo Dropout approximates the posterior distribution of weights by performing a Monte Carlo integral on the posterior distribution through multiple samplings of different Dropout configurations, thus obtaining the predicted posterior distribution. The Monte Carlo method is a numerical method that uses random sampling to estimate mathematical expectations or integrals, offering advantages such as simplicity, versatility, and parallelizability. The prediction variance of Monte Carlo Dropout primarily reflects the cognitive uncertainty of the model, i.e., the uncertainty of model parameters due to the limited training data. In addition, there is another type of uncertainty called stochastic uncertainty or data uncertainty, caused by noise in the data itself. A complete uncertainty quantification should consider both types of uncertainty simultaneously, but in practice, cognitive uncertainty is often more important and can be reduced by increasing the amount of data.

[0178] In the soil moisture estimation scenario of this invention, the Monte Carlo Dropout mechanism performs 100 forward propagations for each sample during the inference phase, with the Dropout ratio set to 0.2 (i.e., each neuron has a 20% probability of being inactivated). For a typical sample, it is assumed that the predicted values ​​obtained from 100 propagations are distributed between 0.22 and 0.28. Between these values, the predicted mean is 0.25. The prediction standard deviation is 0.015. Then the 95% confidence interval is This confidence interval indicates that there is a 95% probability that the actual soil moisture content is between 0.221 and 0.279. Between these parameters, users are provided with a confidence assessment of the estimation results. Experiments show a positive correlation between the prediction variance and the actual error in Monte Carlo Dropout; that is, samples with large prediction variance tend to have larger actual errors, validating the effectiveness of uncertainty quantification. By setting a variance threshold, samples with high uncertainty can be identified, allowing for additional validation or data supplementation, thereby improving the overall reliability of the estimation. Furthermore, uncertainty information can be used for active learning, prioritizing the labeling and training of samples with the greatest model uncertainty to maximize data utilization efficiency.

[0179] The Monte Carlo Dropout mechanism brings the following significant technical advantages to this invention. First, credibility assessment: By providing confidence intervals, the soil moisture estimation results include not only point estimates but also a credibility assessment of those estimates, enhancing the transparency and interpretability of the results. Second, risk management: In risk-sensitive applications such as drought early warning, irrigation decisions, and flood forecasting, uncertainty information can help decision-makers assess risks and formulate more robust strategies. Third, model diagnosis: By analyzing the spatial distribution and temporal variations of prediction variance, weak points in the model and deficient areas in the data can be identified, providing guidance for model improvement and data collection. Fourth, computational efficiency: Compared to other uncertainty quantification methods (such as fully variational inference in Bayesian neural networks and ensemble learning), Monte Carlo Dropout is simple to implement, has controllable computational overhead, and only requires multiple forward propagations during inference, without the need to retrain multiple models or perform complex posterior inference. Experiments show that Monte Carlo Dropout can achieve high-quality uncertainty estimation at the cost of increasing inference time by about 10-20 times (100 propagations compared to 1 propagation), which is acceptable in practical applications, especially for non-real-time batch processing scenarios.

[0180] In summary, this invention constructs a high-precision, highly generalizable, and reliable soil moisture inversion model through an innovative combination of multi-source fusion, dual-path collaboration, automatic optimization, and uncertainty quantification. It provides an effective technical means for the accurate monitoring and intelligent application of soil moisture, and has significant practical value and prospects for promotion.

[0181] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing soil moisture inversion by fusing deep learning and machine learning, characterized in that, The method comprises the following steps: a multi-source heterogeneous data acquisition step, comprising: acquiring GNSS-R reflected signals, multispectral remote sensing image data, synthetic aperture radar backscattering coefficient data, meteorological parameter data, and topographic data of the sample; a multi-source heterogeneous data spatio-temporal adaptive fusion step, comprising: based on the multi-source heterogeneous data, calculating the contribution weight of each data source through a cross-modal attention fusion module to obtain an adaptively weighted fusion feature; constructing spatially adjacent sampling points into a graph structure, performing graph convolution propagation through a graph neural network layer to capture spatial correlation, and obtaining a fusion feature representation containing spatial context information; a deep learning and machine learning dual-path collaborative inversion step, comprising: inputting the fusion feature representation into a deep learning branch and a machine learning branch respectively; the deep learning branch extracts high-dimensional nonlinear features through a convolution-recurrent hybrid network; the machine learning branch establishes a piecewise linear model with physical constraints through a gradient boosting decision tree ensemble; according to the characteristics of the input data, the weights of the two branches are dynamically adjusted through an adaptive fusion module to obtain a collaborative inversion feature; a Bayesian optimized adaptive hyperparameter adjustment and uncertainty quantification step, comprising: based on the Bayesian optimization framework, the mapping relationship between hyperparameters and model performance is modeled through a Gaussian process to intelligently search for the optimal hyperparameter configuration; in the inference stage, the prediction distribution is obtained through multiple forward propagations by the Monte Carlo Dropout mechanism; based on the prediction distribution, the soil sample moisture estimation value and its confidence interval are calculated to realize uncertainty quantification. 2.The method of claim 1, wherein, In the multi-source heterogeneous data spatio-temporal adaptive fusion step, the calculation of the cross-modal attention fusion module comprises: performing linear transformation on the features of each data source to obtain query vectors, key vectors, and value vectors; calculating the similarity between the query vectors and the key vectors, and normalizing the attention weights through a softmax function; based on the attention weights, the value vectors are weighted and summed to obtain the fused feature representation. 3.The method of claim 1, wherein, The construction of the graph neural network layer comprises: taking spatial sampling points as nodes of a graph, calculating the connection relationship between nodes based on geographical distance and environmental similarity, and constructing an adjacency matrix; aggregating the neighborhood feature information of each node through graph convolution operation; after multi-layer graph convolution propagation, a node representation containing global spatial context is obtained.

4. The method of claim 1, wherein the method is characterized by: The convolution-recurrent hybrid network used in the deep learning branch comprises: a convolution layer is used to extract spatial features, and multiple convolution kernels are used to capture spatial patterns of different scales; a long short-term memory network layer is used to capture temporal dependence and model the time evolution of soil moisture; a fully connected layer performs nonlinear transformation on the extracted features to output high-dimensional feature representation.

5. The method of claim 1, wherein the method is characterized by: The gradient boosting decision tree ensemble used in the machine learning branch comprises: training multiple decision trees based on historical observation data; introducing physical constraints, taking the upper and lower bounds of soil moisture as hard constraints for the predicted values; iteratively optimizing through a gradient boosting algorithm, each tree fitting the residual of the previous tree; integrating the prediction results of all decision trees to obtain the final output.

6. The method of claim 1, wherein the method is characterized by: The method for dynamically adjusting the weights of the two branches according to the characteristics of the input data comprises: Sample density and feature complexity of the input data are calculated; When the sample density is high and the feature complexity is large, the weight of the deep learning branch is increased; When the sample density is low or the feature complexity is simple, the weight of the machine learning branch is increased; The final collaborative inversion feature is obtained by weighted summation.

7. The method of claim 1, wherein the method is characterized by: The method for modeling the mapping relationship between hyperparameters and model performance by Gaussian process in the Bayesian optimization framework includes: Defining the hyperparameter search space, including learning rate, network layer number, hidden unit number, etc. Using Gaussian process to probabilistically model the functional relationship between hyperparameter configuration and validation set performance; Balancing exploration and utilization by collecting functions to select the next set of hyperparameters to be evaluated; Iteratively update the Gaussian process model until the optimal hyperparameter configuration is found.

8. The method of claim 1, wherein the method is characterized by: The method for obtaining the prediction distribution by multiple forward propagations of the Monte Carlo Dropout mechanism includes: Maintain the Dropout layer activation during the inference stage; Multiple forward propagations are performed on the same input sample, and different neurons are randomly inactivated in each propagation; Collect the prediction results of multiple propagations, and calculate the prediction mean as the soil moisture estimation value; Calculate the prediction variance and obtain the confidence interval based on the normal distribution assumption.

9. The method of claim 1, wherein the method is characterized by: The method also includes a multi-task learning mechanism that simultaneously optimizes soil moisture estimation, confidence prediction, and anomaly detection, and improves overall performance through feature sharing and joint training between tasks.

10. The method of claim 1, wherein the method is characterized by: The GNSS-R reflected signal feature data includes signal-to-noise ratio, phase, amplitude, delay-Doppler plot; the multispectral remote sensing image data includes reflectivity of visible and near-infrared bands; the meteorological parameter data includes temperature, precipitation, evapotranspiration; the topographic data includes elevation, slope, aspect.

Citation Information

Patent Citations

  • Ground-based GNSS-R data soil moisture estimation method based on corrected phase

    CN115980317A

  • Soil Moisture Estimation Method Based on Corrected Phase Ground-Based GNSS-R Data

    CN115980317B

Cited By

  • Vegetation canopy water content inversion method, device, equipment, product and medium

    CN122112538A

  • Vegetation canopy water content inversion method, device, equipment, product and medium

    CN122112538B

  • Fusion intervention targeting small sample marker enhancement method and system

    CN122117339A

  • Method for retrieving farmland soil moisture content based on multi-satellite dual-frequency GNSS signal intensity

    CN122194192A

  • Intelligent monitoring system for coastal wetland ecological parameters based on deep learning

    CN122198098A