Soil humidity multi-source data assimilation method based on dynamic interval analysis and EnKF
Through the multi-source data assimilation method of dynamic interval analysis and EnKF, the problem of insufficient weight adjustment and interval analysis in traditional soil moisture estimation is solved, and high-precision soil moisture data fusion is achieved, which enhances the robustness of data assimilation.
Patent Information
- Application Number
- CN202510336203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Traditional soil moisture estimation methods are difficult to dynamically adjust the weight of multi-source data, and lack effective interval analysis methods, resulting in inaccurate data fusion results.
The multi-source data assimilation method based on dynamic interval analysis and EnKF is adopted to optimize data weight allocation and covariance adjustment through data preprocessing, state prediction, dynamic interval analysis, dynamic weight adjustment and data assimilation update, and optimize the Kalman gain matrix by combining Newton's iterative method and improved covariance matrix.
It significantly improves the accuracy and robustness of data fusion, ensures that different data sources contribute reasonably during the fusion process, and improves the accuracy of soil moisture estimation.
Smart Images

Figure CN120256867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrological multi-source data processing methods, and specifically relates to a multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF. Background Art
[0002] Soil moisture is a key variable in hydrological, agricultural, and meteorological models, and the accuracy of its estimation is of great significance for drought monitoring, irrigation management, and climate prediction. However, there are many problems with traditional soil moisture estimation methods:
[0003] (1) It is difficult to dynamically adjust the weights of multi-source data to reflect changes in uncertainty. In practical applications, the reliability of soil moisture data provided by different data sources varies with factors such as time and space. Traditional methods cannot reasonably adjust the weights of each data source in the data fusion process according to the real-time changes in data source uncertainty, resulting in the data fusion result not being able to accurately reflect the true soil moisture situation.
[0004] (2) There is a lack of effective interval analysis methods, making it difficult to accurately characterize the error ranges of different data sources. Due to the lack of effective interval analysis means, it is impossible to accurately determine the error intervals of the data of each data source. This makes it impossible to fully consider the uncertainty of the data during data fusion, thus affecting the accuracy of soil moisture estimation.
[0005] The Ensemble Kalman Filter (EnKF) is an optimization method for non-linear and high-dimensional data assimilation. It approximates the state distribution of the system by introducing a set of "ensemble" states, thereby performing data fusion in complex systems. This makes it particularly suitable for applications in problems such as soil moisture that involve high dimensions and complex non-linear changes. EnKF overcomes the limitations of the traditional Kalman Filter (KF) in dealing with non-linear problems and high-dimensional data. It significantly improves the model's adaptability to observational data and computational efficiency by using multiple prediction ensembles as state estimates. In data processing, EnKF can effectively process and fuse data from multiple sources, especially in the presence of large uncertainties or noise. Summary of the Invention
[0006] To solve the problems of the traditional soil moisture estimation method, such as the difficulty in dynamically adjusting the weights of multi-source data and the lack of effective interval analysis methods, the present invention proposes a multi-source data assimilation method based on dynamic interval analysis and EnKF, aiming to optimize the data weight allocation and covariance adjustment mechanism to improve the accuracy of data fusion.
[0007] The present invention is implemented as follows: A multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF, the method comprising the following steps:
[0008] Step 1, data preprocessing: Obtain multi-source data, including measured soil moisture data and ERA-5 reanalysis data, and perform standardization processing on the obtained multi-source data to eliminate the dimensional differences between different data sources, so that the multi-source data is consistent in the time and space dimensions, thereby providing consistent data input for Step 2;
[0009] Step 2, state prediction: The distributed soil hydrological model generates an initial moisture state prediction set for the target area, including the initial moisture state and its uncertainty estimate, to capture the soil moisture changes within the target area;
[0010] Step 3, dynamic interval analysis: Use Newton iteration to calculate the upper and lower bounds of the uncertainty interval of each data source in the multi-source data described in Step 1, and combine the propagation mechanism of the simulation error to generate the error interval of each predicted value, thereby adjusting the interval range of the initial moisture state prediction set;
[0011] Step 4, dynamic weight adjustment: Optimize the weight allocation function according to the dynamic interval range, and use a non-linear function to dynamically adjust the weights of multi-source data fusion;
[0012] Step 5, data assimilation and update: Use the dynamically improved covariance matrix and weights to calculate the Kalman gain matrix and update the initial moisture state prediction set, generating the optimal moisture estimate after multi-source data fusion.
[0013] Furthermore, the standardization processing described in Step 1 includes time alignment, spatial resolution matching, and normalization processing. The formula for normalization processing is:
[0014]
[0015] In the formula, x i " is the measured soil moisture value after normalization processing, x i is the original measured soil moisture value, x min is the minimum moisture among all measured values, x max is the maximum moisture among all measured values.
[0016] Furthermore, the distributed soil hydrological model described in Step 2 is the GBHM geomorphology-based hydrological model.
[0017] Furthermore, the expression of the initial moisture state prediction set described in Step 2 is:
[0018] x t(1 = f(x t + u t ) + ∈ t (2)
[0019] In the formula, xt(1 is the predicted value of the next state of soil moisture, x t is the current soil moisture state, u t is the control input, f(x t +u t ) function is the evolution law of soil moisture from the current state to the next state, ∈ t is Gaussian white noise.
[0020] Furthermore, the Newton iteration formula described in step 3 is:
[0021]
[0022] In the formula, x n(1 is the approximation of the (n + 1)-th iteration, x n is the approximation of the n-th iteration, f(x n ) is the uncertainty of soil moisture, f"(x n ) is the derivative value of the function at x n ;
[0023] For the i-th data source, the upper and lower bounds of the interval are [L i , U i , and the calculation formula for the dynamic upper and lower bounds of the interval is:
[0024] L i = x - α i , U i = x - β i (4)
[0025] In the formula, x is the predicted value of the initial state of soil moisture, α i , β i are respectively the upper and lower boundaries of the error interval of this data source calculated by Newton iteration.
[0026] Furthermore, the dynamic adjustment of the weights for multi-source data fusion described in step 4 is: for data sources with smaller uncertainty intervals, higher weights are given; for data sources with larger uncertainties, their weights are correspondingly reduced to ensure that different data sources contribute reasonably during the fusion process and improve the accuracy of the data assimilation results. For N data sources, the calculation formula for the weight w i is:
[0027]
[0028] In the formula, i and j are the indices of the data sources, Δ i is the interval range of the i-th data source, Δ i = U i - L i , Δ j is the interval range of the j-th data source, Δj = U j - L j 。
[0029] Furthermore, for the data source with a smaller uncertainty interval, the range of the given weight is 0.7–1.0; for the data source with a larger uncertainty, the range of the given weight is 0.1–0.3.
[0030] Furthermore, in order to ensure that the sum of all weights is 1, each weight is normalized, and the formula for weight normalization is:
[0031]
[0032] where w i " is the weight after normalization, and w i is the weight calculated by formula (5).
[0033] Furthermore, in step 5, the Kalman gain matrix is optimized, and the soil moisture state formula and the updated error covariance matrix are updated:
[0034] Optimization of the Kalman gain matrix: The calculation formula for the Kalman matrix K is:
[0035] K = PH T (HPH T + R) -1 (7)
[0036] where P is the error covariance matrix of the predicted state, H is the measured matrix, H T is the transpose of the measured matrix, and R is the covariance matrix of the measurement noise;
[0037] An improved covariance matrix is introduced to obtain the optimized Kalman gain matrix
[0038]
[0039] where is the optimized error covariance matrix of the predicted state;
[0040] The updated soil moisture state formula is:
[0041] x t(1 = x t + K(Z t - Hx t ) (9)
[0042] where x t(1 is the soil moisture state at the next moment, x t is the soil moisture state at the current moment, Zt is the measured value of soil moisture at the current moment;
[0043] The formula for updating the covariance matrix is as follows:
[0044] P t(1 =(I - KH)P t (10)
[0045] In the formula, P t(1 is the covariance matrix at the next moment, P t is the covariance matrix at the current moment, and I is the identity matrix.
[0046] The beneficial effects of the present invention are as follows: By introducing a dynamic interval analysis and dynamic weight optimization mechanism into the traditional method, the accuracy of data fusion is significantly improved and the robustness is enhanced. Combining with the Newton iteration method, the uncertainty interval of the data source is dynamically analyzed and the weight is adjusted, improving the accuracy of data fusion; Based on the improved covariance matrix, the Kalman gain matrix is optimized to achieve efficient assimilation of multi-source data and enhance the robustness of data assimilation.
[0047] The following elaborates on the present invention in detail in conjunction with the accompanying drawing description and specific embodiments. Accompanying Drawing Description
[0048] Figure 1 is the flow block diagram of the method of the present invention. Specific Embodiments
[0049] Embodiment 1:
[0050] This embodiment provides a multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF, as Figure 1 shown, including the following steps:
[0051] Step 1, data preprocessing:
[0052] Obtain multi-source data, including measured soil moisture data and ERA-5 reanalysis data. The ERA-5 reanalysis data is the fifth-generation global high-resolution meteorological, climate, and environmental information dataset provided by the European Centre for Medium-Range Weather Forecasts (ECMWF), covering data from 1950 to the present. Perform standardization processing on it to eliminate the dimensional differences between different data sources. Use linear interpolation to align the time steps and bilinear interpolation to match the spatial resolution to ensure the consistency of multi-source data in the time and space dimensions.
[0053] Time alignment is to unify the time axes of all data sources to the same time step and start and end points. Spatial matching is to unify the spatial information of different data sources to the same standard based on spatial resolution and coordinate axes. Normalization is to normalize the units and quantities of different data sources to a unified range for subsequent analysis. For multi-source data, normalization can reduce the impact of different dimensions on the assimilation result. The formula for normalization processing is:
[0054]
[0055] where x i " is the measured value of soil moisture after normalization processing, x i is the original measured value of soil moisture, x min is the minimum moisture in all measured values, x max is the maximum moisture in all measured values.
[0056] Step 2, state prediction:
[0057] Use the distributed soil hydrological model to generate a prediction ensemble of soil moisture, including the initial moisture state and its uncertainty estimate, providing a basis for subsequent uncertainty analysis and data assimilation. The initial ensemble contains multiple possibilities to capture the soil moisture changes in the target area.
[0058] x t(1 = f(x t + u t ) + ∈ t (2)
[0059] where x t(1 is the predicted value of the next state of soil moisture, x t is the current soil moisture state (i.e., the predicted soil moisture value), u t is the control input, which is usually an external influencing factor in the hydrological model, such as precipitation, temperature, evaporation, etc., and these are all important factors affecting soil moisture changes. The function f(x t + u t ) is the evolution law of soil moisture from the current state to the next moment state, and ∈ t is Gaussian white noise, which refers to unpredictable random perturbations or uncertain factors, and it is used to represent the random uncertainty or error in model prediction.
[0060] Calculate the uncertainty interval of the initial state using the Newton iteration method, and combine the error propagation mechanism of the simulation (i.e., the accumulation of errors over time and the spread of errors in space) to generate the error interval for each predicted value, including the upper and lower bounds, and then generate the adjusted dynamic interval range. The accumulation of errors over time means that since the change in soil moisture is dynamic, any initial error will gradually increase or change over time. The simulation error will propagate and accumulate in the subsequent simulation process through the prediction, update, and correction steps at each time step. The spread of errors in space means that if the prediction of soil moisture is based on a spatial model (such as the moisture distribution in a certain area), the initial error will not only propagate in the time dimension but also affect the moisture estimation of the entire area. Especially in places where there is not enough observational data, the error will spread to adjacent areas. The dynamic interval is used to describe the uncertainty range of each data source under different conditions, thereby improving the robustness of data assimilation.
[0061] Newton iteration formula: Assume that f(x n ) describes the uncertainty of soil moisture, then use Newton iteration to determine the root of this function. The iteration formula is:
[0062]
[0063] This process is used to continuously adjust the upper and lower bounds of the uncertainty interval until the convergence condition is met.
[0064] Calculation of the upper and lower bounds of the dynamic interval:
[0065] Assume that for the i-th data source, the upper and lower bounds of the interval are [L i , U i , then update after each iteration:
[0066] L i = x - α i , U i = x - β i (4)
[0067] In the formula, x is the predicted value of the initial state of soil moisture, and α i , β i are the dynamic adjustment values calculated through Newton iteration (the upper and lower boundaries of the error interval of this data source).
[0068] Step 4, dynamic weight adjustment:
[0069] Based on the accuracy and reliability of model prediction data, ERA-5 reanalysis data, and measured data, weights are dynamically assigned to provide optimized weights for data assimilation. Using a non-linear optimization function, the weights of multi-source data are adjusted based on a dynamic interval range. Specifically, for data sources with a smaller uncertainty interval, a higher weight is given, and the weight can be 0.7 - 1.0; for data sources with a larger uncertainty, the weight is correspondingly reduced, and the weight can be 0.1 - 0.3. This adjustment mechanism ensures that different data sources contribute reasonably during the fusion process, improving the accuracy of the data assimilation results.
[0070] For N data sources, the weight w i is calculated by the formula:
[0071]
[0072] where i and j are the indices of the data sources, Δ i is the interval range of the i-th data source, Δ i = U i - L i , Δ j is the interval range of the j-th data source, Δ j = U j - L j .
[0073] Weight normalization: To ensure that the sum of all weights is 1, each weight is normalized:
[0074]
[0075] Step 5, Data assimilation and update:
[0076] Using the improved covariance matrix and the Kalman gain matrix K, multi-source data are assimilated to update the humidity state set. Through the dual improvement of dynamic weights and the improved covariance matrix, the optimal soil moisture estimate after fusion is obtained, i.e., formula (9).
[0077] Calculation of the Kalman gain matrix: The calculation formula of the Kalman gain matrix K is:
[0078] K = PH T (HPH T + R) -1 (7)
[0079] where P is the error covariance matrix of the predicted state, H is the measured matrix of soil moisture data and ERA-5 reanalysis data, and R is the covariance matrix of the measured Gaussian white noise. By introducing the improved covariance matrix the optimized Kalman gain matrix is obtained:
[0080]
[0081]
[0082] State update formula: The formula for updating the soil moisture state is as follows:
[0083] x t(1 = x t + K(Z t - Hx t ) (9)
[0084] In the formula, x t(1 is the soil moisture state at the next moment, x t is the soil moisture state at the current moment, and Z t is the measured value of the soil moisture at the current moment;
[0085] The Kalman gain matrix K is used to adjust the difference between the model prediction and the measured value.
[0086] Update of the covariance matrix: The update formula for the error covariance matrix P t is as follows:
[0087] P t(1 = (I - KH)P t (10)
[0088] In the formula, P t(1 is the covariance matrix at the next moment, P t is the covariance matrix at the current moment, and I is the identity matrix.
[0089] In the dynamically optimized assimilation method, the covariance matrix P t is optimized to P t(1 to better describe the uncertainty of the system state.
[0090] Example 2:
[0091] This example is the application example data of Example 1, as shown in Table 1:
[0092] Table 1 Application example data
[0093]
[0094]
[0095]
[0096] In the table, GBHM represents the soil moisture results output by the hydrological model, ERA-5 represents the soil moisture results extracted from ERA-5 reanalysis data, SVWC represents the measured soil moisture results, and Kalman_Filter_Estimate represents the results after data fusion based on the improved Kalman filter. According to the RMSE results compared with the measured soil moisture data (SVWC): GBHM vs SVWC: RMSE is 0.144; ERA-5 vs SVWC: RMSE is 0.037; Kalman_Filter_Estimate vs SVWC: RMSE is 0.0014. The error between Kalman_Filter_Estimate and SVWC is the smallest, with an RMSE of 0.0014, indicating that the results after Kalman filter fusion are closest to the measured values and have the highest accuracy. According to the calculation, the accuracy improvement of the Kalman filter method (Kalman_Filter_Estimate) compared with GBHM and ERA-5 is as follows: compared with GBHM, the Kalman filter improves the accuracy by 99.06%. Compared with ERA-5, the Kalman filter improves the accuracy by 96.32%.
[0097] Finally, it should be noted that the above is only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred solution, those of ordinary skill in the art should understand that the technical solution of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present invention.
Claims
1. A multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF, characterized in that The method includes the following steps: Step 1, data preprocessing: Obtain multi-source data, including measured soil moisture data and ERA-5 reanalysis data, and perform standardization processing on the obtained multi-source data to eliminate the dimensional differences between different data sources, so that the multi-source data is consistent in the time and space dimensions, thereby providing consistent data input for Step 2; Step 2, state prediction: The distributed soil hydrological model generates an initial humidity state prediction set for the target area, including the initial humidity state and its uncertainty estimate, to capture the soil moisture changes in the target area; Step 3, dynamic interval analysis: Use Newton iteration to calculate the upper and lower bounds of the uncertainty interval of each data source in the multi-source data described in Step 1, and combine the propagation mechanism of the simulation error to generate the error interval of each predicted value, thereby adjusting the interval range of the initial humidity state prediction set; Step 4, dynamic weight adjustment: Optimize the weight allocation function according to the dynamic interval range, and use a non-linear function to dynamically adjust the weights of the multi-source data fusion; Step 5, data assimilation and update: Use the dynamically improved covariance matrix and weights to calculate the Kalman gain matrix and update the initial humidity state prediction set, generating the optimal humidity estimate after multi-source data fusion.
2. The soil moisture multi-source data assimilation method based on dynamic interval analysis and EnKF according to claim 1, characterized in that The standardization processing described in Step 1 includes time alignment, spatial resolution matching, and normalization processing. The formula for normalization processing is: where x i " is the measured value of soil moisture after normalization, x i is the original measured value of soil moisture, x min is the minimum moisture in all measured values, x max is the maximum moisture in all measured values.
3. The multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF according to claim 1, wherein The distributed soil hydrological model described in Step 2 is the GBHM geomorphology-based hydrological model.
4. The multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF according to claim 1, characterized in that The expression of the initial humidity state prediction set described in Step 2 is: x t(1 = f(x t + u t ) + ∈ t (2) where, x t(1 is the predicted value of the next state of soil moisture, x t is the current state of soil moisture, u t is the control input, f(x t + u t ) function is the evolution law of soil moisture from the current state to the next moment state, ∈ t is Gaussian white noise.
5. The multi-source data assimilation method of soil moisture based on dynamic interval analysis and EnKF according to claim 1, characterized in that, The Newton iteration formula described in Step 3 is: where x n(1 is the approximation of the (n + 1)-th iteration, x n is the approximation of the n-th iteration, f(x n ) is the uncertainty of soil moisture, f"(x n ) is the derivative value of the function at x n ; For the i-th data source, the upper and lower bounds of the interval are [L i , U i , and the calculation formula for the dynamic upper and lower bounds of the interval is: L i = x - α i , U i = x - β i (4) where x is the predicted value of the initial state of soil moisture, and α i , β i are the upper and lower boundaries of the data source error range calculated by Newton iteration, respectively.
6. The multi-source data assimilation method for soil moisture based on dynamic interval analysis and EnKF according to claim 1, characterized in that The dynamic adjustment of the weights for multi-source data fusion described in Step 4 is as follows: for data sources with a smaller uncertainty interval, a higher weight is given; for data sources with a larger uncertainty, their weights are correspondingly reduced to ensure that different data sources contribute reasonably during the fusion process and improve the accuracy of the data assimilation results. For N data sources, the weight w i is calculated by the formula: where i and j are the indices of the data sources, and Δ i is the range of the i-th data source, and Δ i = U i - L i , and Δ j is the range of the j-th data source, and Δ j = U j - L j .
7. The soil moisture multi-source data assimilation method based on dynamic interval analysis and EnKF according to claim 6, characterized in that, For the data source with a smaller uncertainty interval, the range of the weight given is 0.7–1.0; for the data source with a larger uncertainty, the range of the weight given is 0.1–0.
3.
8. The soil moisture multi-source data assimilation method based on dynamic interval analysis and EnKF according to claim 6, characterized in that To ensure that the sum of all weights is 1, each weight is normalized. The formula for weight normalization is: where w i " is the weight after normalization, and w i is the weight calculated by formula (5).
9. The soil moisture multi-source data assimilation method based on dynamic interval analysis and EnKF according to claim 1, characterized in that In Step 5, optimize the Kalman gain matrix, update the soil moisture state formula, and update the error covariance matrix: Optimization of the Kalman gain matrix: The calculation formula of the Kalman matrix K is: K = PH T (HPH T +R) -1 (7) where P is the error covariance matrix of the predicted state, H is the measurement matrix, and H T is the transpose of the measurement matrix, and R is the covariance matrix of the measurement noise; Introduce an improved covariance matrix to obtain an optimized Kalman gain matrix Wherein, is the error covariance matrix of the optimized predicted state; The formula for updating the soil moisture state is: x t(1 = x t + K(Z t - Hx t ) (9) where x t(1 is the soil moisture state at the next moment, and x t is the soil moisture state at the current moment, and Z t is the measured value of the soil moisture at the current moment; The formula for updating the covariance matrix is: P t(1 = (I - KH)P t (10) where P t(1 is the covariance matrix at the next moment, and P t is the covariance matrix at the current moment, and I is the identity matrix.
Citation Information
Patent Citations
Soil temperature and humidity data assimilation method based on EnPF
CN107545121A
Data assimilation method for soil water content initial value
CN110390168A
Method for improving wind speed prediction through observation data of Nudging wind power plant
CN114483485A
Medium and long term weather prediction method and system based on variable mesh model, and medium
CN117076738A
Infectious disease prediction mixed data assimilation method based on time-varying set Kalman filtering and KNN classifier
CN119153125A
Cited By
Distributed flood forecasting method based on multi-layer soil humidity assimilation
CN121050001A