Method for assimilation of multi-source soil moisture data based on dynamic interval analysis and EnKF

The integration of dynamic interval analysis and EnKF optimizes weight distribution and covariance adjustment to enhance the accuracy and robustness of soil moisture data fusion, addressing the limitations of conventional methods.

JP7832739B1Active Publication Date: 2026-03-18CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

Conventional methods struggle to dynamically adjust the weights of multi-source data and lack effective interval analysis methods, leading to inaccurate soil moisture estimation due to unaccounted uncertainty and error ranges.

Method used

A method combining dynamic interval analysis and Ensemble Kalman Filter (EnKF) to optimize weight distribution and covariance adjustment, using normalization, state prediction, and iterative methods to adjust uncertainty intervals and weights for improved data fusion.

Benefits of technology

Enhances the accuracy and robustness of soil moisture data fusion by dynamically adjusting weights and optimizing Kalman gain matrices, reducing uncertainty intervals and improving data assimilation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832739000001_ABST
    Figure 0007832739000001_ABST
Patent Text Reader

Abstract

This invention provides a method for assimilating multi-source soil moisture data based on dynamic interval analysis and aggregate Kalman filtering (EnKF). [Solution] The method includes five steps: data preprocessing, state prediction, dynamic interval analysis, dynamic weight adjustment, and data assimilation update. The present invention introduces dynamic interval analysis and dynamic weight optimization mechanisms to conventional methods, significantly improving the accuracy of data fusion and reinforcing robustness. It combines with Newton's iterative method to dynamically analyze uncertainty intervals of data sources and adjust weights to improve the accuracy of data fusion, optimizes the Kalman gain matrix based on an improved covariance matrix, achieves highly efficient assimilation of multi-source data, and reinforces the robustness of data assimilation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technical field for processing hydrological multi-source data, and more specifically to a method for assimilation of soil moisture multi-source data based on dynamic interval analysis and EnKF. [Background technology]

[0002] Soil moisture is a key variable in hydrological, agricultural, and meteorological models, and the accuracy of its estimation is of crucial importance for drought monitoring, irrigation management, and climate forecasting. However, conventional methods for estimating soil moisture have many problems, including the following: (1) It is difficult to dynamically adjust the weights of multi-source data to reflect changes in uncertainty. In actual applications, the reliability of soil moisture data provided by different data sources changes depending on factors such as time and space. Conventional methods cannot rationally adjust the weights of each data source in the data fusion process based on the real-time changes in the uncertainty of the data sources, and as a result, the data fusion results cannot accurately reflect the true soil moisture conditions. (2) There are few effective interval analysis methods, making it difficult to accurately represent the error ranges of different data sources. Due to the lack of effective interval analysis methods, it is not possible to accurately identify the error ranges of each data source. This means that when merging the data, the uncertainty of the data cannot be adequately considered, which affects the accuracy of soil moisture estimation.

[0003] The Ensemble Kalman Filter (EnKF) is an optimization method used for assimilating non-linear and high-dimensional data. It approximates the state distribution of a system by introducing a set of "ensemble" states, thereby performing data fusion in complex systems. This is particularly suitable for problems related to high-dimensional and complex non-linear changes such as soil moisture. The EnKF overcomes the limitations of the conventional Kalman Filter (KF) when dealing with non-linear problems and high-dimensional data. By using multiple prediction ensembles as state estimates, it significantly improves the model's adaptability to observed data and computational efficiency. In data processing, especially when there is significant uncertainty or noise, the EnKF can effectively process and fuse data from multiple sources.

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to solve the problems that it is difficult to dynamically adjust the weights of multi-source data existing in the above conventional soil moisture estimation method and there is a lack of effective interval analysis methods, the present invention provides a method for assimilating multi-source data based on dynamic interval analysis and EnKF, aiming to improve the accuracy of data fusion by optimizing the data weight distribution and covariance adjustment mechanism.

Means for Solving the Problems

[0005] The present invention is realized as follows. A method for assimilating soil moisture multi-source data based on dynamic interval analysis and EnKF, the method comprising: Step 1, data preprocessing: obtaining multi-source data including measured soil moisture data and ERA-5 reanalysis data, performing normalization processing on the obtained multi-source data to eliminate the dimensional differences between different data sources, and maintaining the consistency of the multi-source data in the time and space dimensions, thereby providing data input that conforms to Step 2; Step 2, State Prediction: The distributed soil hydrological model generates an initial humidity state prediction set for the target area including the initial humidity state and its estimated uncertainty value, and captures the soil humidity changes within the target area; Step 3, Dynamic Interval Analysis: Using Newton iteration to calculate the upper and lower boundaries of the uncertainty intervals of each data source in the multi-source data described in Step 1, and generating the error intervals of each predicted value in combination with the propagation mechanism of the simulation error, thereby adjusting the interval range of the initial humidity state prediction set; Step 4, Dynamic Weight Adjustment: Optimizing the weight distribution function based on the dynamic interval range, and performing dynamic adjustment on the fusion weights of the multi-source data using a non-linear function; Step 5, Data Assimilation and Update: Utilizing the dynamically improved covariance matrix and weights, calculating the Kalman gain matrix and updating the initial humidity state prediction set to generate the optimal humidity estimation after multi-source data fusion. It includes the following steps:

[0006] Furthermore, the normalization process described in Step 1 includes time alignment, spatial resolution matching, and normalization processing. The formula for the normalization processing is as follows:

Number

Number

[0007] Furthermore, the distributed soil hydrological model described in Step 2 is a hydrological model based on the terrain of GBHM.

[0008] Furthermore, the expression formula for the initial humidity state prediction set described in Step 2 is as follows:

Number

Number

[0009] Furthermore, the Newton iteration formula described in Step 3 is as follows,

Number

Number

[0010] Furthermore, dynamically adjusting the fusion weights of multi-source data as described in Step 4 is as follows: relatively high weights are assigned to data sources with relatively small uncertainty intervals, and their weights are correspondingly reduced to data sources with relatively large uncertainty, ensuring that the contributions of different data sources in the fusion process are reasonable, improving the accuracy of the data assimilation results, and assigning weights w to N data sources. i The formula for calculating this is as follows:

number

[0011] Furthermore, for data sources with a relatively small uncertainty interval, the range of weights to be assigned is 0.7 to 1.0, and for data sources with a relatively large uncertainty, the range of weights to be assigned is 0.1 to 0.3.

[0012] Furthermore, in order to ensure that the sum of all weights is 1, normalization is performed on each weight, and the weight normalization formula is as follows:

number

number

[0013] Furthermore, in step 5, the formula for optimizing the Kalman gain matrix and updating the soil moisture state and error covariance matrix is ​​updated: Optimization for the Kalman gain matrix: The formula for calculating the Kalman matrix K is as follows:

number

number

number

number

number

number

[0014] The beneficial effects of this invention are as follows: By introducing dynamic interval analysis and dynamic weight optimization mechanisms to conventional methods, the accuracy of data fusion is significantly improved and robustness is enhanced. By combining it with Newton's iterative method, uncertainty intervals of data sources are dynamically analyzed and weights are adjusted to improve the accuracy of data fusion, and the Kalman gain matrix is ​​optimized based on the improved covariance matrix, enabling efficient assimilation of multi-source data and enhancing the robustness of data assimilation.

[0015] The present invention will be described in detail below with reference to the drawings and specific embodiments. [Brief explanation of the drawing]

[0016] [Figure 1] This is a flowchart of the method of the present invention. [Modes for carrying out the invention]

[0017] [Example 1] This embodiment provides a method for assimilating multi-source soil moisture data based on dynamic interval analysis and EnKF, and includes the following steps, as shown in Figure 1. That is, Step 1, Data Preprocessing: Multi-source data, including measured soil moisture data and ERA-5 reanalysis data, is acquired. The ERA-5 reanalysis data is a fifth-generation global high-resolution meteorological, climate, and environmental information dataset provided by the European Centre for Medium-Range Weather Forecasts (ECMWF), containing data from 1950 to the present. This data is standardized to eliminate dimensional differences between different data sources. Linear interpolation is used to align the time step size, and bilinear interpolation is employed to match the spatial resolution, ensuring consistency in the time and spatial dimensions of the multi-source data. Time alignment unifies the time axes of all data sources to the same time step size and start / stop points; spatial matching unifies the spatial information of different data sources to the same standard based on spatial resolution and coordinate axes; and normalization normalizes the units and quantities of different data sources to a unified range, facilitating subsequent analysis. For multi-source data, normalization can reduce the impact of different dimensions on the assimilation results. The formula for the normalization process is as follows:

number

number

[0018] Step 2, State prediction: A distributed soil hydrological model is used to generate a predicted set of soil moisture, including initial moisture conditions and their uncertainty estimates, providing a basis for subsequent uncertainty analysis and data assimilation. The initial set includes various possibilities and captures soil moisture changes within the target area.

number

number

[0019] The uncertainty interval of the initial state is calculated using the Newton iteration method, and combined with the simulation error propagation mechanism (i.e., temporal accumulation of errors and spatial expansion of errors) to generate error intervals for each predicted value, including the upper and lower boundaries, and further to generate the adjusted dynamic interval range. Temporal accumulation of errors means that, because soil moisture changes are dynamic, each initial error gradually increases or changes over time. Simulation errors are propagated and accumulated in the subsequent simulation process through prediction, update, and correction steps of each time step size. Spatial expansion of errors means that, if the soil moisture prediction is based on a single spatial model (e.g., the moisture distribution of a region), the initial error not only propagates in the time dimension but also affects the moisture estimate for the entire region, and especially where observational data is insufficient, the error expands to adjacent regions. Dynamic intervals are used to describe the uncertainty range under different conditions for each data source, thereby improving the robustness of data assimilation. Newton's iterative formula: Let's assume f(x n If ) explains the uncertainty of soil moisture, we use Newtonian iteration to determine the square root of the function. The iteration formula is as follows:

number

number

[0020] Step 4, Dynamic Weight Adjustment: Based on the accuracy and reliability of model prediction data, ERA-5 reanalysis data, and measured data, weights are dynamically allocated to provide optimized weights for data assimilation. Based on the dynamic interval range, the weights of multi-source data are adjusted using a nonlinear optimization function. Specifically, relatively high weights are given to data sources with relatively small uncertainty intervals, which may be between 0.7 and 1.0, while the weights of data sources with relatively large uncertainty are correspondingly reduced, which may be between 0.1 and 0.3. This adjustment mechanism ensures that the contributions of different data sources in the fusion process are reasonable, improving the accuracy of the data assimilation results. For N data sources, weight w i The formula for calculating this is as follows:

number

number

[0021] Step 5, Data Assimilation and Update: Improved covariance matrix

number

number

number

number

number

number

number

[0022] [Example 2] This embodiment is an application example data of Example 1, and is shown in Table 1.

[0023] Application example data [Table 1] JPEG0007832739000033.jpg172161

[0024] In the table, GBHM represents the soil moisture result output from the hydrological model, ERA-5 represents the soil moisture result extracted from ERA-5 reanalysis data, SVWC represents the measured soil moisture result, and Kalman_Filter_Estimate represents the result fused based on improved Kalman filter data. Based on the RMSE results compared with the measured soil moisture data (SVWC): GBHM vs SVWC: RMSE is 0.144, ERA-5 vs SVWC: RMSE is 0.037, and Kalman_Filter_Estimate vs SVWC: RMSE is 0.0014. The error between Kalman_Filter_Estimate and SVWC is minimal, and the RMSE is 0.0014, explaining that the result fused by the Kalman filter is closest to the measured value and has the highest accuracy. Based on calculations, the Kalman filter method (Kalman_Filter_Estimate) improves accuracy compared to GBHM and ERA-5 as follows: Compared to GBHM, the Kalman filter improves accuracy by 99.06%. Compared to ERA-5, the Kalman filter improves accuracy by 96.32%.

[0025] Finally, it should be noted that the above description is not limited to explaining the technical solutions of the present invention, and although the present invention has been described in detail with reference to preferred solutions, as those skilled in the art will understand, modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for assimilating multi-source soil moisture data based on dynamic interval analysis and EnKF, wherein the method is: Step 1, Data Preprocessing: This step involves acquiring multi-source data, including measured soil moisture data and ERA-5 reanalysis data, performing standardization on the acquired multi-source data to eliminate dimensional differences between different data sources, maintaining consistency in the temporal and spatial dimensions of the multi-source data, and thereby providing data input that matches Step 2. Step 2, State Prediction: A dispersed soil hydrological model generates a set of initial humidity state predictions for the target area, including the initial humidity state and its uncertainty estimate, and captures soil humidity changes within the target area. Step 3, Dynamic Interval Analysis: This step involves using Newton iterations to calculate the upper and lower boundaries of the uncertainty intervals for each data source in the multi-source data described in Step 1, and generating error intervals for each predicted value in combination with the simulation error propagation mechanism, thereby adjusting the interval range of the initial humidity state prediction set. Step 4, Dynamic Weight Adjustment: The weight allocation function is optimized based on the dynamic interval range, and dynamic adjustments are made to the fusion weights of multi-source data using a nonlinear function. Data sources with relatively small uncertainty intervals are given relatively high weights, with weights ranging from 0.7 to 1.

0. Data sources with relatively large uncertainty have their weights correspondingly reduced, with weights ranging from 0.1 to 0.

3. This ensures that the contributions of different data sources in the fusion process are reasonable, improving the accuracy of the data assimilation results. For N data sources, the formula for calculating the weights w i is as follows: [Math 1] In the formula, i and j are the indices of the data sources, Δi is the interval range of the i-th data source, Δi = U i - L i, and Δj is the interval range of the j-th data source, Δj = U j - L j, and the step, Step 5, Data Assimilation and Updating: This step involves using dynamically improved covariance matrices and weights to calculate the Kalman gain matrix and update the initial humidity state prediction set to generate the optimal humidity estimate after multi-source data fusion. A method for assimilation of multi-source soil moisture data based on dynamic interval analysis and EnKF, characterized by including the above.

2. The standardization process described in Step 1 includes time alignment, spatial resolution matching, and normalization, and the formula for the normalization process is as follows: [Math 2] In the formula, [Math 3] x is the measured soil moisture value after normalization. i This is the original measured soil moisture value, x min x is the minimum humidity for all measured values, max The method for assimilation of multi-source soil moisture data based on dynamic interval analysis and EnKF, as described in claim 1, characterized in that is the maximum humidity in all measured values.

3. The method for assimilation of multi-source soil moisture data based on dynamic interval analysis and EnKF, as described in claim 1, characterized in that the dispersed soil hydrological model described in step 2 is a hydrological model based on the topography of GBHM.

4. The expression for the initial humidity state prediction set described in Step 2 is as follows: [Math 4] In the equation, x t+1 x is the predicted value for the next state of soil moisture, t This represents the current soil moisture state, u t is the control input, and f(x t +u t The function is the evolutionary rule for soil moisture from its current state to its state at the next time point. [Math 5] The method for assimilation of soil moisture multi-source data based on dynamic interval analysis and EnKF according to claim 1, characterized in that is Gaussian white noise.

5. The Newton iterative formula described in Step 3 is as follows: [Math 6] In the formula, x n+1 is the approximation of the (n + 1)-th iteration, x n is the approximation of the n-th iteration, f(x n ) is the uncertainty of soil moisture, f′(x n ) is the derivative value of the function at x n , For the i-th data source, the upper and lower boundaries of the interval are [L i ,U i The formula for calculating the upper and lower boundaries of the dynamic interval is as follows: [Number 7] In the formula, x is the initial predicted value of soil moisture, and α i , β i The dynamic interval analysis and assimilation method for soil moisture multi-source data based on EnKF according to claim 1, characterized in that the upper and lower boundaries of the data source error interval obtained by Newton iterative calculation are respectively.

6. To ensure that the sum of all weights is 1, we perform normalization on each weight, and the weight normalization formula is as follows: [Number 8] In the formula, [Number 9] is the weight after normalization, and w i The method for assimilation of multi-source soil moisture data based on dynamic interval analysis and EnKF according to claim 1, characterized in that is a weight calculated by formula (1).

Citation Information

Cited By

  • Soil humidity spatial characteristic evaluation method and storage medium

    CN121935862A