Soil water content prediction method under non-consistency condition

By constructing multiple vine structures in soil moisture content prediction and integrating the prediction results using Bayesian model average method, the prediction inaccuracy problem of traditional models under non-consistent conditions is solved, and a higher accuracy and reliable soil moisture content prediction is achieved.

CN120372148APending Publication Date: 2025-07-25NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510450037.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing soil moisture content prediction models are prone to overestimation or underestimation under non-consistent conditions. Traditional vine structure fitting is subjective and cannot accurately reflect the true correlation between soil moisture content and various potential influencing factors.

Method used

The optimal edge distribution was selected using the minimum AICc criterion, all possible vine structures were constructed, probability intervals were generated through Monte Carlo simulations, and prediction results of all potential vine structures were integrated using the Bayesian model averaging method, taking into account soil moisture content prediction under non-consistent conditions.

Benefits of technology

It improves the accuracy and reliability of soil moisture content prediction, can accurately reflect the interaction mode of each predictor under non-consistent conditions, and reduces the uncertainty of the prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372148A_ABST
    Figure CN120372148A_ABST
Patent Text Reader

Abstract

The invention discloses a soil water content prediction method under a non-consistency condition. According to the method, for all fitting edge distributions passing the K-S test, the optimal edge distribution is selected by using the minimum AICc criterion; all possible rattan structures are constructed, and copula fitting and parameter estimation are carried out on each binary variable pair; for each rattan structure, representing conditional distribution of each binary variable pair by using an h function, performing quantile sampling through Monte Carlo simulation to obtain a cumulative distribution function of a target variable on each quantile, and selecting a median as a prediction result of the cumulative distribution function; restoring the cumulative distribution function and the confidence interval based on the fitted edge distribution to obtain a corresponding predicted value of the rattan structure; and integrating the prediction sequences of all the rattan structures by using a Bayesian model averaging method to obtain a final prediction sequence. According to the method, the influence of inconsistency is considered in edge distribution fitting, and the results of all possible rattan structures are integrated, so that the prediction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil water content prediction, and specifically provides a method for predicting soil water content under non-stationary conditions. Background Art

[0002] Soil water content is a key hydrological variable in the land-atmosphere mass and energy exchange. Its abnormal changes may lead to extreme hydrological events such as floods or droughts, posing threats to social and economic development and human life safety. Accurate and reliable soil moisture prediction is crucial for drought prediction, flood warning, agricultural irrigation management, and rational water resources planning.

[0003] Currently, soil water content prediction mainly relies on two categories: physical models and data-driven models. Physical models are abstractions of the real physical relationships between hydrometeorological variables and have strong interpretability, but require a large amount of data as input. Data-driven models represented by machine learning and deep learning regard each variable as a simple time series and can simulate complex non-linear relationships. However, the mechanism of their internal interactions is unknown, and they are prone to overfitting problems. At the same time, during the prediction process, hydrometeorological sequences often show non-stationarity under the influence of time-varying environmental factors. The traditional assumption of stationarity may no longer be applicable, resulting in overestimation or underestimation of the prediction results. Moreover, the correlation between soil water content and various potential influencing factors is not clear.

[0004] Due to the flexibility of vine copula in decomposing multivariate correlation relationships, the vine copula quantile regression model is becoming a possible alternative to machine learning or deep learning models and is applied to soil water content prediction. However, the traditional vine copula quantile regression model uses stationary marginal distributions. Under the background of climate change, the rationality of this assumption is challenged. Considering only stationary margins is likely to cause overestimation or underestimation of the prediction results and is not conducive to guiding actual production. At the same time, due to the complex interactions between soil water content and various influencing factors, the pre-defined single vine structure is often subjective and cannot well represent the true correlation relationship, easily resulting in unreasonable prediction results. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for predicting soil water content under non-stationary conditions, which considers the influence of non-stationarity in marginal distribution fitting and integrates the results of all possible vine structures to improve the prediction performance.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] A method for predicting soil water content under non-stationary conditions, the steps of which include:

[0008] Perform marginal distribution fitting on the soil water content prediction variables and target variables; the prediction variables include, but are not limited to, soil water content, precipitation, air temperature, evapotranspiration data, and the target variable can be soil water content; eight distribution types such as normal distribution, lognormal distribution, gamma distribution, Weibull distribution, Gumbel distribution, logistic distribution, inverse Gaussian distribution, and P-III type distribution are used as candidate marginal distribution types, and the maximum likelihood estimation method is used for parameter estimation respectively;

[0009] For all the fitted marginal distributions that pass the K-S test, select the optimal marginal distribution using the minimum AICc criterion; the marginal distributions include consistent marginal distributions and inconsistent marginal distributions; there is no covariate term in the consistent marginal distribution, that is, the marginal distribution parameters are fixed as constants;

[0010] Construct all possible vine structures, and perform copula fitting and parameter estimation for each pair of binary variables; among them, there are (n - 1)! decomposition methods for n-dimensional variables; The Copula function can construct a high-dimensional joint distribution by connecting each marginal distribution, without being restricted by the type of marginal distribution. For n-dimensional variables, its cumulative joint probability F(x1, x2, ···, x n ) can be defined as:

[0011] F(x1, x2, ···, x n ) = C[F1(x1), F2(x2), ···, F n (x n ) ; θ] = C[u1, u2, ···, u n ; θ]

[0012] where C(·): [0, 1] n → [0, 1] is the copula function with parameter θ; F1(x1) = u1 represents the cumulative distribution function 1 of the margin,

[0013] F2(x2) = u2 represents the cumulative distribution function 2 of the margin, F n (x n ) = u n represents the cumulative distribution function n of the margin; F i (x i ) = u i represents the cumulative distribution function of each margin. There are (n - 1) trees in the vine structure, and the i-th tree has (n - i) edges connecting (n - i + 1) nodes.

[0014] For each vine structure, the conditional distribution of each pair of binary variables is represented using the h function. Quantile sampling is performed through Monte Carlo simulation to obtain the cumulative distribution function of the target variable at each quantile. The 97.5% and 2.5% quantiles are extracted as the upper and lower limits of the confidence interval, and the median of the Monte Carlo simulation sampling is selected as the prediction result of the cumulative distribution function;

[0015] Based on the fitted marginal distributions, the cumulative distribution function and the confidence interval are restored to obtain the predicted values of the corresponding vine structure;

[0016] The Bayesian model averaging method is used to integrate the prediction sequences of all vine structures to obtain the final prediction sequence.

[0017] The Nash-Sutcliffe efficiency coefficient NSE and the coverage ratio CR are respectively used to evaluate the results of deterministic prediction and probabilistic prediction.

[0018] Among them, Monte Carlo simulation is used for multiple samplings. Each sampling result represents a value of the cumulative distribution function. The median of the sampling, which is the 50% quantile, represents the predicted value of the cumulative distribution function. This median also needs to be restored according to the form of the previously fitted consistent or inconsistent margins to obtain the predicted value of soil moisture content.

[0019] According to the above technical solution, the marginal distribution fitting includes inconsistent fitting and consistent fitting;

[0020] The formula for inconsistent fitting of the marginal distributions of the soil moisture content prediction variable and the target variable:

[0021]

[0022] In the formula, g k (·) is the link function of the k-th parameter θ k in the marginal distribution; cov i represents the i-th covariate, and the total number of them is m; β i represents the regression coefficients of each item; β0 represents the regression coefficient; among them, if the parameter k ∈ R, then g k (θ k ) = θ k ; if the parameter k ∈ [0, +∞], then g k (θ k ) = lnθ k . Specifically, both β0 and β i represent the parameters to be fitted and are both fixed values.

[0023] According to the above technical solution, the covariates are selected from time, the El Niño index (NINO3.4), the North Atlantic Oscillation index (NAO), the Arctic Oscillation (AO), the Pacific Decadal Oscillation (PDO), and the Indian Ocean Dipole Mode Index (DMI) and their linear combinations, a total of 63 kinds.

[0024] According to the above technical solution, based on the consistent marginal distribution, after non-consistent fitting, the optimal marginal distribution is selected using the minimum AICc criterion, and then the likelihood ratio test LR is further used to verify the existence of non-consistency. If LR is greater than the (1-α) quantile of the chi-square distribution, it is considered that the marginal distribution is a non-consistent marginal distribution; otherwise, it is considered that the likelihood ratio test is not passed and it is a consistent marginal distribution.

[0025] Among them, α represents the cumulative probability of the chi-square distribution. Generally, 0.05 is taken, that is, when LR is greater than the 95% quantile of the chi-square distribution, it is considered a non-consistent margin.

[0026] According to the above technical solution, the likelihood ratio test formula LR is:

[0027] LR = 2(NLL S -NLL NS );

[0028] In the formula, NLL S represents the negative log-likelihood of the consistent marginal distribution, and NLL NS represents the negative log-likelihood of the non-consistent marginal distribution.

[0029] According to the above technical solution, the posterior distribution p(y|F; Y) of the final prediction sequence is obtained by integration:

[0030]

[0031] Among them, p k (y|f k ; Y) is the prior distribution of the prediction sequence; y represents the final prediction sequence obtained by integration; K is the number of all vine structures; w k = p(f k |Y) is the posterior probability of the vine structure f k under the given training set data Y, and is also regarded as the contribution degree of the k-th vine structure; w k is non-negative and the sum is 1.

[0032] According to the above technical solution, the Nash coefficient efficiency coefficient NSE is:

[0033]

[0034] The coverage ratio CR is:

[0035]

[0036] where y i is the i-th real soil water content sequence, represents the i-th predicted soil water content sequence, represents the mean value of the i-th measured soil water content sequence, n is the sequence length, and C i represents the number of real soil water content values falling within the predicted confidence interval.

[0037] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention incorporates time and atmospheric circulation factors into the marginal distributions of hydrometeorological variables, considers different combinations of covariates, obtains soil water content prediction results under non-uniform conditions, and generates probability intervals based on Monte Carlo simulation; models based on high-dimensional vine copulas, uses different vine structures to explain the interaction patterns of each prediction factor, conducts soil water content prediction from a statistical perspective, and extends the application of quantile regression of copulas to 5 dimensions; at the same time, combines the Bayesian model averaging method to weight-average and integrate the prediction results of all potential vine structures to solve the problem of uncertain vine structure selection, which can improve the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0039] Figure 1 is a flowchart of the steps of a method for predicting soil water content under non-uniform conditions according to the present invention;

[0040] Figure 2 is a statistical chart of the proportion of consistent and non-consistent margins in each variable;

[0041] Figure 3 is the vine copula structure diagram of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Taking 18 stations in Yunnan Province as an example, the method for predicting soil water content under non-uniform conditions according to the present invention is applied to predict the soil water content.

[0044] Monthly soil moisture content, precipitation, air temperature, and evapotranspiration data from 1951 to 2020 at 18 stations in Yunnan Province were selected. The basic situation of the measured soil moisture content values at each station is shown in Table 1.

[0045] Table 1 General situation of research stations

[0046]

[0047] Precipitation, air temperature, evapotranspiration, and antecedent soil moisture content were used as predictive variables, and soil moisture content was used as the target variable. After performing consistent and inconsistent fittings on the predictive variables and target variable of soil moisture content, a goodness-of-fit test for the marginal distribution of the fitting was conducted based on the K-S test, and the optimal marginal distribution was selected using the AICc criterion. For the inconsistent marginal distribution, a likelihood ratio test was further used to verify the existence of its inconsistency. If the likelihood ratio test was not passed, it was considered a consistent marginal. The final marginal distribution fitting results are as Figure 2 .

[0048] Figure 2 In (a) of [], the statistical results of the proportion of consistent and inconsistent margins in each variable are shown. Figure 2 In (b) of [], the statistical results of the covariates of the location parameter in the inconsistent marginal distribution of each variable are shown. Figure 2 In (c) of [], the statistical results of the covariates of the scale parameter in the inconsistent marginal distribution of each variable are shown. Among them, the covariates of the scale parameter in the inconsistent marginal distribution are covariates from time (t), El Niño index (NINO3.4), North Atlantic Oscillation index (NAO), Arctic Oscillation (AO), Pacific Decadal Oscillation (PDO), and Indian Ocean Dipole Mode Index (DMI).

[0049] It can be seen from Figure 2 that most of the margins of the four variables show inconsistency, and the inconsistency of the location parameter is stronger than that of the scale parameter, and both time and environmental factors have an impact. Therefore, it is necessary to consider inconsistency in the fitting of the marginal distribution.

[0050] According to the predictive variables (precipitation, air temperature, evapotranspiration, and antecedent soil moisture content) and the target variable (soil moisture content), a total of 24 different vine structure combinations were determined. The vine copula structure ( Figure 3 ) was selected. For each vine structure, the conditional distribution of each bivariate pair was represented by the h function, and quantile sampling was performed through Monte Carlo simulation to obtain the cumulative distribution function of the target variable at each quantile. The 97.5% and 2.5% quantiles were extracted as the upper and lower limits of the confidence interval. At the same time, the median of multiple samplings in the Monte Carlo simulation, that is, the 50% quantile, represents the predicted value of the cumulative distribution function. Each sampling result represents a value of the cumulative distribution function.

[0051] For example, given the marginal cumulative distribution functions u2, u3, u4, u5 of four predictor variables and the cumulative distribution function u1 of the target variable, one of the decomposition equations of the C-vine can be expressed as:

[0052]

[0053] For a given quantile τ ∈ [0, 1], the predicted cumulative distribution of u1 can be represented by the inverse function of the h function:

[0054]

[0055] In this embodiment, the Monte Carlo simulation sampling technique is adopted to randomly extract 800 quantiles from the uniform distribution. Therefore, 800 groups of predicted sequences of cumulative distributions are obtained. The 97.5% and 2.5% quantiles are extracted as the upper and lower limits of the confidence interval, and the median of the multiple sampling results of the Monte Carlo simulation is selected as the prediction result of the cumulative distribution function.

[0056] According to the marginal distributions of the predictor variables and the target variable, the predicted values of the cumulative distributions of each quantile are restored to the predicted sequences. The same operation is performed on all vine structures, and a total of 24 groups of predicted sequences and their confidence intervals are obtained.

[0057]

[0058] In the formula, x1 represents the target variable, x2 represents the predictor variable precipitation, x3 represents the predictor variable temperature, x4 represents the predictor variable evapotranspiration, x5 represents the predictor variable antecedent soil moisture content, F(.) represents the fitted marginal distribution; F(.) -1 represents the inverse function of the fitted marginal distribution, u1 represents the cumulative distribution function of the target variable at the quantile, represents the predicted sequence obtained at the given quantile τ.

[0059] The Bayesian model averaging method is used to integrate the predicted sequences of all vine structures to obtain the final predicted sequence.

[0060] Two methods, the optimal prediction result (VQR-NS) of a single vine structure and the random forest model (RF), are selected and compared with the EVQR-NS method proposed in the present invention. The Nash coefficient efficiency coefficient NSE and the coverage ratio CR are respectively used as the evaluation indexes for deterministic prediction and probabilistic prediction. The comparison evaluation results of the Nash efficiency coefficient NSE of each model are shown in Table 2, and the comparison of the coverage rate CR of each model is shown in Table 3. Among them, the closer NSE and CR are to 1, the higher the model accuracy.

[0061] Table 2 Comparison of the Nash efficiency coefficient NSE of each model

[0062]

[0063] Table 3 Comparison of Coverage Rate CR of Each Model

[0064]

[0065]

[0066] As can be seen from the two tables, the VQR-NS model considering the non-stationarity of marginal distribution has a certain improvement in the prediction accuracy of the random forest model, and the EVQR-NS model considering vine structure integration can further improve the performance of VQR-NS, showing the best effect among the three models.

[0067] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0068] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for predicting soil water content under non-uniform conditions, characterized in that The steps include: Fitting the marginal distributions of the soil water content prediction variables and the target variables; For all the fitted marginal distributions passing the K-S test, selecting the optimal marginal distribution using the minimum AICc criterion; the marginal distributions include consistent marginal distributions and inconsistent marginal distributions; Constructing all possible vine structures, and performing copula fitting and parameter estimation for each pair of binary variables; For each vine structure, using the h function to represent the conditional distribution of each pair of binary variables, conducting quantile sampling through Monte Carlo simulation, obtaining the cumulative distribution function of the target variable at each quantile, extracting the 97.5% and 2.5% quantiles as the upper and lower limits of the confidence interval, and selecting the median of the Monte Carlo simulation sampling as the prediction result of the cumulative distribution function; Based on the fitted marginal distributions, restoring the cumulative distribution function and the confidence interval to obtain the predicted values of the corresponding vine structure; Using the Bayesian model averaging method to integrate the prediction sequences of all vine structures to obtain the final prediction sequence; Evaluating the deterministic prediction and the probabilistic prediction using the Nash-Sutcliffe efficiency coefficient NSE and the coverage ratio CR respectively.

2. The soil water content prediction method under inconsistent conditions according to claim 1, characterized in that The marginal distribution fitting includes inconsistent fitting and consistent fitting; The formula for inconsistent fitting of the marginal distributions of the soil water content prediction variables and the target variables: where, g k (·) is the link function of the k-th parameter θ k in the marginal distribution; cov i represents the i-th covariate, and the total number thereof is m; β i represents the regression coefficients of each item; β0 represents the regression coefficient; where, if the parameter k ∈ R, then g k (θ k ) = θ k ; if the parameter k ∈ [0, +∞], then g k (θ k ) = lnθ k .

3. The soil moisture content prediction method under non-uniform conditions according to claim 2, wherein The covariates are selected from time, El Niño index, North Atlantic Oscillation index, Arctic Oscillation, Pacific Decadal Oscillation, and Indian Ocean Dipole Mode Index and their linear combinations.

4. A method for predicting soil water content under non-uniform conditions according to claim 1, characterized in that, After inconsistent fitting, using the minimum AICc criterion to select the optimal marginal distribution and further verifying the existence of inconsistency using the likelihood ratio test LR. If the likelihood ratio test is passed, it is considered that the marginal distribution is an inconsistent marginal distribution; otherwise, it is considered that the likelihood ratio test is not passed and it is a consistent marginal distribution.

5. The soil moisture content prediction method under non-uniform conditions according to claim 4, wherein The likelihood ratio test formula LR: LR = 2(NLL S - NLL NS ); where, NLL S represents the negative log-likelihood of the consistent marginal distribution, and NLL NS represents the negative log-likelihood of the inconsistent marginal distribution.

6. The soil water content prediction method under non-uniform conditions according to claim 1, characterized in that, Integrating to obtain the posterior distribution p(y|F; Y) of the final prediction sequence: where p k (y|f k ; Y) is the prior distribution of the prediction sequence; y represents the finally predicted sequence obtained by integration; K is the number of all vine structures; w k = p(f k |Y) is the posterior probability of the vine structure f k under the given training set data Y, and is also regarded as the contribution degree of the k-th vine structure; w k is non - negative and the sum is 1.

7. A method for predicting soil water content under non-uniform conditions according to claim 1, characterized in that, The Nash-Sutcliffe efficiency coefficient NSE: The coverage ratio CR: where y i is the i-th real soil water content sequence, represents the i-th predicted soil water content sequence, represents the mean value of the i-th measured soil water content sequence, n is the sequence length, and C i represents the number of true soil water content values falling within the predicted confidence interval.