A method and system for predicting soil salinization based on altitude

By using component-by-component Metropolis-Hastings algorithm and logistic regression model in soil salinization prediction, the problem of effective utilization of altitude data in soil salinization prediction is solved, and efficient random prediction of soil salinization is achieved.

CN115618720BActive Publication Date: 2025-05-30SHANDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211180898.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-05-30
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

There is still a lack of effective solutions for the prediction methods of soil salinization based on altitude in the prior art, especially when dealing with the random characteristics of soil salinization.

Method used

The component-by-component Metropolis-Hastings algorithm is used to generate random numbers from the posterior distribution, and the logistic regression parameters are updated one by one according to the components in combination with the proposed distribution, and the a priori distribution is updated through the likelihood function to realize the random prediction and forecast of soil salinization.

Benefits of technology

The mixing efficiency of Markov chains is improved, and the accuracy of soil salinization is achieved, and the adjustment parameters are not required to be considered, making it easy to apply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618720B_ABST
    Figure CN115618720B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of soil salinization prediction, and provides a method and system for predicting soil salinization based on altitude, including: obtaining the altitude of the area to be predicted; based on the altitude, using a logistic regression model to predict the binary status of soil salinization in the area to be predicted; wherein, the construction method of the logistic regression model is: based on the altitudes and the binary status of soil salinization of a number of sampling points collected, generate random numbers from the posterior distribution by the component-by-component Metropolis-Hastings sampling method, and update the logistic regression parameters component by component in combination with the proposal distribution, and update the prior distribution through the likelihood function. To achieve random prediction and forecasting of soil salinization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of soil salinization prediction, and particularly relates to a method and system for predicting soil salinization based on altitude. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] The occurrence of salinization and the formation of saline soil are mainly affected by various factors such as nature and human activities, especially controlled by the water-salt movement process of the migration and redistribution of water and salt in the soil body. The process of soil water-salt movement and its regulation mechanism are one of the core issues in the current research on saline soil. The generation of saline soil is the result of the combined action of factors such as climate, soil, topography, and human conditions. Soil salinization prediction and forecasting, in a broad sense, is to predict and forecast the development direction of salinized soil that has already occurred and the possibility and degree of salinization (commonly known as secondary salinization) of non-salinized soil based on the dynamic monitoring data of soil and groundwater water-salt, topography, and river irrigation comprehensive conditions in a region; in a narrow sense, it refers to the prediction and forecasting of secondary salinization. Generally speaking, the research on salinization prediction and forecasting can be divided into three levels:

[0004] (1) Qualitative forecasting is carried out through the study of natural environmental conditions and the occurrence and development laws of soil salinization. By comparing and analyzing the natural conditions of the forecast area and the already salinized area, the possibility of soil salinization is predicted based on expert knowledge and experience. The commonly used methods are the geographical similarity method and the expert forecasting method. The geographical similarity method mainly makes a qualitative forecast based on the available information and proposes the possibility of secondary salinization. This method is difficult to make a more detailed quantitative forecast, so it is a relatively rough forecasting method; the expert forecasting method mainly estimates and preliminarily predicts the possibility of salinization based on expert experience and the characteristics of local natural conditions.

[0005] (2) Semi - quantitative prediction is carried out by combining methods such as water - salt balance and probability statistics. Methods such as water - salt balance and probability statistics are important methods for the transitional research of regional water - salt dynamics from qualitative to quantitative research. Since these methods require relatively simple and easily obtainable input data, they are often used in the prediction of regional soil salinization. The water - salt balance method is based on the theory of mass conservation. Regional water - salt balance can quantitatively characterize the development direction of regional soil salinization. The deficiency of regional water - salt balance research is that it is relatively weak in revealing the mechanism of the movement of regional water and salt, and there are also deficiencies in elaborating the internal relationship between regional soil water and salt, and it cannot accurately predict the distribution of regional water and salt well. Due to the variability of soil properties itself and the randomness of various factors affecting water - salt movement, such as precipitation, evaporation, groundwater level, and groundwater depth, regional water - salt movement has obvious randomness. Therefore, methods such as probability statistics and genetic analysis methods are also often used to study the random characteristics of water - salt movement. This type of method considers the randomness characteristics of water - salt movement and has a certain degree of flexibility, but its portability is poor and it is mostly used for research. In the actual application of saline - alkali land improvement, it still needs to be strengthened.

[0006] (3) On the basis of regional water - salt dynamics research, a mathematical model is established, and with the help of a computer, quantitative prediction of soil salinization is carried out. "Salt goes with water and comes with water". Soil salinization and potential soil salinization are affected by complex factors. Although the groundwater level and depth are important for the prediction of soil salinity, the difficulty of real - time dynamic monitoring of the groundwater level and depth has not been popularized to a certain extent for the prediction of soil salinization and potential salinization. Topography and geomorphology are more practically significant for the prediction of salinity. Altitude is relatively easy to obtain and has strong feasibility for the prediction of soil salinization. In the existing technology, there is still a lack of effective solutions for identifying and dynamically monitoring soil salinization and potential salinization with obvious randomness characteristics based on altitude. Summary of the Invention

[0007] To solve the technical problems existing in the above - mentioned background technology, the present disclosure provides a method and system for predicting soil salinization based on altitude. The Metropolis - Hastings algorithm with component - by - component is used to generate random numbers from the posterior distribution, which improves the mixing efficiency of the Markov Chain. In the Metropolis - Hastings algorithm, updates are carried out one by one according to components without considering adjustment parameters, realizing the random prediction of soil salinization.

[0008] To achieve the above object, the present disclosure adopts the following technical solutions:

[0009] The first aspect of the present disclosure provides a method for predicting soil salinization based on altitude, which includes:

[0010] Obtain the altitude of the area to be predicted;

[0011] Based on the altitude, use the logistic regression model to predict the binary status of soil salinization in the area to be predicted;

[0012] Among them, the construction method of the logistic regression model is: based on the altitudes and the binary status of soil salinization of a number of sampling points collected, generate random numbers from the posterior distribution by the component-by-component Metropolis-Hastings sampling method, and update the logistic regression parameters one by one according to the components in combination with the proposal distribution, and update the prior distribution through the likelihood function.

[0013] Furthermore, the specific method for generating random numbers from the posterior distribution is:

[0014] (301) Let β = (β 0 , β 1 ) T , β 0 and β 1 be the logistic regression parameters obtained in the (t - 1)-th iteration;

[0015] (302) Generate a candidate point β′ 0 from the first proposal distribution;

[0016] (303) Let β′ = (β′ 0 , β 1 ) T , and calculate the first acceptance probability;

[0017] (304) Accept β = β′ with the first acceptance probability; otherwise, β remains unchanged;

[0018] (305) Generate a candidate point β′ 1 from the second proposal distribution;

[0019] (306) Let β′ = (β 0 , β′ 1 ) T , and calculate the second acceptance probability;

[0020] (307) Accept β = β′ with the second acceptance probability; otherwise, β remains unchanged;

[0021] (308) Let t = t + 1, return to (301), and until t reaches the set value, output β.

[0022] Furthermore, the first acceptance probability is:

[0023] α 0 (β, β′0 ) = min{1, A}

[0024] Wherein, β = (β 0 , β 1 ) T , β 0 and β 1 are logistic regression parameters, β′ 0 is the candidate point generated by the first proposal distribution, y is the binary status vector of soil salinization of all sample points, π() represents the prior distribution, and f() represents the target distribution.

[0025] Furthermore, the second acceptance probability is:

[0026] α 1 (β, β′ 1 ) = min{1, B}

[0027] Wherein, β = (β 0 , β 1 ) T , β 0 and β 1 are logistic regression parameters, β′ 1 is the candidate point generated by the second proposal distribution, y is the binary status vector of soil salinization of all sample points, π() represents the prior distribution, and f() represents the target distribution.

[0028] Furthermore, the proposal distribution makes the generated Markov chain satisfy irreducibility, positive recurrence, and aperiodicity.

[0029] Furthermore, the prior distribution is an independent normal distribution.

[0030] The second aspect of the present disclosure provides a soil salinization prediction system based on altitude, which includes:

[0031] A model construction module configured to: generate random numbers from the posterior distribution by the component-by-component Metropolis-Hastings sampling method based on the altitude and the binary status of soil salinization of a plurality of sampled points collected, and update the logistic regression parameters component by component in combination with the proposal distribution, and update the prior distribution through the likelihood function;

[0032] A data acquisition module configured to: acquire the altitude of the area to be predicted;

[0033] A prediction module configured to: predict the binary status of soil salinization of the area to be predicted based on the altitude by using a logistic regression model.

[0034] Further, the prior distribution is an independent normal distribution.

[0035] The third aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in a method for predicting soil salinization based on altitude as described above are implemented.

[0036] The fourth aspect of the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a method for predicting soil salinization based on altitude as described above are implemented.

[0037] Compared with the prior art, the beneficial effects of the present disclosure are:

[0038] The present disclosure provides a method for predicting soil salinization based on altitude, which realizes the stochastic prediction and forecasting of soil salinization based on altitude through the component-by-component method of Markov Chain Monte Carlo (MCMC).

[0039] The present disclosure provides a method for predicting soil salinization based on altitude. This method only requires relatively few altitude samplings to obtain a relatively accurate approximation of whether salinization occurs. The target posterior distribution obtained according to the prior distribution, likelihood function, and proposal distribution transforms the process of integrating the posterior into the process of summing the vector formed by the sampling chain. The posterior high-density map of soil salinization risk prediction provides the possibility of soil salinization prediction in a concise and intuitive manner.

[0040] The present disclosure provides a method for predicting soil salinization based on altitude, which generates random numbers from the posterior distribution using the component-by-component Metropolis-Hastings algorithm, improving the mixing efficiency of the chain; in the Metropolis-Hastings algorithm, updates are performed one by one by component, automatically performing inference without the need to consider adjustment parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The specification drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.

[0042] Figure 1 is a flowchart of a method for predicting soil salinization based on altitude according to Embodiment 1 of the present disclosure;

[0043] Figure 2(a) is the component-by-component M-H algorithm logistic regression parameter β of Embodiment 1 of the present disclosure 0Sample path diagram;

[0044] Figure 2(b) is the sample path diagram of the component-wise M-H algorithm logistic regression parameter β in the first embodiment of the present disclosure; 1 Sample path diagram;

[0045] Figure 3(a) is the component-wise M-H algorithm logistic regression parameter β in the first embodiment of the present disclosure; 0 Traversal mean diagram;

[0046] Figure 3(b) is the component-wise M-H algorithm logistic regression parameter β in the first embodiment of the present disclosure; 1 Traversal mean diagram;

[0047] Figure 4(a) is the component-wise M-H algorithm logistic regression parameter β in the first embodiment of the present disclosure; 0 Autocorrelation diagram;

[0048] Figure 4(b) is the component-wise M-H algorithm logistic regression parameter β in the first embodiment of the present disclosure; 1 Autocorrelation diagram;

[0049] Figure 5 It is the schematic diagram of the posterior density of the soil salinization prediction β 0 component parameter and the 99%, 95%, 50% highest density regions;

[0050] Figure 6 It is the soil salinization prediction β in the first embodiment of the present disclosure; 1 Schematic diagram of the posterior density of the component parameter and the 99%, 95%, 50% highest density regions. Detailed implementation manners

[0051] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.

[0052] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0053] Term explanation:

[0054] Geographical similarity method: Forecast is made by comparing with the situation of the already salinized areas with environmental conditions similar to those of the forecast area. Using this method, it is necessary to conduct a comprehensive investigation and analysis and a detailed comparative study on the conditions of the forecast area and the areas where salinization has occurred, so as to make a forecast. The main environmental factors, natural conditions and human activity conditions of the geographical area need to be considered. The characteristics of natural conditions include: (1) meteorological and hydrological conditions, (2) topographic and geomorphic conditions, (3) geological and hydrogeological conditions, (4) soil conditions (including soil types, distributions, soil physical properties, soil physical chemistry and biological characteristics and their spatial distributions), (5) geochemical characteristics. The human activity conditions mainly refer to: (1) land use conditions, (2) agricultural technical measures, (3) soil improvement measures, (4) water conservancy technical measures (including irrigation, drainage, water storage and other measures).

[0055] Water-salt balance method: By calculating the water-salt balance of the forecast area, the possibility of soil salinization occurrence is forecast. The commonly used salt balance calculation formula is:

[0056] Δs = [P + I + R + G + W + F] - [L p + L i + r + g + u]

[0057] In the formula, Δs is the change in the total salt storage in a certain area; P is the salt amount brought by the atmosphere (atmospheric precipitation, wind transportation, etc.); I is the salt amount brought by irrigation water; R is the salt amount brought by surface water (runoff, flood, waterlogging); G is the salt amount input from below the surface (groundwater, deep soil water); W is the salt amount increased by the local weathering process; F is the salt amount brought by fertilizers and chemical amendments; L p is the salt amount leached by atmospheric precipitation; L i is the salt amount leached by irrigation or flushing water; r is the salt amount carried away by surface runoff; g is the salt amount carried away by the horizontal outflow of soil water; u is the salt amount carried away with agricultural harvests.

[0058] Bayesian inference method based on prior and posterior information: It is a probability method developed based on Bayes' theorem for systematically expounding and solving stochastic problems. The most difficult part of Bayesian statistics is the integral difficulty encountered in seeking the posterior distribution. The integral related to the posterior distribution is very difficult to calculate numerically, especially in the high-dimensional case. The MCMC method provides an effective way to solve such problems, that is, the Monte Carlo method approximates the integral with the average value; the Markov Chain solves the sampling problem of samples.

[0059] The Metropolis-Hastings algorithm (M-H algorithm) is one of the most commonly used Markov chain Monte Carlo methods. This method samples from a general posterior distribution. The sampling strategy is to establish an irreducible and aperiodic Markov chain, and its stationary distribution is the target posterior distribution of interest. The core problem is to determine the rule for transferring from the current value to the next value, and its main task is to generate a Markov chain {x; i = 1,..., J} that satisfies the above regular conditions.

[0060] Example 1

[0061] The mathematical simulation research based on the mechanism of water-salt movement mainly conducts numerical simulations on the water-salt movement in the vadose zone and the non-vadose zone according to the principles of hydrodynamics, and solves the equations with the help of computer technology. The advantage of this method is that it can predict the distribution of groundwater, soil salinity, and soil moisture at any point, any depth, and any time period in the study area.

[0062] As a key factor restricting the salt content in the soil surface layer, the lowering of the groundwater level makes it possible for the surface soil to desalinate. The decrease in the salt content in the soil surface layer makes saline-alkali land reclaimable. The increase in input after reclamation and the differences in input for different land use types promote the nutrient balance to develop towards nutrient accumulation, and then the soil nutrients increase significantly and the nutrient contents of different land types vary. Although the soil fertility quality has been significantly improved and the soil quality has been greatly enhanced, the current bioclimatic conditions causing soil salinization have not changed. The change in the groundwater level still restricts the salt content in the soil surface layer, and soil potential salinization still exists. The recurrence of salinization, droughts and floods, and the instability of agricultural production still exist. Paying attention to the dynamic changes of obstacle factors such as soil potential salinization and realizing the prediction and forecast of soil salinization is not only beneficial to understanding the fertility characteristics, soil obstacles, and health status, but also beneficial to reshaping the soil ecological function of medium and low-yield farmland, improving the soil quality and sustainable production capacity of medium and low-yield farmland, meeting the development and utilization of medium and low-yield farmland soil resources, and is of great significance to the sustainable development of agriculture.

[0063] Based on this, combined with the mechanism of dynamic regulation of groundwater level and water-salt movement, it is necessary to formulate practical measures for preventing and controlling soil salinization. The groundwater depth, groundwater salinity, and soil conductivity (using the regression relationship between conductivity and soil salt content, the salt content in the soil at the sampling point can be obtained) are important indicators for quantitative simulation of soil water-salt movement. The determination of these indicators is of great significance for the dynamic simulation and prediction of soil water-salt movement in terms of time and space. In fact, the spatial pattern of soil salinity in the study area is closely related to the topographic and geomorphic features, soil types, and river distribution in this area. Generally, in areas with higher terrain, the groundwater depth is relatively deeper, and the probability of salinization threat is relatively smaller. Due to the variability of soil properties itself and the randomness of various factors affecting water-salt movement, such as tillage, management measures, planting systems, precipitation, evaporation, groundwater conditions, etc., the regional water-salt movement has the characteristics of randomness and complexity. Therefore, probability statistics methods are often used to study the prediction of water-salt movement and soil salinization.

[0064] In recent years, although probabilistic programming / Bayesian programming calculations have given Bayesian statistics a strong development momentum, there is no report on predicting soil salinization based on altitude using Bayesian inference.

[0065] This embodiment provides a method for predicting soil salinization based on altitude, as Figure 1 shown, including the following steps:

[0066] Step 1: Obtain the altitude x of the i-th area to be predicted i ;

[0067] Step 2: Based on the altitude x of the i-th area to be predicted i , use the logistic regression model to predict the binary soil salinization status y of the i-th area to be predicted i .

[0068] The construction process of the logistic regression model is as follows:

[0069] (1) Collect the altitudes x i and the binary soil salinization status y i of several sampling points, where i = 1,...., n, and n represents the number of sampling points, which can take the value of 100.

[0070] (2) Select the following logistic regression model:

[0071] y i ~B(1, π i ),

[0072] where B represents the binomial distribution, and πi Denotes the parameter of the binomial distribution, y i Denotes the binary status of soil salinization at the i-th sampling point (taking the value of 1 indicates the presence of soil salinization, and taking the value of 0 indicates the absence of salinization), x i Denotes the altitude of the i-th sampling point, β 0 And β 1 Are regression parameters.

[0073] The likelihood function of the logistic regression model is:

[0074]

[0075] Among them, Denotes the posterior distribution.

[0076] Considering that the prior distribution π of β 0 And β 1 Is an independent normal distribution, that is, π(β 0 , β 1 ) = π 1 (β 0 )π 2 (β 1 ), and j = 0, 1; when And Is very large, the prior distribution approaches the non-informative prior, where N is the normal distribution, Is the mean, Is the variance, representing the parameters of the normal distribution. Therefore, the posterior distribution is:

[0077]

[0078] Among them, σ 0 And σ 1 Are the parameters of the posterior distribution.

[0079] (3) Based on the altitudes x i And the binary status of soil salinization y i Of a number of collected sampling points, random numbers are generated from the posterior distribution π(β 0 , β 1 |y) through the component-by-component Metropolis-Hastings sampling method, and the logistic regression parameters are updated one by one according to the component in combination with the proposal distribution. The prior distribution is updated by selecting the likelihood function f(y|β 0 , β 1 ).

[0080] The proposal distribution is:

[0081]

[0082] where β′ = (β′ 0 , β′ 1 ); T and β = (β 0 , β 1 ); T .

[0083] Among them, is the parameter of the proposed distribution.

[0084] In the M-H algorithm, individual updates are performed component by component. Its advantage lies in its convenient application without the need to consider adjustment parameters.

[0085] The specific algorithm for generating random numbers from the posterior distribution π(β 0 , β 1 |y) through the component-by-component Metropolis-Hastings sampling method is as follows: for t = 1,..., T:

[0086] (301) Let That is, let the two components (logistic regression parameters) β 0 and β 1 in β obtained from the (t - 1)-th iteration be used as the two components β 0 and β 1 in β of the (t - 1)-th iteration;

[0087] (302) Generate a candidate point β′ from the first proposed distribution 0 ;

[0088] (303) Let β′ = (β′ 0 , β 1 ), and calculate the first acceptance probability α T (β, β′ 0 ) = min{1, A};

[0089] Among them, y is the binary soil salinization status vector of all sample points, π() represents the prior distribution, and f() represents the target distribution;

[0090] (304) Accept β = β′ with the first acceptance probability α 0 (β, β′ 0 ), that is, update the first component β 0 of β to β′ 0 0 ; otherwise, β remains unchanged;

[0091] (305) Generate a candidate point β′ from the second proposed distribution 1 ;

[0092] (306) Let β′ = (β 0 , β′ 1 ), T , the second acceptance probability α 1 (β, β′ 1 ) = min{1, B};

[0093] Among them,

[0094] (307) Accept β = β′ with the second acceptance probability α 1 (β, β′ 1 ), that is, update the second component β 1 of β to β′ 1 ; otherwise β remains unchanged;

[0095] (308) Let β (t) = β, and t = t + 1, return to (301) until t reaches the set value T, and output β.

[0096] Generating random numbers from the posterior distribution using the component-by-component M-H algorithm improves the mixing efficiency of the chain. In the M-H algorithm, updates are performed one by one by component, and its advantage lies in its convenient application without the need to consider adjustment parameters.

[0097] In this embodiment, f() is used to represent the target distribution, and g() is used to represent the proposal distribution. Suppose we hope to sample from the target distribution f(). The M-H algorithm starts from the initial value X 0 , and specifies a rule for transferring from the current value X t to the next value X t+1 , thereby generating a Markov chain {X t = 0, 1,...}.

[0098] Specifically, given the current value X t , a candidate point x′ is generated from a proposal distribution g(·|x). If the candidate point x′ is accepted, the chain transfers from the state x′ to the next moment t + 1 of the chain, and let X t+1 = x′; otherwise, the chain remains in the state X t , and let X t+1 = X t . Whether the candidate point x′ is accepted as the next value of the chain is determined according to the expression of the acceptance probability α(X t , x′) = min{1, A}, where:

[0099]

[0100] One point to note is that the density functions in the numerator and denominator of the above formula can be replaced by the "kernel of the density function" respectively, so the regularization constant factor in the density function can be omitted to simplify the calculation.

[0101] The iterative process of the Markov chain generated by the M-H algorithm that satisfies the regularity conditions is as follows:

[0102] (1) Select a proposal distribution g(·|X t );

[0103] (2) Generate an initial value x from the proposal distribution 0 ;

[0104] (3) Repeat steps (a)-(d) for t = 1, 2,...:

[0105] (a) Generate a candidate value x' from the proposal distribution g(·|X t );

[0106] (b) Generate a random number U from the uniform distribution U(0, 1);

[0107] (c) Calculate the acceptance probability. If U ≤ A, then accept x' and set X t+1 = x', otherwise set X t+1 = X t ;

[0108] (d) Increment t and return to (a).

[0109] The acceptance probability mentioned in the above steps is not necessarily the larger the better, because this may lead to slower convergence. When the parameter dimension is 1, the acceptance probability should be slightly less than 0.5 for optimality. When the dimension is greater than 5, the acceptance probability should be reduced to about 0.25. To show that the Markov chain generated by the M-H sampling method has a stationary distribution f, it can be shown that the transition kernel (or transition probability) of this chain and f both satisfy the balance equation. In addition, the choice of the proposal distribution g should satisfy, in addition to making the generated Markov chain satisfy the regularity conditions such as irreducibility, positive recurrence, aperiodicity, and having a stationary distribution f:

[0110] (1) The support set of the proposal distribution contains the support set of the target distribution;

[0111] (2) It is easy to sample from it, and it is often taken as a known distribution, such as the normal or t distribution, etc.;

[0112] (3) The proposal distribution should make the acceptance probability easy to calculate;

[0113] (4) The tail of the proposal distribution should be thicker than the tail of the target distribution;

[0114] (5) The frequency of new candidate points being rejected is not high.

[0115] Moreover, when the target distribution is the posterior distribution, the previous M-H algorithm can be directly implemented within the Bayesian framework. Simply replace \(x\) with the parameter \(\theta\) of interest and replace the target distribution \(f(x)\) with the posterior distribution \(\pi(\theta|x)\). The method is summarized as follows:

[0116] (1) Select a proposal distribution \(g(\cdot|\theta\) t )

[0117] (2) Generate an initial value \(\theta\) from the proposal distribution 0 ;

[0118] (3) Repeat steps (a)-(d) for \(t = 1, 2,\cdots\):

[0119] (a) Generate a candidate value \(\theta'\) from the proposal distribution \(g(\cdot|\theta\) t );

[0120] (b) Generate a random number \(U\) from the uniform distribution \(U(0, 1)\);

[0121] (c) Calculate the acceptance probability. If

[0122]

[0123] then accept \(\theta'\) and set \(\theta\) t+1 =\(\theta'\), otherwise set \(\theta\) t+1 =\(\theta\) t ;

[0124] (d) Increment \(t\) and return to (a).

[0125] In the Metropolis sampling method, the proposal distribution is symmetric, i.e., \(g(\cdot|X\) n ) satisfies \(g(X|Y)=g(Y|X)\). Therefore, the acceptance probability is:

[0126]

[0127] According to different choices of the proposal distribution \(g\), the componentwise Metropolis-Hastings sampling method is one of the variants of the Metropolis-Hastings sampling method and has a wide range of applications.

[0128] When the state space is multi-dimensional, instead of updating \(X_n\) as a whole, its components are updated one by one, which is called the componentwise M-H sampling method (componentwise Metropolis-Hastings sampler). This is more convenient and efficient.

[0129] Denote as: \(X_n=(X\) n,1 ,\cdots,X\) n,k ), \(X_{n,-i}=(X\)n,1 ,...., X n,i-1 ,...., X n,k ), respectively represent the state of the chain at the n-th step and the state of other components except the i-th component at the n-th step. f(x) = f(x 1 ,..., x k ) is the target distribution, and f(x i |x -i ) = f(x) / ∫f(x 1 ,..., x k )dx i represents the conditional density of X i with respect to other components.

[0130] The component-wise M-H sampling method consists of k steps: Let X n,i represent the state of the i-th component of Xn after the n-th iteration. Then, in the i-th step of the (n + 1)-th iteration, update Xn,i using the M-H algorithm as follows:

[0131] First, for i = 1,...., k, generate Y i (·|X n,i , X * n,-i ) from the i-th proposal distribution q i , where X * n,-i = (X n+1,1 ,....., X n+1,i-1 ,..., X n,k );

[0132] Then, with probability If Y i is accepted, then let X n+1,i = Y i ; otherwise, let X n+1,i = X n,i .

[0133] The two time trajectory plots of β 0 and β 1 in Figure 2(a) and Figure 2(b) show that the mixing of the chain is good. The plotted curve of the number of iterations and the sample values shows that the values in the trajectory plot appear within a certain range, without showing trends or periodic changes. The sample path plot stabilizes and cannot be distinguished from each other (mixed together), indicating that it has reached a convergence state and the mixing of the chain is good. Among them, Value of beta0 represents the value of the parameter β 0 ; Value of beta1 represents the value of the parameter β 1 ; Iterations represents the number of iterations.

[0134] The diagnosis of convergence characteristics is very important for MCMC simulations. Both the trace plot and the ergodic mean plot can be used to diagnose the convergence of MCMC. The trace plot is obtained by sampling the generated MCMC according to the number of iterations. To avoid the chain getting stuck in a local region of the target distribution, multiple Markov chains are usually generated simultaneously from different initial points. After running for a period of time, if their trace plots stabilize and have good mixing, with no recognizable patterns, including upward or downward curves, and the curves oscillate around a certain value, it is considered that the sampling has converged. The ergodic mean plot of a Markov chain is obtained by plotting the cumulative mean of the generated Markov chain against the number of iterations, which is used to determine whether the ergodic mean has converged. The ergodic mean of a Markov chain after reaching stationarity will tend to a horizontal line. Similarly, to avoid the chain getting stuck in a local region of the target distribution, the convergence of the ergodic means of multiple Markov chains starting from different initial points can also be examined. If only one Markov chain is used, a sufficient number of iterations is required to ensure that the chain can reach every part of the support. The middle two show the ergodic mean plots of β 0 and β 1 After removing the first 1000 iterations (i.e., the burn-in period), they tend to be stable, indicating good convergence of the chain.

[0135] In addition, the autocorrelation coefficient is also a relatively effective method for testing the convergence of a Markov chain. A smaller correlation coefficient indicates faster convergence. In the autocorrelation function plot of the soil salinization prediction chain, it shows that the sampling step size L of the Markov chain is 25 iterations, that is, one sample is taken every 25 samples, and the obtained sample autocorrelation is relatively low, indicating good convergence.

[0136] The diagnostic results of the three groups of figures, namely Figure 2(a) and Figure 2(b), Figure 3(a) and Figure 3(b), and Figure 4(a) and Figure 4(b), all show good convergence of the chain. After removing the burn-in period, the sample means of the two components β 0 and β 1 of β obtained from the chain are 1.8715 and -0.0611 respectively, and the sample standard deviations are 0.7862 and 0.0418 respectively. Among them, Lag represents lag, and Thin-25 iterations represents a sampling step size of 25 iterations. Convergence also refers to whether the samples generated by the algorithm reach the equilibrium distribution, that is, whether the samples are generated from the target distribution. In the autocorrelation function plot of the soil salinization prediction chain, it shows that the sampling step size L of the Markov chain is 25 iterations, that is, one sample is taken every 25 samples, and the obtained sample autocorrelation is relatively low, indicating good convergence.

[0137] The parameters β 0 and β 1The posterior distribution contains all the information and is also a better and complete estimate. Refer to Figure 5 and Figure 6 Soil salinization prediction β 0 and β 1 The posterior density of the parameters and the 99%, 95%, and 50% highest density regions. Among them, Densitydefault represents the density distribution under the default condition, N = 50000 represents the number of iterations is 50000 times, Bandwidth = 0.1417 represents the bandwidth is 0.1417, and Bandwidth = 0.006143 represents the bandwidth is 0.006143. In the case where the posterior distribution of the parameter is unimodal, there may be many groups of intervals (H α / 2 L α / 2 ) that satisfy the following conditions:

[0138]

[0139] The smallest interval (H α / 2 L α / 2 ) that satisfies the above conditions is called the highest density region (highest density region) or bayesian incredible interval of 1-α. For a multimodal posterior distribution, the highest density interval can be composed of more than one continuous interval.

[0140] The convergence status of the parameters and the high-density region of the posterior distribution of the parameters both indicate that the prediction of soil salinization can be achieved based on altitude.

[0141] Example 2

[0142] This embodiment provides a soil salinization prediction system based on altitude, which specifically includes the following modules:

[0143] A model construction module, which is configured to: based on the altitude and soil salinization binary status of a number of sampling points collected, generate random numbers from the posterior distribution through the component-by-component Metropolis-Hastings sampling method, and update the logistic regression parameters one by one according to the component in combination with the proposal distribution, and update the prior distribution through the likelihood function;

[0144] A data acquisition module, which is configured to: acquire the altitude of the area to be predicted;

[0145] A prediction module, which is configured to: based on the altitude, use the logistic regression model to predict the soil salinization binary status of the area to be predicted.

[0146] Among them, the prior distribution is an independent normal distribution.

[0147] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1 one by one, and the specific implementation process is the same, so it will not be repeated here.

[0148] Embodiment 3

[0149] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a method for predicting soil salinization based on altitude as described in Embodiment 1 above.

[0150] Embodiment 4

[0151] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a method for predicting soil salinization based on altitude as described in Embodiment 1 above.

[0152] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0153] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.

[0156] Those of ordinary skill in the art will understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0157] The foregoing is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for predicting soil salinization based on altitude, characterized in that, it includes: Obtaining the altitude of the area to be predicted; Based on the altitude, using a logistic regression model to predict the binary status of soil salinization in the area to be predicted; Among them, the construction method of the logistic regression model is: based on the altitudes and binary soil salinization statuses of a number of sampling points collected, generating random numbers from the posterior distribution by the component-by-component Metropolis-Hastings sampling method, and updating the logistic regression parameters one by one according to the component in combination with the proposal distribution, and updating the prior distribution through the likelihood function; the proposal distribution makes the generated Markov chain satisfy irreducibility, positive recurrence, and aperiodicity; The specific method for generating random numbers from the posterior distribution is: (301) Let , and be the logistic regression parameters obtained in the (t - 1)-th iteration; (302) Generate candidate points from the first proposal distribution ; (303) Let = , calculate the first acceptance probability; the first acceptance probability is: = min{1, A}; where A = = , , and are logistic regression parameters, generates candidate points for the first proposal distribution, is the binary status vector of soil salinization for all sample points, π( ) represents the prior distribution, and f( ) represents the target distribution; (304) Accept with the first acceptance probability ; Otherwise Remain unchanged; (305) Generate candidate points from the second proposal distribution ; (306) Let = , calculate the second acceptance probability; the second acceptance probability is: =min{1,B}; where B = = , , and are logistic regression parameters, is the candidate point generated by the second proposal distribution, is the binary soil salinization status vector of all sample points, π( ) represents the prior distribution, and f( ) represents the target distribution (307) Accept with the second acceptance probability ; Otherwise Remain unchanged; (308) Let t = t + 1, return to (301) until t reaches the set value, and output .

2. The method for predicting soil salinization based on altitude according to claim 1, characterized in that, the prior distribution is an independent normal distribution.

3. A system for predicting soil salinization based on altitude, characterized in that, it includes: A model construction module, which is configured to: based on the altitudes and binary soil salinization statuses of a number of sampling points collected, generate random numbers from the posterior distribution by the component-by-component Metropolis-Hastings sampling method, and update the logistic regression parameters one by one according to the component in combination with the proposal distribution, and update the prior distribution through the likelihood function; A data acquisition module, which is configured to: obtain the altitude of the area to be predicted; A prediction module, which is configured to: based on the altitude, use a logistic regression model to predict the binary status of soil salinization in the area to be predicted; the proposal distribution makes the generated Markov chain satisfy irreducibility, positive recurrence, and aperiodicity; The specific method for generating random numbers from the posterior distribution is: (301) Let , and be the logistic regression parameters obtained in the (t - 1)-th iteration; (302) Generate candidate points from the first proposal distribution ; (303) Let = , calculate the first acceptance probability; the first acceptance probability is: =min{1,A}; where A = = , , and are logistic regression parameters, is the candidate point generated by the first proposal distribution, is the binary soil salinization status vector of all sample points, π( ) represents the prior distribution, and f( ) represents the target distribution; (304) Accept with the first acceptance probability ; Otherwise Remain unchanged; (305) Generate candidate points from the second proposal distribution ; (306) Let = , calculate the second acceptance probability; the second acceptance probability is: = min{1, B}; where B = = , , and are logistic regression parameters, is the candidate point generated by the second proposal distribution, is the binary status vector of soil salinization for all sample points, π( ) represents the prior distribution, and f( ) represents the target distribution (307) Accept with the second acceptance probability ; Otherwise Remain unchanged; (308) Let t = t + 1, return to (301) until t reaches the set value, and output 。 4. The system for predicting soil salinization based on altitude according to claim 3, characterized in that, the prior distribution is an independent normal distribution.

5. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the steps in the method for predicting soil salinization based on altitude according to any one of claims 1-2.

6. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the steps in the method for predicting soil salinization based on altitude according to any one of claims 1-2.

Citation Information

Patent Citations

  • Soil salinity degradation estimation by regression algorithm using agricultural internet of things

    AU2020102098A4

  • Chrysanthemum salt-tolerance associated molecular marker and obtaining method and application thereof

    CN105112403A