A soil ph evolution element modeling method and device based on proton budget balance mechanism guided machine learning
By combining the proton budget balance mechanism and machine learning methods, the problems of high spatiotemporal resolution and mechanism interpretability in regional-scale soil pH history reconstruction were solved, and efficient and accurate prediction of large-scale soil pH spatiotemporal evolution was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to simultaneously achieve high spatiotemporal resolution and mechanistic interpretability in soil pH history reconstruction at the regional scale, especially in areas or periods with scarce data where model generalization ability is insufficient, leading to uncertainty in long-term prediction results.
A VSD+ model based on the proton budget balance mechanism was combined with machine learning methods. A random forest model was trained using the pH value of soil samples and multidimensional environmental variables to reconstruct the spatiotemporal variation characteristics of soil pH. The MetHyd and RothC modules were used to calculate the proton generation and buffering mechanisms in the soil acidification process. Key parameters were calibrated using the Markov chain Monte Carlo method.
It achieves high-resolution spatiotemporal evolution distribution of soil pH over large-scale regions, breaks through the limitations of model prediction on soil sample points, and improves the prediction efficiency and accuracy of soil pH spatiotemporal evolution sequence.
Smart Images

Figure CN122491013A_ABST
Abstract
Description
Technical Field
[0001] This invention pertains to methods for spatiotemporal prediction of soil pH, and particularly relates to a method and apparatus for modeling soil pH evolution based on machine learning guided by the proton balance mechanism. Background Technology
[0002] Arable land is the foundation of agricultural production, playing a vital role in ensuring food supply, maintaining the stability of agricultural ecosystems, and supporting regional economic development. However, long-term excessive application of chemical fertilizers, continuously increasing atmospheric nitrogen deposition, and intensive farming practices have led to the continuous accumulation of acidic substances in the soil, resulting in a decrease in soil pH and inducing problems such as nutrient imbalance, activation of harmful metals, changes in soil microbial community structure, and restricted crop growth, adversely affecting sustainable agricultural production and farmland ecological security. Therefore, accurately characterizing and analyzing the long-term evolution of soil pH at a regional scale is of great significance for soil acidification risk assessment, early warning, and the formulation of control measures.
[0003] Currently, the reconstruction of soil pH history at the regional scale still faces the challenge of achieving both "mechanistic depth" and "spatial accuracy." On the one hand, simulation methods based on process models (such as the VSD model and its extended versions) can better characterize key biogeochemical mechanisms such as proton production, consumption, migration, and neutralization during acidification, based on physicochemical principles such as mass conservation and chemical equilibrium, and have clear mechanistic interpretability.
[0004] Patent application CN116882625A discloses a method for identifying the spatial vulnerability of rural settlements in karst mountainous areas. The method includes: acquiring zoning data of rural settlements in karst mountainous areas; preprocessing the zoning data; constructing a Virtual Segmentation Model (VSD) based on the preprocessed zoning data; constructing a spatial vulnerability evaluation index system based on the VSD model; analyzing and calculating the spatial vulnerability evaluation index system using a geographic detector model and a global autocorrelation model combined with a natural discontinuity grading method; obtaining spatial analysis and identification results of the spatial vulnerability of rural settlements in karst mountainous areas; and obtaining the spatial distribution characteristics of the vulnerability of rural settlements in karst mountainous areas and the interaction mechanism of its influencing factors based on the spatial analysis and identification results. This invention provides important technical basis for measuring the main influencing factors and interactions between factors of spatial vulnerability in rural settlements in karst mountainous areas.
[0005] However, when the VSD model mentioned in the above patent application is applied to large-scale regional simulation, the model is usually difficult to directly generate high spatiotemporal resolution soil pH evolution maps due to the limited spatial availability and heterogeneity characterization ability of the key input parameters required by the model (such as mineral weathering rate, cation exchange characteristics, nutrient cycling flux, etc.), which limits its application in regional precision management.
[0006] On the other hand, data-driven spatial statistics and machine learning mapping methods can efficiently integrate multi-source remote sensing data, environmental covariates, and historical survey samples to generate high-spatial-resolution soil property distribution maps through complex nonlinear relationship learning. Although this method performs excellently in static mapping, the model heavily relies on the spatiotemporal coverage of training samples when reconstructing long-term historical evolution. However, measured soil samples from historical periods are often scarce and unevenly distributed, leading to uncertainty in the model's generalization ability in data-scarce periods or regions. Furthermore, purely data-driven models lack constraints on the basic physicochemical processes of soil pH evolution, exhibiting "black box" characteristics, weak mechanistic interpretability, and affecting the reliability of long-term prediction results.
[0007] Therefore, there is an urgent need to develop a new method that integrates the mechanistic characteristics of process models with the advantages of spatial representation in machine learning, in order to accurately reveal the long-term spatiotemporal evolution of regional soil pH. Summary of the Invention
[0008] This invention provides a meta-modeling method for soil pH evolution based on proton budget balance mechanism-guided machine learning. This method can reconstruct the soil pH evolution process based on a mechanistic framework when soil pH time series data is missing, and combine machine learning to achieve spatial extrapolation, thereby accurately characterizing the spatiotemporal variation features of large-scale farmland soil pH.
[0009] This invention provides a meta-modeling method for soil pH evolution based on proton balance mechanism-guided machine learning, comprising: Historical or current pH values, initial soil properties, annual climate data, and agricultural management data of different soil samples in the test area are collected. Then, standardized preprocessing is performed. The preprocessed initial soil properties and annual climate data are input into the MetHyd model (Meteorological and Hydrological Model) to obtain nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus. The preprocessed initial soil properties, annual climate data, and agricultural management data, as well as the nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, are input into the VSD+ (Very Simple Dynamic Model Plus) model to obtain the soil pH evolution sequence of different soil samples in the test area within a set time period. The soil pH evolution sequence within a set time period of the area to be tested is used as a label, and a sample set is constructed with the corresponding multidimensional environmental variable set. The machine learning model is trained through the sample set to obtain the soil pH evolution model. When applied, the environmental variables of the area to be tested are input into the soil pH evolution model to obtain the annual spatiotemporal evolution distribution dataset of soil pH in the area to be tested.
[0010] Preferably, the pre-processed initial soil properties and annual climate data are input into the MetHyd model to obtain nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, including: Based on initial soil properties and annual climate data, the environmental reduction function and water flux were calculated using the MetHyd model. The environmental reduction function is a rate reduction factor (usually a dimensionless function between 0 and 1) for soil nitrogen transformation processes (such as nitrification, denitrification, and mineralization) relative to the optimum conditions under actual hydrothermal conditions, determined by both temperature and soil moisture content (or relative water saturation) functions. Water flux refers to the vertical migration of water in the soil profile per unit time (mainly reflected in precipitation surplus or leaching), calculated from the difference between precipitation and evapotranspiration combined with soil water holding capacity parameters. These variables were input as external drivers into the VSD+ model.
[0011] Preferably, the initial soil properties include cation exchange capacity, soil bulk density, organic carbon content, clay content, base saturation, and calcium carbonate content; The annual climate data refers to daily precipitation, average temperature, and solar radiation data for a set number of years.
[0012] Preferably, the pre-processed initial soil properties, annual climate data, and agricultural management data, along with nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, are input into the VSD+ model to obtain the soil pH evolution sequence of different soil samples in the test area over a set time period, including: The VSD+ model includes the RothC (Rothamsted Carbon Model) module, which incorporates soil physicochemical properties (initial organic carbon content, initial clay content, bulk density, etc.), dynamics of exogenous carbon and nitrogen inputs (such as crop residues and root inputs), and climate data (temperature, precipitation / evapotranspiration) as input parameters. This module divides soil organic carbon into five independent carbon pools and sets corresponding C / N stoichiometry and decomposition kinetic parameters for each pool. It simulates processes such as organic carbon mineralization, nitrogen mineralization, and nitrification, and, combined with hydrological leaching processes, characterizes the dynamic changes in proton production, alkaline cation release, and migration within the soil system, providing key flux inputs for subsequent soil solution chemical equilibrium calculations. To characterize the multi-stage buffering system in farmland soil acidification, the VSD+ model further comprehensively describes three types of buffering mechanisms under different pH ranges: First, based on the initial cation exchange capacity (CEC), base saturation (BS), and soil solution ion concentration, the Gapon equation is used to dynamically solve the rapid equilibrium of exchangeable cations between the solid and liquid phases to characterize the short-term ion exchange buffering process; second, the mineral weathering rate, constrained by both parent material and climatic conditions, is introduced to characterize the long-term acid neutralization capacity of the soil; in the strongly acidic stage, the dissolution-precipitation equilibrium mechanism of gibbsite is coupled, and the inherent thermodynamic solubility product (Ksp) is used as a constraint to describe the Al content in the soil solution. 3+ With OH - The equilibrium relationship was used to simulate the activation and release of aluminum under low pH conditions and its effect on excess protons (H+). + The model integrates carbon and nitrogen transformation, ion exchange, mineral weathering, aluminum buffering, and hydrological leaching processes on an annual scale, and iteratively solves the H value using the soil solution charge balance equation. + The concentration was determined, and the temporal evolution of soil pH was calculated accordingly.
[0013] Preferably, the agricultural management data is calculated and then used as input to the VSD+ model. The calculation of the agricultural management data includes calculating the flux of external nutrient input, the flux of crop nutrient removal, and the carbon and nitrogen flux of straw returning to the field.
[0014] Further preferably, the calculation of the exogenous nutrient input flux includes: comprehensive calculation of atmospheric deposition, biological nitrogen fixation, agricultural fertilizer input and ammonia volatilization loss. The agricultural fertilizer input calculation is obtained by correcting for ammonia volatilization loss based on nitrogen application intensity, soil clay content, temperature and fertilizer type coefficient to determine the net acid-causing nitrogen input.
[0015] Further preferred methods include calculating crop nutrient removal fluxes, including: based on crop planting patterns and grain yield statistics, combined with the straw-to-grain ratio, element content characteristics, and straw return rate of different crops, calculating the net flux of phosphorus, sulfur, basic cations, and nitrogen removed from the farmland system with crop harvest.
[0016] Further optimization involves calculating the carbon and nitrogen fluxes of straw returning to the field, including: quantitatively calculating the organic carbon and total nitrogen fluxes entering the soil mineralization cycle through straw returning to the field based on the straw returning ratio, crop moisture content, root-to-shoot ratio, and biomass characteristics.
[0017] This invention uses the quantitative calculation of organic carbon and total nitrogen fluxes entering the soil mineralization cycle through straw return as a key driver of soil background property evolution.
[0018] Preferably, the pH values of different historical or current soil samples from the area to be tested are used as observations, and the Markov Chain Monte Carlo (MCMC) method is used to perform Bayesian calibration on key sensitive parameters. Specifically, the logarithmic values of the aluminum hydrolysis equilibrium constant, the empirical index of aluminum dissolution, the logarithmic values of the aluminum-base cation exchange selectivity coefficient, and the logarithmic values of the hydrogen-base cation exchange selectivity coefficient can be set with prior distributions based on empirical ranges in the literature; the mineral weathering rate parameter is set with prior ranges based on factors affecting weathering intensity such as parent material type, soil mineral composition, soil layer thickness, soil texture, temperature, and precipitation / runoff, referring to regional empirical values or weathering model results. Subsequently, in each MCMC sampling, the candidate parameter combination is input into the VSD+ model to simulate the soil pH of the corresponding sample point and year, and a likelihood function is constructed using the residuals between the simulated value and the measured pH. The posterior distribution of the parameters is updated step by step to obtain the optimal parameter combination and its uncertainty range, so that the root mean square error between the simulated soil pH evolution sequence and the observed pH value does not exceed a set threshold.
[0019] Preferably, the multidimensional environmental variable set includes climate variables, topographic variables, soil background attribute variables, agricultural management variables, and remote sensing features, wherein: The climate variables include one or more of the following: total annual precipitation, average annual temperature, total annual solar radiation, wind speed, relative humidity, nitrogen deposition, and sulfur deposition; The terrain variables include one or more of the following: elevation, slope, curvature, and terrain humidity index extracted based on a digital elevation model (DEM); The soil background attribute variables include one or more of the following: soil texture, bulk density, cation exchange capacity, initial base saturation, and bedrock thickness. The agricultural management variables include one or more of the following: fertilizer application rate, irrigation amount, and crop planting pattern; The remote sensing image features include: Sentinel-1 radar data VV and VH polarizations and their ratios.
[0020] Preferably, a soil pH evolution model is obtained by training a machine learning model using a sample set, including: Using the soil pH evolution sequence of the area to be tested within a set time period as a label, combined with a multidimensional set of environmental variables, a random forest machine learning model is trained. Hyperparameter optimization was performed using grid search, with root mean square error (RMSE) and coefficient of determination (R²) used as the parameters. 2 Accuracy verification was conducted to screen and determine soil pH evolution models.
[0021] On the other hand, the present invention also provides a soil pH evolution meta-modeling device based on proton balance mechanism-guided machine learning, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the soil pH evolution meta-modeling method based on proton balance mechanism-guided machine learning.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention uses the VSD+ model to obtain soil pH evolution sequences from multiple representative soil sample points in the test area. The obtained soil pH evolution sequences are used as labels, and the machine learning model trained with the corresponding environmental variables can directly and accurately predict the pH evolution sequence values of other soil sample points. This results in a high-resolution annual soil pH spatiotemporal evolution distribution dataset for a large area, breaking through the limitations of the VSD+ model in predicting soil sample points and improving the prediction efficiency of large-area soil pH spatiotemporal evolution sequences. Attached Figure Description
[0023] Figure 1 A flowchart of a soil pH evolution meta-modeling method based on proton budget balance mechanism-guided machine learning provided in this embodiment of the invention; Figure 2 The scatter plot showing the correlation between measured and predicted soil pH values provided in this embodiment of the invention represents the pH prediction accuracy obtained by the VSD+ process model and the random forest model. Figure 3 This is a spatiotemporal distribution map of pH in the 0-20 cm surface soil layer of Chengmai County, Hainan Province, obtained according to an embodiment of the present invention. Detailed Implementation
[0024] The present invention will be further described and illustrated below with reference to the accompanying drawings and specific embodiments.
[0025] This invention provides a meta-modeling method for soil pH evolution based on proton budget balance mechanism-guided machine learning. By constructing this method, a process model is used to reconstruct the long-term soil pH evolution trajectory. Combined with a machine learning model, high-resolution spatiotemporal mapping of regional-scale soil pH is achieved. The implementation process is as follows: Figure 1 As shown, it includes: (1) Soil sample collection, screening, and physicochemical analysis, along with initial soil properties and meteorological data acquisition: In order to characterize the spatiotemporal variability of cultivated land soil properties, historical soil testing and fertilizer recommendation data were integrated with county-level field survey data conducted in the current year, specifically including: The specific operation of soil sampling and collection provided in this invention is as follows: For historical soil testing and fertilizer recommendation data, firstly, based on spatial analysis technology, the coordinates of historical sampling points are overlaid with the regional land use status map for analysis. Sampling points for the main cultivated land types of paddy fields and dry land are retained, while records from non-agricultural production areas such as forest land, grassland, and construction land are removed to reduce the interference of non-agricultural land on soil acidification trend analysis. Secondly, samples with incomplete data or obvious anomalies are removed, ultimately constructing a high-quality historical farmland soil pH dataset covering the main agricultural areas of the county. For the soil status survey of that year, a stratified sampling design is adopted, stratifying the data based on topography and soil type, focusing on paddy field and dry land farmland ecosystems. Sampling points are evenly distributed throughout the county to obtain 0-20 cm surface soil samples, ensuring good representativeness of the spatial distribution of the samples.
[0026] The specific operation of the soil standardization sample pretreatment process provided in this invention is as follows: First, the soil sample is naturally air-dried at room temperature, during which plant roots, gravel, insects, and other non-soil impurities are manually removed. Then, the soil sample is ground and passed through a 2 mm (10 mesh) nylon sieve. Finally, the sieved, air-dried soil sample is weighed, and deionized water is added at a soil-to-water mass-volume ratio of 1:2.5. The mixture is thoroughly shaken to ensure uniform dispersion of soil particles. After settling and equilibration, the pH value of the soil-water suspension is measured using a calibrated glass electrode pH meter.
[0027] The specific operations for initial soil properties, meteorological driving forces, and agricultural data acquisition provided in this invention are as follows: To construct a simulation of soil pH evolution starting in the 1980s, initial soil properties of sample points are obtained based on the Comprehensive World Soil Database (HWSD v1.2), including cation exchange capacity, soil bulk density, organic carbon content, clay content, base saturation, and calcium carbonate content. Meteorological data are obtained from the National Qinghai-Tibet Plateau Scientific Data Center, selecting daily precipitation, average temperature, and solar radiation data from 1980 to 2020, and resampled from 1 km resolution to a monthly scale sequence using a Python script. To accurately characterize the water conditions of intensive farmland, the monthly average farmland irrigation water volume is extracted based on the China Monthly Water Consumption Raster Dataset (HSWUD) from 1965 to 2022, and superimposed with the monthly precipitation during the same period to construct a total precipitation sequence. Finally, the preprocessed annual and monthly total precipitation, monthly average temperature, and monthly average solar radiation data are input into the MetHyd sub-model to calculate and obtain parameters such as mineralization, nitrification, and denitrification reduction coefficients and precipitation surplus as inputs to the VSD+ model.
[0028] (2) Obtain the soil pH evolution sequence through the VSD+ model: Input the pre-processed initial soil properties and annual climate data into the MetHyd model (Meteorological and Hydrological Model) to obtain the nitrification, denitrification and mineralization reduction coefficients and precipitation surplus. Input the pre-processed initial soil properties, annual climate data and agricultural management data, as well as the nitrification, denitrification and mineralization reduction coefficients and precipitation surplus into the VSD+ (Very Simple Dynamic Model Plus) model to obtain the soil pH evolution sequence of different soil samples in the test area within a set time period.
[0029] In one specific embodiment, the pre-processed initial soil properties and annual climate data are input into the MetHyd model to obtain nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, including: Based on initial soil properties and annual climate data, the environmental reduction function and water flux are calculated using the MetHyd model. The environmental reduction function is a rate reduction factor (usually a dimensionless function between 0 and 1) for soil nitrogen transformation processes (such as nitrification, denitrification, and mineralization) relative to the optimum conditions under actual hydrothermal conditions, determined by both temperature and soil moisture content (or relative water saturation) functions. Water flux refers to the vertical migration of water in the soil profile per unit time (mainly reflected in precipitation surplus or leaching), calculated from the difference between precipitation and evapotranspiration combined with soil water holding capacity parameters. These variables are input as external drivers into the VSD+ model. In the specific implementation, the preprocessed initial soil properties and annual / monthly climate data are first input into the MetHyd sub-model. The simulation year is set, and the calculation is run to obtain the annual nitrification reduction coefficient, denitrification reduction coefficient, mineralization reduction coefficient, and precipitation surplus parameter.
[0030] The initial soil properties provided in the specific embodiments of the present invention include cation exchange capacity, soil bulk density, organic carbon content, clay content, base saturation and calcium carbonate content; the annual climate data are daily precipitation, average temperature and solar radiation data for a set number of years.
[0031] In one specific embodiment, the preprocessed initial soil properties, annual climate data, and agricultural management data, along with nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, are input into the VSD+ model to obtain the soil pH evolution sequence of different soil samples in the test area over a set time period, including: The VSD+ model includes the RothC (Rothamsted Carbon Model) module, which incorporates soil physicochemical properties (initial organic carbon content, initial clay content, bulk density, etc.), dynamics of exogenous carbon and nitrogen inputs (such as crop residues and root inputs), and climate data (temperature, precipitation / evapotranspiration) as input parameters. This module divides soil organic carbon into five independent carbon pools and sets corresponding carbon-to-nitrogen ratios (C / N) stoichiometry and decomposition kinetic parameters for each pool. This is used to simulate processes such as organic carbon mineralization, nitrogen mineralization, and nitrification. Furthermore, it combines hydrological leaching processes to characterize the dynamic changes in proton production, alkaline cation release, and migration within the soil system, providing key flux inputs for subsequent soil solution chemical equilibrium calculations. To characterize the multi-stage buffering system in the process of farmland soil acidification, the VSD+ model further comprehensively describes three types of buffering mechanisms under different pH ranges: First, based on the initial cation exchange capacity (CEC), base saturation (BS), and soil solution ion concentration, the Gapon equation is used to dynamically solve the rapid equilibrium of exchangeable cations between the solid and liquid phases to characterize the short-term ion exchange buffering process; second, the mineral weathering rate, which is constrained by both parent material and climatic conditions, is introduced to characterize the long-term acid neutralization capacity of the soil; in the strongly acidic stage, the dissolution-precipitation equilibrium mechanism of gibbsite is coupled, using the inherent thermodynamic solubility product (K0)... sp Using ) as a constraint, describe the Al in the soil solution 3+ With OH - The equilibrium relationship was used to simulate the activation and release of aluminum under low pH conditions and its effect on excess protons (H+). + The model integrates carbon and nitrogen transformation, ion exchange, mineral weathering, aluminum buffering, and hydrological leaching processes on an annual scale, and iteratively solves the H value using the soil solution charge balance equation. + The concentration was determined, and the temporal evolution of soil pH was calculated accordingly.
[0032] In a specific embodiment of the present invention, agricultural management data is calculated and used as input to the VSD+ model. The calculation of the agricultural management data includes the calculation of external nutrient input flux, crop nutrient removal flux, and straw return carbon and nitrogen flux.
[0033] In one specific embodiment, the input flux of the exogenous nutrient is calculated, including: the comprehensive calculation and quantification of atmospheric deposition, biological nitrogen fixation and agricultural fertilizer input. The agricultural fertilizer input calculation is obtained by correcting for ammonia volatilization loss based on nitrogen application intensity, soil clay content, temperature and fertilizer type coefficient to determine the net acid-causing nitrogen input.
[0034] Specifically, atmospheric deposition is quantified, including: extracting NO3 based on a national-scale atmospheric deposition dataset with a 5 km resolution. - NH4+ Potential acid-causing precursors such as SO2 and Ca 2+ Na + Mg 2+ Annual sedimentation flux of iso-acid buffers, including extracted SO2 and Ca. 2+ K + Na + Mg 2+ The values are directly input into the VSD+ model.
[0035] Specifically, biological nitrogen fixation flux was quantified, including calculations based on crop type, with a target of 25 kg N·m·h for rice and non-leguminous dryland crops. -1 and 5 kg N ha -1 The non-symbiotic nitrogen fixation rate is input as the amount of nitrogen fixed into the VSD+ model.
[0036] Specifically, agricultural fertilizer inputs and ammonia volatilization losses were quantified, including: extracting application rates of 13 fertilizer types based on a global crop-specific nitrogen fertilizer dataset from 1961 to 2020, and subtracting ammonia volatilization losses from the annual fertilizer application rate for each grid cell. The ammonia volatilization loss rate was calculated using an empirical equation: Vol NH3 = exp (1.612+0.007 ×N rate +0.043 ×clay +0.038 × Tem + coff ),in N rate Nitrogen application intensity, clay Soil clay content, Tem For temperature, coff This is the fertilizer type coefficient. Therefore, the total nitrogen input for VSD+ is set to atmospheric NO3. - NH4 + The sum of sedimentation and nitrogen fertilizer application rate minus ammonia volatilization; total K + The input setting is atmospheric deposition K. + The sum of potassium fertilizer; Cl - The input is set to be equivalent to the amount of potassium fertilizer applied (assuming it is all KCl).
[0037] Specifically, regarding the quantification of element fluxes caused by crop harvesting and straw return to the field, this embodiment uses a 10 m resolution crop distribution map generated from Sentinel-2 imagery to identify major planting systems such as "double-cropping rice" and "water-dryland rotation." Combined with county-level crop yield data from statistical yearbooks, a rasterization algorithm is used to allocate these values to each pixel. The amount of phosphorus, sulfur, and basic cations absorbed by the crops is also considered. T ps Defined as: T ps =Y c ×P c +Y c ×K jc ×(1-K r )×P j , in Y c Indicates grain yield; K jc The ratio of grass to grain; K r Represents the rate of straw return to the field; P z and P j The values represent the nutrient concentrations of the corresponding elements in the grain and straw, respectively. The calculated amounts of phosphorus, sulfur, and basic cations absorbed by the crop are then input into the VSD+ model.
[0038] At the same time, calculate the total nitrogen uptake of crops. T N Nitrogen returned by straw R N and return of carbon R C : T N =Y c ×N c +Y c ×K jc ×N j R N =Y c ×K jc ×K r ×C x R C =Y c × ( 1-WC ) ×SG ratio × ( RSratio +RR ) × 0.45 in, N c and N j These represent the nitrogen concentrations in grains and straw, respectively. C x This represents the concentration of the corresponding element in the straw. WC Moisture content, SG ratio For grass and grain ratio, RS ratio For the root-to-crown ratio, RR The percentage of straw returned to the field is 0.45, and the conversion coefficient from biomass to carbon content is 0.45. The calculated total crop nitrogen uptake, straw nitrogen return, and straw carbon return are input into the VSD+ model.
[0039] In a specific embodiment of this invention, the pH values of different historical or current soil samples from the area to be tested are collected as observations. The Markov chain Monte Carlo method is used to perform Bayesian calibration on key sensitive parameters, including: The logarithmic values of the aluminum hydrolysis equilibrium constant, the empirical index of aluminum dissolution, the logarithmic values of the aluminum-base cation exchange selectivity coefficient, and the logarithmic values of the hydrogen-base cation exchange selectivity coefficient were set according to empirical ranges in the literature. The mineral weathering rate parameters were combined with factors affecting weathering intensity such as parent material type, soil mineral composition, soil layer thickness, soil texture, temperature, and precipitation / runoff, and the prior ranges were set with reference to regional empirical values or weathering model results. Subsequently, in each MCMC sampling, the candidate parameter combination with the set prior distribution and prior range was input into the VSD+ model to simulate the soil pH of the corresponding soil sample points and years. The likelihood function was constructed using the residuals between the simulated values and the measured pH, and the posterior distribution of the parameters was updated step by step. Finally, the optimal parameter combination and its uncertainty range were obtained, so that the root mean square error between the simulated soil pH evolution sequence and the observed pH value was not higher than the set threshold.
[0040] Specifically, regarding the parameterization and calibration of geochemical processes, the logarithm of the aluminum hydrolysis equilibrium constant (log K) Al-OH The initial value is set to 8.5, the initial value of the aluminum dissolution empirical index (expAl) is set to 3, and the logarithm of the aluminum-base cation exchange selectivity coefficient (log K) is... Al-BC The logarithm of the hydrogen-base cation exchange selectivity coefficient (log K) H-BC The initial value was set to 1. Using historical soil testing and fertilizer recommendation data points and current year's measured pH data as observations, the Markov chain Monte Carlo method was employed to analyze log K. Al-OH log K Al-BC log KH-BC Parameters such as expAl and mineral weathering rate are subjected to Bayesian calibration and posterior distribution sampling to ensure that the root mean square error of the pH values collected from the simulated soil pH evolution sequence does not exceed the set threshold.
[0041] (3) Construction of multidimensional environmental variable sets and machine learning modeling: The specific operation for constructing the multidimensional environmental variable set provided in the specific embodiment of the present invention is as follows: the soil pH evolution sequence with consistent mechanism and continuous time generated by the VSD+ model is used as the target variable, and combined with multi-source environmental covariates of climate, topography, agricultural management and soil properties to construct a machine learning model.
[0042] Specifically, in order to characterize the regional hydro-climatic features and their regulatory role on biogeochemical processes, the climate covariates selected included climate indicators such as precipitation, average temperature, maximum temperature, minimum temperature, total solar radiation, wind speed, relative humidity, potential evapotranspiration, and air pressure.
[0043] Specifically, atmospheric deposition variables were included to characterize natural acid input, including nitrogen oxide deposition, sulfur dioxide deposition, and ammonia deposition data.
[0044] Specifically, a series of terrain prediction factors are derived from the digital elevation model, including: elevation, aspect, plane curvature, profile curvature, slope length and slope factor, terrain humidity index, terrain location index, terrain ruggedness index, specific catchment area, total catchment area, multi-resolution valley bottom flatness index, multi-resolution ridge top flatness index, and vertical distance to river channel or valley depth related index.
[0045] Specifically, agricultural management factors were incorporated to characterize the impact of human activities on soil pH, including the application rates of inorganic nitrogen fertilizer, organic nitrogen fertilizer, and phosphate fertilizer, as well as irrigation water usage.
[0046] Specifically, soil baseline attribute factors are incorporated, derived from the Integrated World Soil Database (HWSD) and the High-Resolution National Soil Information Grid, including: soil texture components (sand, silt, clay content, texture classification, bulk density, soil type, soil layer thickness, coarse debris content, carbon pool, soil organic carbon, carbon-nitrogen ratio, cation exchange capacity, initial base saturation, parent material calcium content, bedrock depth, as well as total nitrogen, total phosphorus, total potassium and their corresponding density indices.
[0047] Specifically, to characterize the spatial variability of surface conditions, Sentinel-1 radar remote sensing variables, including VV and VH backscattering coefficients and their ratios, were incorporated. In addition, soil moisture and soil temperature were included as state variables to capture the combined effects of seasonal hydrothermal conditions and tillage on the acidification process.
[0048] Specifically, to ensure consistent spatial scale and grid alignment across different data sources, all covariates are unified to a common coordinate reference system and resampled to a consistent spatial resolution. For continuous variables such as temperature and elevation, bilinear interpolation is used for resampling; for categorical variables such as soil type and texture classification, nearest neighbor assignment is used for resampling.
[0049] The specific machine learning modeling operation provided in the specific embodiment of the present invention is as follows: Random forest is selected as the training model, and grid search combined with 10-fold cross-validation is used to optimize the hyperparameters, with R... 2 As a preferred criterion.
[0050] Specifically, the hyperparameter search space of random forest is set as follows: number of trees: 100 to 500; number of candidate features evaluated when splitting a node: 1 to 20; maximum tree depth: 10 to 40; minimum number of samples per leaf node: 1 to 10.
[0051] This embodiment uses the root mean square error (RMSE) and the coefficient of determination (R²). 2 RMSE is used to evaluate model performance. RMSE quantifies the average magnitude of the prediction error in units of the response variable, while R0... 2 It measures the proportion of the variance of the target variable explained by the model. The calculation formula is as follows: in, n The number of sample data. and For the first i Observed and predicted values for each sample. It is the average of all observations.
[0052] (4) Construction of a dataset on the spatiotemporal evolution of soil pH The specific operation for constructing the soil pH spatiotemporal evolution distribution dataset provided in the specific embodiment of the present invention is as follows: select the random forest model that performs best in cross-validation, and perform pixel-by-pixel extrapolation of soil pH values in farmland areas.
[0053] Specifically, the environmental covariates from 1980 to 2020 were input into the optimal random forest model to extrapolate the soil pH value of farmland area pixel by pixel, generating a dataset of the spatiotemporal evolution distribution of farmland soil pH with a spatial resolution of 30 m.
[0054] The present invention also provides a soil pH evolution meta-modeling device based on proton balance mechanism-guided machine learning, comprising: a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the soil pH evolution meta-modeling method based on proton balance mechanism-guided machine learning.
[0055] Example 1 In this embodiment, Chengmai County in northwestern Hainan Province, China, was selected as the study area. Using soil testing and fertilization data from 918 farmland surface soil samples collected in the 2010s and 144 farmland surface soil samples collected in 2023, a soil acidification history reconstruction and digital mapping method based on a proton balance mechanism-guided machine learning approach was employed. This resulted in the spatiotemporal evolution distribution of farmland surface soil pH from 1980 to 2020 at a spatial resolution of 30 m. The specific data and implementation details are shown below: Step (1): Integrate historical farmland soil test data from 2010 onwards, and construct a cross-period soil property verification dataset after spatial filtering and outlier removal. The sampling strategy in 2023 adopted a stratified sampling design, focusing on paddy fields and dryland areas, and collected 0-20 cm topsoil samples during the fallow period after crop harvest.
[0056] Step (2): While conducting field sampling, information on the main crop types, fertilization habits, and planting patterns over the past forty years in the study area was systematically collected. The main crop rotation systems in Chengmai County include double-cropping rice, melon-vegetable-rice rotation, and tropical fruit cultivation. There are differences in the geochemical responses of paddy fields and dry land to soil acidification.
[0057] Based on Sentinel-2 and historical Landsat imagery, the peak characteristics of time-series vegetation indices were used to identify paddy fields (bimodal) and drylands (unimodal or irregular peaks), thereby delineating the spatial distribution of regional crop planting. Combined with yield data from statistical yearbooks, the total amount of ions removed by crop harvest and the net acidification flux from atmospheric deposition and fertilization inputs were quantified. Initial soil properties and climate data were then used as input variables for a VSD+ process model. The Markov chain Monte Carlo method was employed to perform Bayesian calibration and posterior distribution sampling of the parameters, thereby simulating the annual soil pH evolution sequence from 1980 to 2020.
[0058] Step (3): The system collected environmental covariates, including Sentinel-1 synthetic aperture radar remote sensing variables, topographic variables, climate variables, soil background properties and agricultural management data.
[0059] SAR imagery was acquired using Sentinel-1 GRD products, with thermal noise cancellation, radiometric calibration, and topographic correction performed using dual-polarization data (VV and VH). The mean backscattering coefficients of VV and VH during the growing season and their ratio (VV / VH) were extracted to characterize surface roughness and dielectric properties related to soil moisture.
[0060] Topographic variables were calculated based on SRTM DEM data using SAGA GIS software, including elevation, slope, aspect, multi-resolution valley floor flatness index, and topographic humidity index, to characterize potential pathways for water accumulation and solute leaching.
[0061] Meteorological data came from the National Tibetan Plateau Scientific Data Center, including precipitation, temperature, solar radiation, and atmospheric deposition data; soil texture and parent material information were extracted from HWSD and high-resolution soil grids.
[0062] The agricultural management data were extracted from the monthly average farmland irrigation water volume data of the China Monthly Water Consumption Raster Dataset (HSWUD) from 1965 to 2022; and from the application data of 13 fertilizer types extracted from the global crop-specific nitrogen fertilizer dataset from 1961 to 2020.
[0063] Step (4): Environmental covariate screening employed a recursive feature elimination method combined with random forest importance scoring to rank the explanatory power of VSD+ simulations of time-series pH and select the feature subset with the highest contribution. The results showed that soil temperature, bedrock thickness, annual rainfall, nitrogen application rate, elevation, and atmospheric acid deposition were the most critical factors affecting the spatial variability of soil pH in this region.
[0064] Step (5): Use the long-term pH data generated by VSD simulation as training labels, and divide the training and validation sets in an 8:2 ratio. Optimize the hyperparameters of the random forest using 10-fold cross-validation. For example... Figure 2 As shown, the results indicate that random forest performs better in handling nonlinear relationships and preventing overfitting, with its R-value on the validation set being higher each year. 2 All reached above 0.6, with a root mean square error (RMSE) of 0.12-0.13 pH units.
[0065] Step (6): The spatiotemporal distribution map of soil acidification was drawn using the rasterio library in Python. The selected environmental covariates from previous years were input into the trained optimal random forest model, and the pH distribution map in GeoTIFF format from 1980 to 2020 was derived and output year by year. Overall, the pH of farmland soil in Chengmai County showed a decreasing trend, the acidification range gradually expanded over time, and the overall soil acidity increased. The soil pH evolution meta-modeling method based on the proton balance mechanism guided by machine learning in this embodiment not only filled the gap in historical monitoring data but also reconstructed the dynamic process of the evolution of farmland soil in Chengmai County from "weakly acidic" to "strongly acidic". This embodiment obtained a spatiotemporal distribution map of surface soil pH with a spatial resolution of 30 m. Figure 3 As shown, in 1980, the northern and central-eastern parts of the study area still had a relatively large number of weakly acidic to near-neutral soils with a pH value higher than 5.5; by 2000, the low pH area began to expand in the central and northern parts, and some areas changed from weakly acidic to strongly acidic; by 2023, the area with a pH value lower than 5.25 or even lower than 5.0 further increased, and formed relatively continuous acidification patches in the central and northern farmland areas, while the area with a pH value higher than 5.75 decreased significantly.
[0066] On the other hand, the present invention also provides a soil pH evolution meta-modeling device based on proton balance mechanism-guided machine learning, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the soil pH evolution meta-modeling method based on proton balance mechanism-guided machine learning.
Claims
1. A method for meta-modeling soil pH evolution based on proton budget balance mechanism-guided machine learning, characterized in that, include: Collect historical or current pH values, initial soil properties, annual climate data, and agricultural management data from different soil samples in the test area. Then, perform standardized preprocessing. Input the preprocessed initial soil properties and annual climate data into the MetHyd model to obtain nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus. Input the preprocessed initial soil properties, annual climate data, and agricultural management data, as well as the nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus into the VSD+ model to obtain the soil pH evolution sequence of different soil samples in the test area within a set time period. The soil pH evolution sequence within a set time period of the area to be tested is used as a label, and a sample set is constructed with the corresponding multidimensional environmental variable set. The machine learning model is trained through the sample set to obtain the soil pH evolution model. When applied, the environmental variables of the area to be tested are input into the soil pH evolution model to obtain the annual spatiotemporal evolution distribution dataset of soil pH in the area to be tested.
2. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The pre-processed initial soil properties and annual climate data were input into the MetHyd model to obtain nitrification, denitrification, and mineralization reduction coefficients, as well as precipitation surplus, including: Based on initial soil properties and annual climate data, the environmental reduction function and water flux are calculated using the MetHyd model. The environmental reduction function is a rate reduction factor of the soil nitrogen transformation process relative to the optimum conditions under actual hydrothermal conditions, and is determined by both the temperature function and the soil moisture content function. The water flux refers to the vertical migration of water in the soil profile per unit time, which is calculated from the difference between precipitation and evapotranspiration combined with soil water holding capacity parameters. These variables are used as external drivers input into the VSD+ model.
3. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The initial soil properties include cation exchange capacity, soil bulk density, organic carbon content, clay content, base saturation, and calcium carbonate content. The annual climate data refers to daily precipitation, average temperature, and solar radiation data for a set number of years.
4. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The preprocessed initial soil properties, annual climate data, and agricultural management data, along with nitrification, denitrification, and mineralization reduction coefficients and precipitation surplus, were input into the VSD+ model to obtain the soil pH evolution sequence of different soil samples in the test area over a set time period, including: The VSD+ model includes the RothC module, which incorporates soil physicochemical properties, exogenous carbon and nitrogen input dynamics, and climate data input parameters. The RothC module divides soil organic carbon into five independent carbon pools and sets corresponding C / N stoichiometry and decomposition kinetic parameters for each pool. This is used to simulate processes such as organic carbon mineralization, nitrogen mineralization, and nitrification. It also combines hydrological leaching processes to characterize the dynamic changes in proton production, alkaline cation release, and migration in the soil system, providing key flux inputs for subsequent soil solution chemical equilibrium calculations. To characterize the multi-stage buffering system in the process of farmland soil acidification, the VSD+ model further comprehensively describes three types of buffering mechanisms under different pH ranges: First, based on the initial cation exchange capacity, base saturation, and soil solution ion concentration, the Gapon equation is used to dynamically solve the rapid equilibrium of exchangeable cations between the solid and liquid phases to characterize the short-term ion exchange buffering process; second, the mineral weathering rate, which is constrained by both parent material and climatic conditions, is introduced to characterize the long-term acid neutralization capacity of the soil; in the strongly acidic stage, the dissolution-precipitation equilibrium mechanism of gibbsite is coupled, and the inherent thermodynamic solubility product is used as a constraint to describe the Al content in the soil solution. 3+ With OH - The model aims to simulate the activation and release of aluminum under low pH conditions and its buffering effect on excess protons by establishing a balance relationship. Finally, the model integrates carbon and nitrogen transformation, ion exchange, mineral weathering, aluminum buffering, and hydrological leaching processes on an annual scale, and iteratively solves the H2O equation based on the soil solution charge balance equation. + The concentration was determined, and the temporal evolution of soil pH was calculated accordingly.
5. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The agricultural management data is calculated and then used as input to the VSD+ model. The calculation of the agricultural management data includes the calculation of external nutrient input flux, crop nutrient removal flux, and straw return carbon and nitrogen flux.
6. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 5, characterized in that, The calculation of the exogenous nutrient input flux includes: comprehensive calculation of atmospheric deposition, biological nitrogen fixation, agricultural fertilizer input and ammonia volatilization loss. Among them, the calculation of agricultural fertilizer input is based on nitrogen application intensity, soil clay content, temperature and fertilizer type coefficient to correct for ammonia volatilization loss, so as to determine the net acid-causing nitrogen input. The calculation of crop nutrient removal fluxes includes: based on crop planting patterns and grain yield statistics, combined with the straw-to-grain ratio, element content characteristics and straw return rate of different crops, the net flux of phosphorus, sulfur, basic cations and nitrogen removed from the farmland system with crop harvest; The carbon and nitrogen fluxes of straw returning to the field are calculated, including: based on the straw returning ratio, crop moisture content, root-to-shoot ratio and biomass characteristics, the organic carbon and total nitrogen fluxes entering the soil mineralization cycle through straw returning to the field are quantitatively calculated.
7. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The pH values of different historical or current soil samples from the area to be tested were collected as observations. The Markov chain Monte Carlo method was used to perform Bayesian calibration on key sensitive parameters, including: The logarithmic values of the aluminum hydrolysis equilibrium constant, the empirical index of aluminum dissolution, the logarithmic values of the aluminum-base cation exchange selectivity coefficient, and the logarithmic values of the hydrogen-base cation exchange selectivity coefficient were set according to empirical ranges in the literature. The mineral weathering rate parameters were combined with factors affecting weathering intensity such as parent material type, soil mineral composition, soil layer thickness, soil texture, temperature, and precipitation / runoff, and the prior range was set with reference to regional empirical values or weathering model results. Subsequently, in each MCMC sampling, the candidate parameter combination with the set prior distribution and prior range was input into the VSD+ model to simulate the soil pH of the corresponding soil sample points and years. The likelihood function was constructed using the residuals between the simulated values and the measured pH, and the posterior distribution of the parameters was updated step by step. Finally, the optimal parameter combination and its uncertainty range were obtained, so that the root mean square error between the simulated soil pH evolution sequence and the observed pH value was not higher than the set threshold.
8. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, The multidimensional environmental variable set includes climate variables, topographic variables, soil background attribute variables, agricultural management variables, and remote sensing features, among which: The climate variables include one or more of the following: total annual precipitation, average annual temperature, total annual solar radiation, wind speed, relative humidity, nitrogen deposition, and sulfur deposition; The terrain variables include one or more of the following: elevation, slope, curvature, and terrain humidity index extracted based on a digital elevation model (DEM); The soil background attribute variables include one or more of the following: soil texture, bulk density, cation exchange capacity, initial base saturation, and bedrock thickness. The agricultural management variables include one or more of the following: fertilizer application rate, irrigation amount, and crop planting pattern; The remote sensing image features include: Sentinel-1 radar data VV and VH polarizations and their ratios.
9. The method for meta-modeling soil pH evolution based on proton budget balance mechanism guided by machine learning according to claim 1, characterized in that, A soil pH evolution model is obtained by training a machine learning model using a sample set, including: Using the soil pH evolution sequence of the area to be tested within a set time period as a label, combined with a multidimensional set of environmental variables, a random forest machine learning model is trained. Hyperparameter optimization was performed using grid search, and accuracy was verified using root mean square error and coefficient of determination. Soil pH evolution models were then screened and determined.
10. A soil pH evolution meta-modeling device based on proton budget balance mechanism-guided machine learning, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the soil pH evolution meta-modeling method based on proton budget balance mechanism-guided machine learning according to any one of claims 1-9.