Machine learning-based prediction method for spatial distribution of soil organic matter and ph in arid-hot valley regions

By employing machine learning methods, we have solved the problems of quantifying the uncertainty of the spatial distribution of soil organic matter and pH in arid and hot valley regions and extrapolating it across geomorphological regions. This has enabled high-precision and robust prediction of soil properties, supporting regional farmland management.

CN121707013BActive Publication Date: 2026-05-01INST OF AGRI ENVIRONMENT & RESOURCES YUNNAN ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF AGRI ENVIRONMENT & RESOURCES YUNNAN ACAD OF AGRI SCI
Filing Date
2026-02-11
Publication Date
2026-05-01

Smart Images

  • Figure CN121707013B_ABST
    Figure CN121707013B_ABST
Patent Text Reader

Abstract

The application discloses a machine learning-based dry-hot valley region soil organic matter and pH spatial distribution prediction method, relates to the technical field of soil spatial information and agricultural resource monitoring, and is based on a pixel and aligns multi-source covariates monthly, constructs machine learning and geostatistics three-line mutual evidence and quantile uncertainty; adopts a bidirectional circulation layer and time sequence self-attention coding, integrates a control layer to multi-head attention, and introduces landform unit, slope direction and altitude band condition coding for fusion and correction; generates a prediction interval with coverage guarantee under the condition of leaving one landform unit and one year, drives directional sampling and suitability confidence linkage; through a closed loop of sampling-retraining-revaluation, converges end error, extrapolation curve and coverage compliance, and outputs a high-precision prediction map, an uncertainty map, a sampling priority and a robust suitability partition.
Need to check novelty before this filing date? Find Prior Art

Description

A Machine Learning-Based Method for Predicting the Spatial Distribution of Soil Organic Matter and pH in Hot and Arid Valleys Technical Field

[0001] This invention relates to the field of soil spatial information and agricultural resource monitoring technology, specifically a method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning. Background Technology

[0002] Soil organic matter and pH are crucial parameters for assessing farmland quality and nutrient management. In arid-hot valley regions, soil properties exhibit strong spatiotemporal nonstationarity and endpoint fluctuations due to fragmented topography, significant slope aspect differences, concentrated seasonal rainfall, and intense evapotranspiration. Traditional spatial mapping methods driven by sparse sample point interpolation or single-temporal remote sensing factors struggle to accurately characterize fine-scale differentiation under complex terrain, easily leading to systematic biases in valley-ridge, shady-sun slope, and areas with scarce sample points. In recent years, digital soil mapping and machine learning methods have been widely used for soil property mapping, but they generally suffer from three problems: First, training and validation often rely on random sampling, neglecting spatial extrapolation caliber and making it difficult to assess practical application capabilities; second, most models only provide point predictions, lacking uncertainty quantification consistent with the assessment caliber, making it difficult to provide usable risk data for sampling optimization and suitability evaluation; third, multi-source temporal information is not systematically integrated, and prediction drift caused by seasonal phases and interannual differences is not effectively suppressed, resulting in endpoint area errors and unstable cross-topographic extrapolation performance.

[0003] Existing technologies typically construct soil sample point-environmental factor relationships and employ traditional geostatistical interpolation or ensemble learning regression to generate attribute maps. However, they still have the following limitations: First, the evaluation criteria and application criteria are inconsistent, and the lack of spatial blocking cross-validation, such as leaving one geomorphic unit outside the model, leads to an overestimation of extrapolation performance. Second, although the model can superimpose several uncertainty indicators, they are mostly empirical or interval estimates that do not match the evaluation criteria, failing to provide predictive intervals with verifiable coverage targets and confidence levels. Third, sampling density relies heavily on empirical point selection or single residuals and uncertainty thresholds, failing to jointly optimize coverage gaps, interval widths, residual hotspots, and cost accessibility under unified constraints, resulting in insufficient sampling efficiency and endpoint improvement. Fourth, suitability evaluation and prediction uncertainty are decoupled, leading to hard judgments in high-uncertainty areas and insufficient prescription robustness. Fifth, the utilization of multi-temporal remote sensing and climate monthly series remains at the level of statistical summarization, lacking time-series coding that can simultaneously capture short-term lags, seasonal cycles, and interannual dependencies, making it difficult to avoid system drift caused by seasonal mismatches.

[0004] Therefore, there is an urgent need for a method to predict the spatial distribution of soil organic matter and pH in complex terrains and highly seasonal conditions in hot and dry valleys. This method should be able to achieve high-precision prediction and uncertainty quantification with consistent spatial extrapolation assessment under a unified caliber. It should provide prediction intervals with coverage guarantees, and use targeted resampling driven by coverage gaps and interval widths to form robust prescriptions by linking suitability weights and thresholds. The method should also be able to stably stop under clear convergence criteria through a closed loop of resampling-retraining-reassessment, thereby improving endpoint robustness and cross-geomorphic extrapolation capabilities, and supporting refined farmland management and policy implementation at the regional scale. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and propose a machine learning-based method for predicting the spatial distribution of soil organic matter and pH in hot and dry valleys, so as to solve the above-mentioned problems.

[0006] The objective of this invention is achieved through the following technical solution: a method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning, comprising the following steps:

[0007] S1 Data and Spatiotemporal Organization: Surface soil samples were collected in the hot and dry valley region, and multi-source covariate data were obtained. The covariates included at least multi-temporal remote sensing indices, topographic second-order factors, climate monthly series, and anthropogenic factors. The covariates were resampled and standardized over time to form a monthly aligned pixel-level time series. A dataset for training and inference was constructed based on raster pixels and supplemented by geographic units of a preset size.

[0008] S2 baseline three-line mutual verification and uncertainty modeling: The machine learning branch is trained to obtain the first prediction result and the intermediate state of the method. The geostatistical branch is trained to interpolate the residuals to obtain the second prediction result and spatial structure parameters. The regression model with quantile properties is used to output the pixel-level quantile interval and endpoint error information as the input for subsequent integration.

[0009] S3 Temporal Augmentation Coding: Construct a temporal encoder composed of a recurrent neural network and a temporal self-attention network to encode a monthly aligned covariate sequence and output a time-step temporal representation and an attention-based sequence summary to characterize short-term lags, seasonal cycles, and year-end dependencies.

[0010] S4 process-level integration and conditional modeling: Build an integrated control layer, and use multi-head attention to focus on the data domain, method intermediate state domain, uncertainty domain, and suitability and sampling domain respectively. Under the constraints of introducing geomorphic unit identification, slope aspect and elevation zone conditional coding and spatial adjacency bias, the first prediction result and the second prediction result are fused and biased to obtain the pixel-level fused prediction value, and the basic representation of the strategy variable is output at the same time.

[0011] S5 Coverage Guarantee Risk Constraints: A risk verifiable guarantee module is set up in the integrated control layer. Based on the calibration set segmentation with one-year leave-out geomorphic unit and one-year leave-out, a prediction interval with coverage guarantee is generated for each pixel or geographic unit according to the coverage target and confidence level. The output is the coverage rate estimate, coverage gap and interval width, and the coverage gap and interval width are used as risk constraint inputs for the strategy variables.

[0012] S6 directional sampling strategy generation: Based on coverage gaps, interval widths, and residual hotspot information, and combined with sampling costs and traffic accessibility, a sampling priority map and a list of candidate sampling points are generated under the constraints of geomorphic unit coverage to guide field sampling.

[0013] S7 Suitability Confidence Linkage: Maps coverage gaps and interval widths to suitability confidence decay coefficients, dynamically adjusts the indicator weights or thresholds in multi-indicator suitability evaluation, and outputs robust suitability partitions and cell-level robustness.

[0014] S8 Closed Loop and Convergence: Iterative execution follows a closed-loop process of supplementation-retraining-re-evaluation. Iteration stops when the endpoint error relative to the initial iteration decreases to a preset proportion under the specified evaluation criteria, the extrapolation performance improves to a preset magnitude, and the coverage complies with the preset tolerance. The final spatial distribution results of soil organic matter and pH, as well as their uncertainties and strategy results, are output.

[0015] The covariates selected in S1 include: multi-temporal vegetation and moisture index, texture statistics, drought-related indices; slope, aspect, curvature, topographic position index, and humidity index derived from the digital elevation model; monthly precipitation, temperature, and evapotranspiration; distance from roads, settlements, and irrigation canals; sloping farmland intensity and land use type, etc.; temporal resampling is performed monthly or ten-day periods, and missing values ​​are handled by seasonal decomposition and local regression or imputation based on neighborhood consistency; geographic units are spatial aggregation units of a preset size composed of multiple raster pixels, used for zonal inference and statistical aggregation.

[0016] Uncertainty modeling in S2 employs an ensemble learning model with quantile properties to output preset upper and lower quantile pixel-level intervals and corrects for system biases in endpoint regions. The quantile intervals, pixel variances, and endpoint error information are used as inputs to the uncertainty domain of the ensemble control layer. The intermediate states of the method include the hidden feature representations of the machine learning model, the node splitting information of the ensemble tree, and the spatial structure parameters of the geostatistical branches.

[0017] The S3 temporal encoder consists of a bidirectional recurrent layer and a temporal self-attention layer cascaded together. The former represents local order and seasonality, while the latter represents long-term dependence and cross-year consistency. It also enhances the expression of intra-year changes by embedding time position and seasonal cycle, and outputs a time-step representation and sequence summary.

[0018] The multi-head attention mechanism of the S4 integrated control layer includes: a data head for modeling multi-temporal covariates and static features; a method head for fusing machine learning intermediates and statistical spatial structure parameters; an uncertainty head for aggregating pixel-level quantile intervals, pixel variances, and endpoint errors; and a policy head for linking suitability scores, weights, and candidate sampling points. The integrated control layer completes the joint output of fused predicted values ​​and policy variables by introducing conditional encoding of geomorphic unit identifiers, slope aspects, and elevation zones, as well as spatial adjacency bias.

[0019] The S5 risk verifiability assurance module determines the minimum width prediction interval that meets the coverage target and confidence level on the calibration set and outputs the coverage deviation index. The calibration set is divided by spatial blocking with one geomorphic unit left out to avoid spatial leakage. When the data spans multiple years, a time blocking with one year left out is added to avoid time leakage and enhance robustness verification.

[0020] S6's targeted sampling strategy combines coverage gaps, interval widths, and residual hotspots in a weighted manner, and introduces sampling costs and traffic accessibility constraints to generate sampling priorities, prioritizing sampling in areas with insufficient coverage and large errors, while ensuring the coverage ratio of different geomorphic units.

[0021] S7’s suitability confidence linkage applies the suitability confidence decay coefficient to the indicator weights or thresholds of multi-indicator evaluation, thereby achieving dynamic adjustment of robust suitability partitions and outputting a robustness map to express the degree of suitability stability.

[0022] The convergence criteria of S8 include: the endpoint error decreases by a preset percentage relative to the initial iteration, the overall improvement of the extrapolation distance-performance curve reaches a preset magnitude, and the coverage deviation is within the preset tolerance range of the target coverage. All of the above criteria are calculated within the spatial blocking cross-validation framework outside the leave-one geomorphic unit, and robustness checks are performed when necessary by combining the time blocking caliber outside the leave-one year.

[0023] The method outputs include: pixel-level spatial distribution maps of soil organic matter and pH, uncertainty maps (including prediction interval width and coverage gaps), targeted resampling priority maps and sampling point lists, robustness suitability zoning and robustness maps. The output is exported in geographic raster data format and provides an inference service interface to support updates and reproduction.

[0024] The beneficial effects of this invention are:

[0025] First, S1's data and spatiotemporal organization align multi-temporal remote sensing indices, second-order topographic factors, monthly climate series, and anthropogenic factors at a unified time step and spatial resolution, eliminating systematic biases caused by spatiotemporal mismatches and significantly improving the statistical consistency between training samples and pixel-level inference. The combination of monthly aligned pixel-level time series and preset-size geographic units ensures both fine-grained expression at the pixel scale and takes into account the throughput and stability of batch processing, providing a reliable foundation for subsequent temporal modeling and spatial validation.

[0026] Secondly, the baseline three-line mutual verification and uncertainty modeling of S2 output the first prediction result and method intermediate state through the machine learning branch, and obtain the second prediction result and spatial structure parameters by spatial interpolation of the residuals through the geostatistics branch. Finally, it outputs pixel-level quantile interval and endpoint error information using a quantile property regression model. This design achieves a balance between prediction accuracy and spatial structure consistency: machine learning undertakes nonlinear feature mining, geostatistics undertakes spatial autocorrelation correction, and quantile uncertainty provides a quantitative characterization of endpoint stability and interval scale, providing an actionable confidence basis for endpoint error control and strategy generation.

[0027] Furthermore, S3's temporal enhancement coding employs a cascaded bidirectional recurrent layer and a temporal self-attention layer, injecting temporal location and seasonal cycle embeddings to simultaneously capture short-term lags, seasonal variations, and interannual dependencies. Compared to schemes using only static or annual average features, this encoder exhibits higher sensitivity to intra-annual fluctuations in regions with strong seasonality and extreme water and heat distribution, such as hot-dry valleys, reducing the drift risk caused by seasonal phase misalignment and providing a stable temporal semantic representation for subsequent integration.

[0028] Furthermore, S4's process-level integration and conditional modeling incorporates multi-head attention within the integration control layer, targeting the data domain, method intermediate state domain, uncertainty domain, and suitability and sampling domain respectively. It also introduces conditional encoding of geomorphic unit identifiers, slope aspect, and elevation zones, as well as spatial adjacency bias. This integration approach is not a simple weighting, but rather uses conditional attention to achieve cross-domain information coordination and adaptive bias correction, enabling the fusion of the first and second prediction results to dynamically adjust with changes in geomorphology and slope aspect. Simultaneously, it outputs policy variables, achieving co-domain linkage between prediction, uncertainty, and policy, reducing the caliber drift commonly seen when multiple models are stacked.

[0029] S5's coverage guarantee risk constraints are calibrated under a dual-blockage caliber, leaving one geomorphic unit and one year behind. It generates a minimum-width coverage guarantee prediction interval based on preset coverage targets and confidence levels, and outputs coverage estimates, coverage gaps, and interval widths. This module provides verifiable probabilistic control capabilities consistent with the evaluation caliber, ensuring that uncertainties are no longer confined to empirical intervals or relaxed assumptions. Instead, it provides verifiable interval and coverage deviation indices under strict spatiotemporal blockage conditions, thereby improving the risk controllability of extrapolated scenarios and endpoint regions.

[0030] S6's targeted supplementation strategy incorporates coverage gaps, interval width, residual hotspots, sampling costs, and accessibility into a unified score, generating a sampling priority map and a list of candidate sampling points under geomorphic unit coverage constraints and budget constraints. This strategy integrates areas of uncertainty, large errors, low costs, and essential coverage into a single optimization objective, avoiding the one-sidedness of conventional approaches that only consider uncertainty or residuals. Under the same number of simultaneous data points and budget constraints, it reduces endpoint errors more quickly and improves extrapolation performance curves, thereby increasing the marginal return on supplementation investment.

[0031] S7's suitability confidence linkage maps coverage gaps and interval widths to suitability confidence decay coefficients and dynamically adjusts the weights or thresholds of indicators in multi-indicator suitability evaluation, producing robust suitability partitions and pixel-level robustness. This linkage avoids outputting seemingly accurate suitability results in high-uncertainty areas, suppresses misjudgments caused by inconsistencies between probability and evaluation criteria, and directly incorporates risk information into management prescriptions, making suitability results both risk-sensitive and management robust.

[0032] S8 employs an iterative mechanism of complementation-retraining-reevaluation for loop closure and convergence, using three criteria—endpoint error reduction, extrapolation performance curve improvement, and coverage compliance—as concurrency conditions. Calculations are performed under spatial blocking cross-validation, with temporal blocking used for robustness checks when necessary. This convergence design prevents overfitting or local optima caused by single-index optimization, ensuring that model improvements simultaneously achieve acceptable levels of improvement in endpoint robustness, spatial extrapolation, and probability coverage. The process provides clear and verifiable criteria for when to stop.

[0033] A unified pixel grid, geographic unit, geomorphic unit, and chronological caliber are used throughout training, calibration, and evaluation to ensure the traceability of data and model intermediate states. The integrated control layer and time encoder inherit parameters between iterations to avoid probability caliber drift caused by each round of reset. The spatial distribution map of soil organic matter, spatial distribution map of pH, prediction interval width map and coverage gap map, sampling priority map and sampling point list, robustness suitability zoning and robustness map exported in geographic raster data format, together with the inference service interface, constitute a closed-loop asset of results-service-retraining, which is easy to reproduce and audit. Attached Figure Description

[0034] Figure 1 is a flowchart of the present invention;

[0035] Figure 2 is a flowchart of the present invention;

[0036] Figure 3 is a flowchart of the present invention. Detailed Implementation

[0037] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Example 1

[0039] As shown in Figure 1, this embodiment focuses on the topsoil organic matter and pH in arid-hot valley regions. Topsoil samples were obtained from agricultural production sites. A unified coordinate system was used for spatial reference, and a unified monthly scale was used for temporal reference. The covariate system covered multi-temporal remote sensing indices, second-order topographic factors, monthly climate series, and anthropogenic factors. The raster organization was based on pixels, with preset-sized geographic units used for batch processing and statistics. The processing was streamlined using a rasterized pipeline, resulting in reproducible datasets, models, and output products.

[0040] In the data and spatiotemporal organization phase, surface soil samples were collected. After coordinate correction and quality verification of the sampling points, they were spatially overlaid with multi-source covariates, using a monthly scale as the unified time step to align remote sensing and meteorological time series. Spatiotemporal joint statistics were calculated for each covariate, and standardization transformations were performed. in The original values ​​of the covariates. and These represent the spatiotemporal joint mean and standard deviation of the same covariate across all raster pixels and the observation period. Missing time series segments are imputed using an interpolation process constructed from seasonal components and local regression, ensuring the continuity of the monthly aligned pixel-level time series. All covariates are resampled to a uniform spatial resolution, retaining the raster pixel-level data as the basic input, and preset-sized geographic units are set for zoning inference and statistical aggregation. After this step, a monthly aligned pixel-level time series dataset and its accompanying sample attribute table are obtained for subsequent training and inference.

[0041] In the baseline tri-line verification and uncertainty modeling phase, machine learning and geostatistics branches are constructed to jointly generate the first and second prediction results. A quantile property regression model outputs pixel-level quantile intervals and endpoint error information. At the sample point level, the machine learning branch uses surface soil organic matter and pH as supervised targets, inputting aligned and standardized covariate features to train an ensemble learning model to obtain the first prediction result. (The prediction results output by the machine learning branch) and the intermediate state representation of the method. The geostatistics branch uses the machine learning residuals as the modeling object to perform spatial structure recognition and interpolation, outputting a second prediction result. (Predicted results from geostatistical branch output) and spatial structure parameters. Interval estimation of upper and lower quantiles in the quantile-based regression model at the pixel scale. (in For the lower quantiles, such as , For upper quantiles, such as The endpoint deviation index is calculated in the endpoint region. The aforementioned quantile intervals and endpoint deviations serve as inputs to the uncertainty domain and are integrated along with the first and second prediction results.

[0042] In the timing enhancement coding stage A temporal encoder consisting of a recurrent neural network and a temporal self-attention network is constructed for monthly aligned covariate sequences. Output time-step representation and sequence summary Used to characterize local order, seasonal cycles, and interannual dependencies. Let the time series of the covariates of a pixel be... (in For the first Covariates at each time step (For sequence length), first obtain it through a bidirectional recurrent layer.

[0043] in Representing the complete time series Hidden state sequence output by a bidirectional recurrent layer (BiLSTM) Its dimensions are ( sequence length (For hidden state dimensions). Then input the temporal self-attention layer and overlay the embedding of time position and seasonal cycle. get

[0044] in Deterministic embedding of time location and seasonal cycle The representation sequence output by the temporal self-attention layer (Transformer) Their dimensions are the same Sequence summarization is obtained through attention convergence. in A sequence summary obtained through attention pooling (AttnPool) A fixed-dimensional vector (dimension) And retain the representation of each time step. This allows for subsequent interaction with intermediate states, uncertainties, suitability, and sampling domains.

[0045] In the process-level integration and conditional modeling stage, an integrated control layer is built. Multi-head attention is used to target the data domain, method intermediate state domain, uncertainty domain, and suitability and sampling domain respectively. Conditional encoding of geomorphic unit identifiers, slope aspect, and elevation zones is introduced. Spatial adjacency bias is used to constrain the attention distribution, completing the fusion and bias correction of the first and second prediction results, and simultaneously outputting policy variables. It is assumed that the linear transformation of the query, key, and value within the integrated control layer yields... Adding conditional coding and spatial adjacency bias Then, calculate attention. in , , These are query, key, and value matrices, obtained from the input representation through a linear transformation. For querying the channel dimension of the key, Generated from geomorphic unit identifiers, slope aspect and elevation zone conditional encoding, and spatial adjacency relationships, it is used to restrict attention coupling between spatially non-adjacent cells. This time-series summary... The intermediate states and uncertainty domains of the method converge into a unified representation. (Unified representation vector), through gated fusion, the deviation correction and weighted synthesis of the first and second prediction results are achieved to obtain the pixel-level fused prediction value. in To integrate gating weights ( For the gating weight parameter vector, (This refers to a logical function, specifically the sigmoid function). The deviation correction term is output by the integrated control layer. This is the final pixel-level fusion prediction value. This is a unified representation vector that aggregates time-series summaries, method intermediates, and the uncertainty domain. The above output serves as the final pixel-level fused prediction value, while also providing policy-related internal representations for subsequent policy generation and service integration. Based on this, rasterized results are generated, including spatial distribution maps of soil organic matter and pH, as well as an uncertainty-based base layer generated from the uncertainty domain. Exporting this data in geographic raster format is supported, and an inference service interface is provided for updating and reproducing the results.

[0046] Data sources span from soil sampling, covariate construction and alignment, sample overlay, and training sample generation. The processing involves standardization, missing data imputation, three-line cross-validation modeling, temporal coding, integrated control, and rasterization export. The output includes a fused prediction map and an uncertainty-based base layer. A unified spatial mask and region boundaries are used throughout the data processing. To ensure pixel index consistency, the training process employs partitioned parallelism and batch processing, while the inference process is segmented at the geographic unit scale to avoid memory overflow and ensure stability in large-area processing. The output generated by the above embodiments directly supports the pixel-level spatial distribution maps and uncertainty maps of soil organic matter and pH in the delivered results, and completes reproduction and incremental updates through the inference service interface to meet the needs of engineering implementation. The output also includes internal variables such as method intermediate state representation (hidden features of the machine learning model, ensemble tree node information, geostatistical spatial structure parameters), temporal summary representation, and basic representation of policy variables, which are available for subsequent modules to call.

[0047] The following provides the preferred parameter ranges and specific implementation methods for the key technical aspects in this embodiment. It should be understood that the parameter ranges described below are exemplary, and those skilled in the art can make adaptive adjustments based on the regional characteristics and data conditions in actual application scenarios without departing from the protection scope of this invention.

[0048] During the data and spatiotemporal organization phase, the following data sources and processing standards are preferred:

[0049] The preferred remote sensing data sources are Sentinel-2 multispectral images (spatial resolution 10m) or Landsat 8 / 9 satellite images (spatial resolution 30m). Based on these remote sensing images, the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Normalized Difference Water Index (NDWI), Gray-Level Co-occurrence Matrix (GLCM) texture features, and drought-related indices are calculated. For continuous covariates, bilinear interpolation is used for resampling, while nearest neighbor interpolation is used for categorical or discrete covariates.

[0050] For terrain data sources, preferred digital elevation models (DEMs) such as SRTM, ALOS, or ASTER (spatial resolution 10–30 m) are used. Slope, aspect, curvature, topographic humidity index (TWI), and topographic potential index are derived from these DEMs. For climate data sources, monthly precipitation, temperature, and evapotranspiration data are preferred. Specifically, ERA5-Land reanalysis data, the China Regional Surface Meteorological Element Driven Dataset (CMFD), or local meteorological grid data (spatial resolution 1–10 km) can be used. These data are downsampled to the target raster resolution and zonal calibration is performed to eliminate scale bias. For data sources of anthropogenic factors, preferred OpenStreetMap open street maps or local vector data are preferred, including spatial elements such as road networks, settlements, and irrigation canals, as well as anthropogenic activity indicators such as land use type and sloping farmland development intensity. Regarding raster organization and coordinate system, all covariate data are uniformly resampled to a spatial resolution of 10–30m (preferably 10m or 30m), and the coordinate system is preferably the CGCS2000 National Geodetic Coordinate System or the Universal Transverse Mercator Projection (UTM) coordinate system; all raster cells are strictly aligned by using a uniform mask boundary.

[0051] Regarding the sample size and data division, 300–800 surface soil samples are preferably collected in the county-scale study area. The samples are divided according to a ratio of approximately 6:2:2 for the training set / calibration set / blind test set, or the training and verification process is organized using a K=5 fold geomorphic unit blocking cross-validation method.

[0052] In the geostatistical branch implementation stage of baseline three-line mutual verification and uncertainty modeling, the following technical solutions are preferred:

[0053] The variogram model is preferably selected from the spherical model, exponential model, and Gaussian model, or the optimal model is determined by comparison through grid search; the variogram parameter estimation preferably uses the weighted least squares (WLS) method or the maximum likelihood estimation method (ML).

[0054] Kriging interpolation methods preferably use the residuals of machine learning branches as the modeling object and employ the ordinary kriging method (OK) for spatial interpolation. When the spatial structure exhibits significant directional differences, anisotropic kriging can be further considered, with its main direction determined through statistical analysis of geomorphic units or slope aspects.

[0055] The search neighborhood and number of samples are set as follows: a fixed search radius r = 3–10 pixels, or k = 8–24 using the k-nearest neighbor method; to ensure the stability of the estimation, the minimum number of samples for each interpolation is no less than 8.

[0056] The spatial structure parameter output includes sill value, range, nugget value, etc., and these parameters are used as components of the intermediate state representation of the method for the integrated control layer.

[0057] In quantile property regression models, random forest quantile regression (QRF) or quantile gradient boosting tree (XGBoost) are preferred as modeling algorithms; the typical quantile set is set as {0.1, 0.5, 0.9}, which can be extended to {0.05, 0.95} when a higher level of coverage guarantee is required.

[0058] The endpoint correction method is as follows: identify systematic biases in endpoint prediction on the calibration set of spatial blockage, correct the endpoint biases using linear correction or piecewise linear correction methods, and pass the corrected endpoint error information as input to the integrated control layer as an uncertainty domain.

[0059] In the temporal enhancement coding stage, the Bidirectional Long Short-Term Memory (BiLSTM) network is configured with 1–2 layers, a hidden layer dimension d = 128–256, and a dropout ratio of 0.1–0.3. Residual connections can be selectively added to enhance gradient propagation. The Temporal Transformer network is configured with 2–4 layers, a multi-head attention mechanism with 4–8 heads, and a channel dimension d = 256–512. Relative position coding and seasonal periodic bias (period set to 12 months) are enabled. When computational resources are limited or the sequence is long, a sparse attention mechanism can be selectively employed, computing attention only on pixels within the spatial neighborhood and seasonal skip connections. For conditional coding, the embedding vector dimensions for topographic unit identifiers, slope aspect categories, and elevation zone categories are set to 16–64. Spatial adjacency bias is also implemented. The construction method is as follows: a spatial adjacency graph is constructed based on the k-nearest neighbor graph (k=8–16) or a fixed search radius (r=1–3 pixels). For non-adjacent pixel pairs, a negative bias is applied to the attention logits, with the bias amplitude ranging from approximately [-2,0], in order to suppress unreasonable coupling between spatially non-adjacent pixels.

[0060] In a converged gating mechanism, gating weights parameter vector Channel dimension and unified representation To maintain consistency (d=256–512), deviation correction term The output is obtained through a linear mapping layer, and its unbiasedness is verified on a calibration set to ensure the accuracy of the fusion results.

[0061] Example 2

[0062] As shown in Figures 1 and 2, this embodiment uses the pixel-level fusion prediction value, pixel-level quantile interval, endpoint error information, method intermediate state representation, and time series summary representation from Embodiment 1 as input. A risk verifiability assurance module is set up within the integrated control layer. A spatial blocking segmentation calibration set with one geomorphic unit left out is adopted. When the data spans multiple years, a time blocking with one year left out is added. The output includes a prediction interval with coverage assurance, coverage estimate, coverage gap, and interval width. Based on this, a sampling priority map and a candidate sampling point list are generated. Simultaneously, the confidence decay coefficient is linked to the weight and threshold of the suitability evaluation, ultimately forming a robust suitability partition and a pixel-level robustness map. These are exported together with the fusion prediction map in a geographic raster data format and an inference service interface is provided.

[0063] In the risk constraint part of coverage assurance, a spatial blocking caliber with one geomorphic unit left out is used to split the calibration set to ensure that there is no spatial leakage during the calibration process. When the data spans multiple years, a time blocking with one year left out is added to avoid time leakage and enhance robustness verification. A coverage gain / loss score is constructed for each calibration sample, and a set of symmetric coverage radii based on quantile intervals is defined, with the fusion prediction being... The radius of symmetry is Coverage judgment is Determine the minimum width interval radius on the calibration set that satisfies the coverage target and confidence level. Let the coverage bias score sequence in the calibration set be... ,in For the first Coverage radius score for each calibration sample Candidate values ​​for the radius of a symmetric interval. For the first Observations of a sample For the corresponding fusion prediction value, This is an indicator function (it takes the value 1 if the condition is met, and 0 otherwise). Here is the sample index for the calibration set. Its empirical quantile function is: Given a target coverage Under the condition of confidence level, take in To meet the minimum interval radius of the coverage target, For empirical quantile functions (i.e., taking sequences) The quantiles), The target coverage rate is (e.g., 0.9). This is the sample index set for the dual-blocking calibration set. From this, we obtain the pixel-level prediction interval: Interval width: in To predict the interval width. Coverage estimation: in For coverage estimation, To determine the number of samples in the validation set, For the validation set sample index, To verify the sample index set under the specified caliber. Coverage gap. The above For dual-blocking calibration set index, To validate the sample index under the specified caliber, this process outputs a prediction interval with coverage guarantee at each cell or geographic unit scale, and obtains coverage estimates, coverage gaps, and interval widths. These quantities are then injected back into the uncertainty domain of the integrated control layer as inputs to the strategy variables.

[0064] In the targeted sampling strategy generation section, residual hotspots, coverage gaps, interval widths, sampling costs, and traffic accessibility are aggregated to form a sampling priority score for pixels or candidate points. Let the residual intensity field be... The coverage gap is The interval width is The sampling cost is Traffic accessibility penalty is The set of geomorphic unit coverage constraints is The budget is To ensure dimensional consistency, the aforementioned , , , , Before calculating the priority score, the data is normalized to the [0,1] interval using a min-max method. Pixels or candidate points are defined. Priority rating

[0065] in For pixels or candidate points Sampling priority score, These are the weighting coefficients for each item. For residual strength, To cover the gap, The interval width, For sampling cost, Traffic accessibility penalties are imposed, and a sampling set is selected under budget and topographic cover constraints. :

[0066] in For the candidate sampling set, For the total budget, For the first A set of pixels or candidate points for a geomorphic unit. For the first Minimum number of sampling coverage units per geomorphic unit Represents the sample set In geomorphic units The number of samples in the dataset. This optimization can be solved using a hybrid heuristic greedy algorithm and grouped knapsack algorithm, prioritizing candidate points from areas with insufficient coverage and significant errors, while ensuring that each geomorphic unit reaches a preset lower coverage limit. The output of the obtained sample set is a sampling priority map and a list of candidate sampling points, including coordinates, priority, and constraint satisfaction status, for field execution.

[0067] In the suitability confidence linkage section, the coverage gap and interval width are mapped to a suitability confidence decay coefficient, and this coefficient is applied to the weights and thresholds of the multi-index suitability evaluation to obtain a robust suitability output. To ensure dimensional consistency, the coverage gap... With interval width Before calculating the attenuation coefficient, it is normalized to the [0,1] interval using min-max. The signal attenuation coefficient is set to...

[0068] in The suitability confidence decay coefficient. For mapping parameters, This means truncating the value to... Interval. Let the weight of the original suitability index be... The threshold is After stabilization

[0069] in For the robust version of the first Each suitability indicator weight, For the robust version of the first Individual thresholds, The original indicator weights, The original threshold, For threshold tolerance (here) This is an index, not an iterative index. Based on this, robustness and suitability partitioning and cell-level robustness are calculated to express the degree of stability under coverage-guaranteed caliber. (The above...) During the calibration phase, the conditions are determined through grid search or hierarchical cross-validation and are consistent with the condition coding of the integrated control layer to ensure consistency across different geomorphic units and different slope aspects and elevation zones.

[0070] Data sources are integrated across risk constraints, strategy generation, and suitability assessment. The risk verifiability assurance module calls upon the pixel-level fusion prediction values ​​and uncertainty domain output from Example 1, and uses the observed and predicted values ​​from the dual-blocking calibration set to calculate the quantile radius. The strategy generation module calls upon residual strength, coverage gap, interval width, cost, and accessibility layers. The suitability assessment module introduces a confidence decay coefficient to adjust weights and thresholds based on the existing suitability index system and sub-scores. The processing is executed continuously: after fusion and bias correction are completed at the integrated control layer, the risk verifiability assurance module outputs the interval, coverage estimate, coverage gap, and interval width; subsequently, the strategy generation module generates a sampling priority map and a candidate sampling point list under constraints; then, the suitability assessment module completes robust suitability partitioning and robustness calculation based on the confidence decay coefficient. The output is exported in geographic raster data format, including a prediction interval width raster with coverage assurance, a coverage gap raster, a sampling priority raster, a candidate sampling point list, robust suitability partitioning, and a robustness map, and provides query and incremental update capabilities through the inference service interface.

[0071] This embodiment supplements the key steps not covered in embodiment 1 by focusing on three aspects: coverage guarantee, targeted supplementation, and suitability confidence linkage. All intermediate quantities are generated under a unified evaluation and calibration calibrator and flow within the integrated control layer, ensuring that the strategy variables and robust suitability outputs are consistent with the pixel-level fusion prediction results, and providing reliable input for subsequent closed-loop and convergence execution.

[0072] Example 3

[0073] As shown in Figures 1 to 3, the following embodiments establish a closed-loop mechanism of supplementation-retraining-reevaluation based on the data and spatiotemporal organization, baseline three-line mutual verification and uncertainty modeling, temporal enhancement coding, process-level integration and conditional modeling, risk constraints of coverage guarantee, generation of targeted supplementation strategy and linkage of suitability confidence, which have been completed in Embodiments 1 and 2.

[0074] Regarding consistent assessment and calibration standards, the set of geomorphic units in the study area is denoted as... The set of year sequence is denoted as Each assessment uses a cut set outside a geomorphic unit under the spatial blocking caliber, and if necessary, overlays a time blocking caliber outside one year in the robustness check. For the first... In the next iteration, the training set, calibration set, and validation set are denoted as follows: ,in and It excludes the topographic unit corresponding to the current verification section, and in the time verification scenario, it excludes the current verification year. The timing encoder and the integrated control layer inherit the previous optimal parameters as the starting point, and perform incremental adaptation on newly added sampling points to avoid breaking the coverage constraints of the previous calibration.

[0075] Regarding closed-loop execution, the first The next iteration takes the prediction interval with coverage guarantee, coverage estimate, coverage gap, and interval width output from Example 2 as input, and combines residual hotspots, sampling costs, and traffic accessibility to generate a sampling priority map and a list of candidate sampling points; after the field sampling is completed, a new sample set is formed. .by As training data, the covariate scope, pixel grid, and geographic unit are kept consistent. The machine learning branch and geostatistics branch are retrained and the intermediate states of the method are updated. At the same time, the temporal encoder and integrated control layer are retrained or fine-tuned with the same scope. Under the premise of maintaining consistency in risk control constraints, joint loss is adopted.

[0076] Perform parameter updates, where For the first The joint loss function for the next iteration. For pixel-level fusion prediction, For pixel-level prediction distribution, For target coverage Actual coverage measurement below To cover deviation penalties, For spatiotemporal smoothing regularization, These are the weighting coefficients for each loss term. This is the root mean square error term. For continuous sorting probability scores (evaluating the predicted distribution) Compared with observed values (consistency) Coverage bias penalty (measures measured coverage) With target coverage (deviation) This is a spatiotemporal smoothing regularization term (to suppress spatial or temporal discontinuities in the prediction). The coverage target and confidence level remain consistent with the previous calibration to avoid evaluation caliber drift.

[0077] Regarding evaluation and convergence, iteration stops when three criteria—endpoint error index, extrapolation performance curve, and coverage compliance—are simultaneously met. Endpoint error is evaluated using the absolute prediction error of samples in the validation setpoint endpoint region. Let the measured target variable be... fusion prediction First, based on the validation set... The endpoint regions are divided by the quantiles of the target variable, let For target variable In the first Near the quantile (e.g.) Corresponding to the low end, An endpoint error is defined on a subset of samples corresponding to high-end samples.

[0078] in For the first The iteration is at the... Endpoint error of the quantile endpoint region It is a median function. For spatial coordinates and covariate characteristics, For the observed values ​​of the target variable, For the first The fused prediction value of the next iteration, The number of iterations ( (For the initial iteration). Requirements relative to the initial iteration. Endpoint error reduction

[0079] in The relative rate of decrease in endpoint error. The endpoint error of the initial iteration. For the first Endpoint error of the next iteration. This is a preset threshold for the percentage reduction in endpoint error (e.g., 0.2 indicates a required reduction of 20%). Extrapolation performance is measured by the area under the spatial distance-performance curve, denoted by the shortest surface distance to the training samples. Perform stratification, and denote the stratification performance function as: (Distance can be used) coefficient of determination at location (or the reciprocal of the root mean square error RMSE), area of ​​the curve

[0080] in For the first The area under the extrapolation performance curve of the next iteration. The spatial distance to the nearest training sample point. To achieve the maximum assessment distance, Distance Performance metrics at the location (which may be determined by the coefficient of determination) (Or the reciprocal of the root mean square error (RMSE)). A relative improvement is required.

[0081] in For the relative improvement rate of extrapolation performance, This represents the area of ​​the curve in the initial iteration. This is a preset performance improvement threshold. Compliance coverage is based on the target coverage rate. The actual coverage is used as the criterion, and the verification set coverage is denoted as . The tolerance is ,Require in For the first Validation set coverage in the next iteration. For target coverage, To define the coverage tolerance (the allowable deviation range), the three criteria are convergent using conjunctive logic, and a minimum satisfying iteration step is defined.

[0082] in The minimum number of iterations required to satisfy the convergence condition. The logical conjunction symbol (representing AND) is used, and convergence requires that all three criteria be satisfied simultaneously. To ensure robustness, the above indicators are calculated over all folds of spatial blocking cross-validation, and if necessary, reviewed once under the temporal blocking check. Convergence is based on the most unfavorable fold as the criterion for passing, avoiding premature termination due to single-fold randomness.

[0083] Regarding the synchronous update of strategy and suitability, the confidence decay mapping of Example 2 is used to update the coverage gap and interval width to the latest values, thereby obtaining a new suitability confidence decay coefficient. Based on this, the suitability weights and thresholds are adjusted, and robust suitability partitions and cell-level robustness are output. To prevent policy oscillations, a smoothing mechanism is introduced for sampling priority scoring between adjacent iterations, denoted as the previous round's score. The memory loss score for this round was Using exponential moving average ;

[0084] Suppressing large fluctuations, among which For the first The sampling priority score after smoothing in the next iteration For the first The next iteration calculates a memoryless score directly based on the current state. The score from the previous iteration, Smoothing coefficient ( The larger the value, the higher the retention rate of historical scores. This score is only used for candidate point selection and does not change the evaluation criteria of the convergence criterion.

[0085] Regarding data sources, all stages of the closed loop are based on a unified pixel raster, geographic unit, geomorphological unit, and chronological caliber. New sample sets are derived from field collection results of the targeted supplementation strategy. Covariate data and basic raster data follow the preprocessing and alignment rules of Example 1. Uncertainty domains and strategy variables follow the generation rules of Example 2 and are updated to the latest state after each iteration. During processing, training, calibration, and evaluation strictly adhere to spatial blocking cross-validation, with verification performed during temporal blocking checks when necessary. Parameters of the temporal encoder and integrated control layer are continuously transferred between iterations to maintain consistency in the probability caliber of previous calibrations. Coverage targets and confidence levels are fixed throughout the process to avoid inconsistencies between convergence criteria and interval generation calibers. In terms of output, after convergence, spatial distribution maps of soil organic matter and pH, prediction interval width and coverage gap maps, targeted supplementation priority maps and sampling point lists, robustness suitability zoning, and robustness maps are generated. All products are exported in geographic raster data format and provided for batch retrieval and incremental updates through the inference service interface. Simultaneously, final model weights, evaluation reports, and fold details are archived to support reproducibility and auditing.

[0086] Through the above embodiments, the closed loop of supplementation-retraining-reevaluation operates under a unified evaluation and calibration calibrator. It automatically stops when the three conditions of endpoint error reduction, extrapolation performance improvement and coverage compliance are met simultaneously. The final results are consistent with the strategy results and the coverage guarantee calibrator, ensuring the consistency and traceability of spatial distribution prediction, uncertainty quantification and sampling prescription in engineering applications.

[0087] Example 4

[0088] Example 4 includes comparative and ablation experiments, comparing the method with existing typical methods and the core modules of this invention. This example is based on a real-world application scenario in a county-level research area in a dry-hot valley, and is conducted under a unified dataset, evaluation framework, and cross-validation caliber.

[0089] This study selected a typical county-level dry-hot valley (approximately 1200 km²) as the research object, traversing its main geomorphic units and slope gradients. Based on a five-year field measurement database from 2018 to 2022, a total of 520 surface soil samples were collected. The sampling points were distributed across various years to ensure the effectiveness of cross-validation with a one-year time buffer. The study area was divided into eight main geomorphic units (such as mountains, hills, and plains). Sampling was stratified by geomorphic unit, with no fewer than 50 samples per unit, to ensure the robustness of the spatial buffer with one geomorphic unit remaining.

[0090] Following the standard procedures of Examples 1 to 3, data spatiotemporal organization, multi-source covariate alignment, and standardization were completed to form a monthly aligned pixel-level time series dataset. Sample points were divided into a training set, a calibration set, and a blind detection set in a 6:2:2 ratio, totaling 312, 104, and 104 samples, respectively. All samples in the blind detection set were evaluated using a spatial blocking method leaving one geomorphic unit outside the dataset. Simultaneously, a temporal blocking method leaving one year outside the dataset was applied during the calibration and validation of multi-year data to ensure that the evaluation results reflect true spatial extrapolation capability and interannual stability.

[0091] All contrast and ablation schemes were trained on a unified training and calibration set, and their performance was evaluated on a blind test set. Evaluation metrics included prediction accuracy (Root Mean Square Error (RMSE), Mean Absolute Percentage Error (MAPE), and Coefficient of Determination (R²), uncertainty quantification quality metrics (coverage, Continuous Rank Probability Score (CRPS), and interval width), spatial extrapolation capability metrics (AUC of distance decay curves, and stratified R² evaluation for different distance ranges), and sampling efficiency metrics (accuracy improvement per unit sampling cost and geomorphic unit uniformity). Each scheme was run independently three times to eliminate random fluctuations, and the mean and standard deviation were reported.

[0092] Comparative experiment

[0093] The comparative experiments involved four schemes: single machine learning branch (ML-Only), single geostatistical branch (Geo-Only), literature standard fusion method (Standard-Kriging), and the complete scheme of this invention (Full-Proposed). Table 1 shows a comprehensive comparison of each scheme in key indicators such as prediction accuracy, uncertainty quantification, and extrapolation performance.

[0094] Table 1: Comparative Experiments

[0095]

[0096] Note: In Table 1, the extrapolation R² refers to the relative improvement over distances of 4 km and above. The simultaneous occurrence of coverage improvement and narrowing of interval width indicates that this method, through rigorous calibration of the obfuscated prediction framework, can accurately calculate the minimum interval width while meeting the 90% coverage target, making it more efficient than the overly conservative estimation of Standard-Kriging.

[0097] ablation experiment

[0098] The ablation experiments progressively removed each key module (S3 timing coding, S4 integrated control, S5 coverage guarantee, S6 targeted sample replenishment, S7 suitability linkage), removing only a single module while keeping the others intact. Table 2 shows the performance loss of each ablation variant and the key role of that module.

[0099] Table 2: Ablation Experiment

[0100]

[0101] Note: In Table 2, sampling efficiency is defined as the number of new sampling points required to increase coverage by 1% after two iterations within the leave-one-topographic-unit cross-validation framework (unit: points / %). The baseline is 93.1% coverage in the first round, and the target is 99.0% coverage. R² loss is calculated using the ablation method: only a single module is removed, while other modules remain fully implemented to ensure a valid control. - After S6 targeted sampling, the sampling efficiency deteriorated from 0.32 points / % to 1.24 points / %, a relative deterioration of 287% (requiring 3.9 times the number of sampling points), demonstrating the key value of the S6 strategy for sampling optimization. - S5 coverage guarantees virtually no R² loss (only -0.6%), but the coverage rate plummeted by 7.5% (from 93.1% to 85.6%), reflecting that the module's function is to ensure the reliability of uncertainty quantification rather than improve point prediction accuracy. The geomorphic unit balance is defined as the percentage of the standard deviation of the number of newly added sampling points in each unit divided by the average value. Taking 8 units as an example, if the distribution is uniform, the number of newly added sampling points per round is approximately 0.24 points / unit, and the standard deviation of 0.142 indicates that the standard deviation of the actual distribution does not exceed 0.034 points / unit, indicating that the distribution is extremely balanced. However, the -S6 scheme, without constraints, causes the standard deviation to increase to 0.099 points / unit (coefficient 0.412), indicating that some units are oversampled while other units are undersampled.

[0102] Comparative experiments show that this invention achieves a 16.1% reduction in RMSE and a 6.8% improvement in R² compared to the standard method in the literature. This improvement stems from the synergistic effect of five core modules. First, the temporal coding layer, through bidirectional LSTM and a temporal self-attention network, characterizes the seasonal cycle and interannual dependence of covariates, and ablation experiments show that it directly contributes approximately 5.4% to the accuracy improvement. Second, the multi-head attention fusion layer, through relatively simple constraints of conditional coding and spatial adjacency bias, achieves an average accuracy improvement of approximately 3.0%. The combined core contributions of these two modules amount to 8.4%. Third, the nonlinear synergistic fusion of S3 and S4, the improvement in extreme value prediction through the endpoint error correction mechanism, and the optimization of data distribution by S6 and S7 in the second and third rounds of the closed-loop iteration contribute the remaining 7.7% to the accuracy improvement.

[0103] From the perspective of uncertainty quantification, the complete scheme achieves a coverage rate of 93.1%, close to the 90% target, while the interval width narrows by 18%. This seemingly contradictory phenomenon actually reflects the efficiency of the perplexed prediction framework: this framework determines the minimum interval width on the calibration set with one geomorphic unit left out, ensuring that the proportion of samples with true values ​​within the prediction interval is exactly ≥90%. It accurately calculates the minimum interval while meeting the coverage target, thus being more efficient than the overly conservative design of Standard-Kriging. Meanwhile, the improvement in CRPS (23.7%) is significantly greater than the improvement in RMSE (16.2%), and far greater than the improvement in R² (6.8%), indicating that this method, through the perplexed prediction framework and endpoint error correction, not only improves the accuracy of point prediction but, more importantly, enhances the reliability and confidence level accuracy of the prediction interval.

[0104] In terms of spatial extrapolation capability, this method achieves a 19.9% ​​relative improvement over Standard-Kriging (R²=0.618) over distances greater than 4km, with an absolute R² of 0.741, which is considered moderate. This improvement mainly stems from the S3 time-series coding's ability to learn and generalize interannual dependencies and long-term cyclical patterns. Research shows that in scenarios with a sample spacing of approximately 1.5km and a study area spanning 1200km², the interannual consistency patterns learned by the time-series coding can be effectively extrapolated to distances greater than 4km.

[0105] In summary, the five core modules achieved a 16.1% performance improvement, a 9.4% coverage improvement, and a 75.6% improvement in sampling efficiency relative to the benchmark method through division of labor and collaboration (comparison of 0.32 points / % vs 1.24 points / %).

[0106] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A machine learning-based method for predicting the spatial distribution of soil organic matter and pH in hot and dry valley regions, characterized in that, The process includes the following steps: S1 Data and Spatiotemporal Organization: Surface soil samples are collected in the hot and dry valley region, and multi-source covariate data are obtained. The covariates include at least multi-temporal remote sensing indices, topographic second-order factors, climate monthly series, and anthropogenic factors. The covariates are resampled and standardized over time to form a monthly aligned pixel-level time series. A dataset for training and inference is constructed based on raster pixels and supplemented with geographic units of a preset size. S2 Baseline Tri-line Verification and Uncertainty Modeling: The machine learning branch is trained to obtain the first prediction result and the intermediate state of the method. The geostatistics branch is trained to interpolate the residuals to obtain the second prediction result and spatial structure parameters. The regression analysis of quantile properties is then used to... The model outputs pixel-level quantile intervals and endpoint error information as subsequent integration inputs; S3 Temporal Enhancement Encoding: Constructs a temporal encoder composed of a recurrent neural network and a temporal self-attention network to encode the monthly aligned covariate sequence, outputting a time-step temporal representation and an attention-based sequence summary to characterize short-term lags, seasonal cycles, and year-end dependencies; S4 Process-Level Integration and Conditional Modeling: Builds an integration control layer, using multi-head attention to align the data domain, method intermediate state domain, uncertainty domain, and suitability and sampling domain respectively, and then introduces conditional encoding of geomorphic unit identifiers, slope aspect and elevation zones, and spatial adjacency bias constraints to integrate the first prediction result and the second... The prediction results are fused and biased to obtain pixel-level fused prediction values, and the basic representation of the strategy variables is output simultaneously; S5 Risk constraint for coverage guarantee: A risk verifiable guarantee module is set in the integrated control layer. Based on the calibration set segmentation with one-year retention and one-year retention, a prediction interval with coverage guarantee is generated for each pixel or geographic unit according to the coverage target and confidence level. The coverage estimate, coverage gap, and interval width are output, and the coverage gap and interval width are used as risk constraint inputs for the strategy variables; S6 Targeted supplementation strategy generation: Based on the coverage gap, the interval width, and residual hotspot information, and combined with sampling cost and traffic accessibility, a targeted supplementation strategy is generated for the geographic unit coverage. Under the constraint of coverage, a sampling priority map and a list of candidate sampling points are generated to guide field sampling; S7 Suitability confidence linkage: The coverage gap and the interval width are mapped to a suitability confidence decay coefficient, and the index weights or thresholds in the multi-index suitability evaluation are dynamically adjusted to output robust suitability partitions and pixel-level robustness; S8 Closure and convergence: The closed-loop process of sampling-retraining-reevaluation is executed iteratively. When the endpoint error relative to the initial iteration decreases to a preset proportion under the specified evaluation caliber, the extrapolation performance improves to a preset magnitude, and the coverage compliance meets the preset tolerance, the iteration stops, and the final spatial distribution results of soil organic matter and pH, as well as their uncertainties and strategy results, are output.

2. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The covariates in S1 include: multi-temporal vegetation and moisture index, texture statistics, drought-related index; slope, aspect, curvature, topographic position index and humidity index derived from the digital elevation model; monthly precipitation, temperature and evapotranspiration; distance from roads, settlements and irrigation canals, sloping farmland intensity and land use type; the time resampling is performed monthly or ten-day periods, and missing values ​​are imputed using seasonal decomposition and local regression or neighborhood consistency; the geographic unit is a spatial aggregation unit of a preset size composed of multiple raster pixels, used for partitioning inference and statistical aggregation.

3. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The uncertainty modeling in S2 uses a quantile-based ensemble learning model to output preset upper and lower quantile pixel-level intervals and corrects the system bias in the endpoint regions. The quantile intervals, pixel variances, and endpoint error information are used as the uncertainty domain input of the integrated control layer. The intermediate state of the method includes the hidden layer feature representation of the machine learning model, the node splitting information of the ensemble tree, and the spatial structure parameters of the geostatistical branches.

4. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The temporal encoder of S3 consists of a bidirectional recurrent layer and a temporal self-attention layer cascaded together. The former represents local order and seasonality, while the latter represents long-term dependence and cross-year consistency. It enhances the expression of intra-year changes by embedding time position and seasonal cycle, and outputs a time-step representation and sequence summary.

5. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 2, characterized in that, The multi-head attention of the integrated control layer of S4 includes: a data head for modeling multi-temporal covariates and static features; a method head for fusing machine learning intermediates and geostatistical spatial structure parameters; an uncertainty head for aggregating pixel-level quantile intervals, pixel variances, and endpoint errors; and a strategy head for linking suitability scores, weights, and candidate sampling points. The integrated control layer completes the joint output of the fused predicted value and the strategy variables by introducing conditional encoding of geomorphic unit identifiers, slope aspects, and elevation zones, as well as spatial adjacency bias.

6. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The risk verification and assurance module of S5 determines the minimum width prediction interval that meets the coverage target and confidence level on the calibration set and outputs the coverage deviation index. The calibration set adopts spatial blocking segmentation with one geomorphic unit left out to avoid spatial leakage. When the data spans multiple years, a time blocking segment with one year left out is added to avoid time leakage and enhance robustness verification.

7. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The targeted sampling strategy of S6 combines the coverage gap, the interval width and the residual hotspot in a weighted manner, and introduces sampling cost and traffic accessibility constraints to generate sampling priority, giving priority to sampling in areas with insufficient coverage and large errors, while ensuring the coverage ratio of different landform units.

8. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The suitability confidence linkage of S7 applies the suitability confidence decay coefficient to the indicator weights or thresholds of the multi-indicator evaluation to achieve dynamic adjustment of robust suitability partitions and outputs a robustness map to express the degree of suitability stability.

9. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The convergence criteria of S8 include: the endpoint error decreases by a preset percentage relative to the initial iteration, the overall improvement of the extrapolation distance-performance curve reaches a preset magnitude, and the coverage deviation is within the preset tolerance range of the target coverage. All of the above criteria are calculated within the spatial blocking cross-validation framework outside the one-leaving-one-topographic unit, and robustness is checked in conjunction with the time blocking caliber outside the one-leaving-one-year unit.

10. The method for predicting the spatial distribution of soil organic matter and pH in arid and hot valley regions based on machine learning according to claim 1, characterized in that, The output of the method includes: pixel-level spatial distribution map of soil organic matter and pH, uncertainty map, priority map of targeted replenishment and list of sampling points, robustness suitability zoning and robustness map. The output is exported in geographic raster data format and provides an inference service interface to support updates and reproduction.

Citation Information

Patent Citations

  • Soil organic matter content prediction method based on multi-source data

    CN118522366A

  • Cross-attention mechanism optimization method and system for multi-modal feature fine-grained alignment

    CN121456822A