Land assessment method, system, electronic device and computer storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]在上述实现过程中,构建多个差异化子模型并基于其预测概率动态确定集成权重进行加权融合,克服了单一模型对特定环境变量响应敏感、易产生系统性偏差的缺陷,提升了土地适宜性评估的精度与稳健性;同时,集成权重依据各子模型在不同空间单元上的预测表现自适应调整,使得融合过程能够有效捕捉环境变量与土地适宜性之间的复杂非线性关系,增强了模型对多源异质环境数据的适应性,从而获得更可靠的土地适宜性空间分布评估结果
[0013] In the above implementation process, the complex ecological responses of target land use types to different environmental gradients can be flexibly characterized, making the model more ecologically realistic. At the same time, the introduction of spatial random fields to represent spatial autocorrelation structure explicitly incorporates the interdependence between neighboring spatial units into the model, effectively solving the residual autocorrelation and biased estimation problems caused by neglecting spatial dependence in traditional non-spatial models, and improving the spatial continuity and rationality of the predicted probability. In addition, the output of the predicted probability of each spatial unit based on the Bayesian inference framework can quantify the prediction uncertainty while providing point estimation, providing richer information support for decision risk management in land suitability assessment.
Smart Images

Figure CN122549974A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of land appraisal technology, and more specifically, to a land appraisal method, system, electronic device, and computer storage medium. Background Technology
[0002] In the field of crop suitability assessment for saline-alkali land, mechanistic models, based on physiological limits, use rigid thresholds (such as critical soil salinity concentration and minimum water requirements) to depict a broad "basic ecological niche." In Northeast China, where hydrothermal conditions are favorable, these models often output "overestimated suitability" results, failing to reflect the nonlinear regulatory effects of local limiting factors such as salt accumulation and differences in micro-topography and drainage, leading to an overestimation of the area of developable land.
[0003] Data-driven models heavily rely on the spatial distribution of training sites and the quality of labels. Their results tend to pinpoint "realized niches" limited by human development history, existing irrigation and drainage facilities, and geographical barriers, rather than the actual potential distribution range of crops. In habitat-edge areas (such as transitional zones with moderate salinity and unstable water-salt dynamics), different models often produce logically divergent predictions (e.g., the confidence levels for the same area being classified as "suitable" and "unsuitable" are diametrically opposed), and these discrepancies lack systematic cross-validation and uncertainty quantification.
[0004] The aforementioned paradigms still have significant shortcomings in practical applications in saline-alkali land, most notably the logical imbalance between mechanistic explanatory power and statistical generalization. Although ensemble learning improves overall accuracy through multi-model combination, it typically employs strategies such as weighted voting or mean fusion, failing to address the fundamental contradictions in niche definitions among different paradigms at the mechanistic level. In other words, existing ensemble methods merely technically reconcile the output results, failing to construct a unified framework capable of simultaneously accommodating physiological constraints and real-world environmental filtering. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a land assessment method, system, electronic device and computer storage medium to improve the above-mentioned problems existing in the prior art.
[0006] In a first aspect, embodiments of this application provide a land assessment method, the land assessment method comprising: acquiring a spatial dataset of environmental variables for an area to be assessed, and environmental variable values of known distribution point samples and background point samples of a target land use type within the area to be assessed, and constructing a modeling sample set; constructing a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set, and obtaining the land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model for each spatial unit; determining integration weights based on the prediction performance indicators of the first sub-model, the second sub-model, and the third sub-model; and weighting and fusing the land suitability prediction probabilities based on the prediction performance indicators to obtain an integrated land suitability assessment result.
[0007] In the above implementation process, multiple differentiated sub-models are constructed and their prediction probabilities are dynamically determined for weighted fusion. This overcomes the shortcomings of single models being sensitive to specific environmental variables and prone to systematic bias, thus improving the accuracy and robustness of land suitability assessment. At the same time, the fusion weights are adaptively adjusted according to the prediction performance of each sub-model in different spatial units, enabling the fusion process to effectively capture the complex nonlinear relationship between environmental variables and land suitability. This enhances the model's adaptability to multi-source heterogeneous environmental data, thereby obtaining more reliable spatial distribution assessment results of land suitability.
[0008] Optionally, the step of constructing a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit includes: wherein the first sub-model is a discriminative ensemble learning model based on a tree model; taking the modeling sample set as input, and under the condition of training each decision tree, randomly selecting an equal number of samples from known distribution point samples and background point samples respectively; and outputting the first land suitability prediction probability of each spatial unit based on the first sub-model.
[0009] In the above implementation process, during the training of each decision tree, an equal number of samples from known distribution points and background points are randomly sampled to keep the number of samples of the two classes balanced. This effectively alleviates the class imbalance problem caused by the fact that the number of background point samples is much larger than the number of known distribution point samples, avoids model biased prediction, and thus significantly reduces the false negative rate and false negative risk of suitable areas in land suitability assessment. This allows the output land suitability prediction probability to more realistically reflect the actual suitability of each spatial unit.
[0010] Optionally, the step of constructing a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit further includes: wherein the second sub-model is a generative distribution model based on the maximum entropy principle; taking known distribution point samples and background point samples in the modeling sample set as input, using regularization and cross-validation to control model complexity, and estimating the potential spatial distribution of the target land use type under the condition of satisfying the environmental constraints of the known distribution point samples based on the maximum entropy principle; and outputting the second land suitability prediction probability of each spatial unit based on the second sub-model.
[0011] In the above implementation process, the environmental constraints of the known distribution point samples are modeled based on the maximum entropy principle, so that the probability estimate closest to the uniform distribution can be output when only limited distribution information is available. This avoids the bias introduced by subjective assumptions and ensures the robustness of the model under sparse data conditions. At the same time, the model complexity is controlled by regularization and cross-validation, which effectively suppresses the overfitting phenomenon caused by too many environmental variables or sample noise. This makes the output second land suitability prediction probability have better generalization ability and spatial continuity, thus more objectively reflecting the potential suitability spatial distribution pattern of the target land use type.
[0012] Optionally, the step of constructing a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit further includes: wherein the third sub-model is a Bayesian spatial hierarchy model that incorporates spatial random effects and a nonlinear response function; taking the modeling sample set as input, using a binomial response distribution and a link function, constructing a nonlinear response function of environmental variables through a random walk model, and characterizing the spatial autocorrelation structure through a spatial random field; and outputting the third land suitability prediction probability for each spatial unit based on the third sub-model.
[0013] In the above implementation process, the complex ecological responses of target land use types to different environmental gradients can be flexibly characterized, making the model more ecologically realistic. At the same time, the introduction of spatial random fields to represent spatial autocorrelation structure explicitly incorporates the interdependence between neighboring spatial units into the model, effectively solving the residual autocorrelation and biased estimation problems caused by neglecting spatial dependence in traditional non-spatial models, and improving the spatial continuity and rationality of the predicted probability. In addition, the output of the predicted probability of each spatial unit based on the Bayesian inference framework can quantify the prediction uncertainty while providing point estimation, providing richer information support for decision risk management in land suitability assessment.
[0014] Optionally, based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit, including: determining the optimal classification threshold of the first sub-model, the second sub-model, and the third sub-model respectively; performing segmented calibration on the prediction probabilities of the first sub-model, the second sub-model, and the third sub-model; and uniformly mapping the optimal classification threshold of the first sub-model, the second sub-model, and the third sub-model to the same preset benchmark value to obtain the land suitability prediction probability.
[0015] In the above implementation process, the optimal classification threshold of each sub-model is determined and segmented calibration is performed to overcome the problem that the predicted probabilities of different models cannot be directly compared due to differences in algorithm principles and probability output characteristics. This makes the probability outputs of each sub-model comparable under a unified scale. At the same time, the optimal classification threshold of each sub-model is uniformly mapped to the same preset benchmark value, ensuring that different models have a consistent judgment standard when making binary decisions on suitability. This eliminates the integration bias caused by threshold differences, thereby making the subsequent weighted fusion integrated land suitability assessment results more accurate and reliable.
[0016] Optionally, the step of determining the integration weights based on the prediction performance indicators of the first sub-model, the second sub-model, and the third sub-model, and weighting and fusing the land suitability prediction probabilities based on the prediction performance indicators to obtain the integrated land suitability assessment result, includes: determining the integration weights of the first sub-model, the second sub-model, and the third sub-model using the ability of each of the first sub-model, the second sub-model, and the third sub-model to distinguish between known distribution point samples and background point samples as the prediction performance indicator, wherein the sub-model with stronger distinguishing ability receives a higher weight; and weighting and summing the land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model based on the integration weights to obtain the integrated land suitability assessment result.
[0017] In the above implementation process, the actual ability of each sub-model to distinguish between known distribution point samples and background point samples is used as the basis for weight allocation. This allows the sub-model with better prediction performance to play a leading role in the integrated results, avoiding the drag effect of weak models on strong models in simple average integration, thereby improving the accuracy of integrated land suitability assessment. At the same time, the data-driven performance weight determination method does not rely on prior assumptions or expert experience, keeping the entire integration process objective and reproducible. It has good adaptability to different areas to be assessed and target land use types, and the output integrated assessment results are superior to single sub-models in terms of spatial robustness and ecological rationality.
[0018] Optionally, the land assessment method further includes: calculating the standard deviation of the calibrated land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model at the spatial unit scale, as the inter-model divergence, which is used to identify the spatial uncertainty of the integrated land suitability assessment results; wherein, if the standard deviation is less than a preset threshold, the assessment results of the spatial unit are determined to have a high degree of consensus; if the standard deviation is greater than the preset threshold, the spatial unit is determined to be a predicted risk area.
[0019] In the above implementation process, the standard deviation of the predicted probability after calibration of each sub-model is calculated to quantify the degree of divergence between models. This provides a spatial unit-by-spatial unit credibility identifier for the integrated evaluation results, making up for the shortcomings of traditional integration methods that only output a single mean result and cannot reflect the reliability of predictions. At the same time, based on a preset threshold, spatial units are divided into high consensus areas and prediction risk areas, so that high uncertainty areas with inconsistent judgments between models can be explicitly identified. This facilitates key review or supplementary investigation of risk areas in actual land planning and management, thereby reducing the risk of decision-making errors caused by blindly accepting the integrated results and improving the practicality and decision support value of land suitability assessment.
[0020] Secondly, embodiments of this application provide a land assessment system, which includes: a data acquisition module, a sub-model construction module, a probability calibration module, and an integrated assessment module; the data acquisition module is used to acquire a spatial dataset of environmental variables of the area to be assessed, as well as the environmental variable values of known distribution point samples and background point samples of the target land use type within the area to be assessed, and to construct a modeling sample set; the sub-model construction module is used to construct a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set, and to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. The probability calibration module is used to determine the optimal classification thresholds for the first sub-model, the second sub-model, and the third sub-model, respectively, to perform segmented calibration on the predicted probabilities of the first sub-model, the second sub-model, and the third sub-model, and to uniformly map the optimal classification thresholds of the first sub-model, the second sub-model, and the third sub-model to the same preset benchmark value to obtain the calibrated land suitability probability; the integrated evaluation module is used to determine the integration weights based on the prediction performance indicators of the first sub-model, the second sub-model, and the third sub-model, and to perform weighted fusion on the calibrated land suitability probability to obtain the integrated land suitability evaluation result.
[0021] In the above implementation process, through the coordinated operation of the data acquisition module, sub-model construction module, probability calibration module, and integration evaluation module, a complete land suitability assessment process from data preparation, multi-model construction, probability calibration to weighted fusion was achieved. The probability calibration module, by determining the optimal classification threshold for each sub-model and performing segmented calibration and unified threshold mapping, solved the problem of incomparable probability outputs caused by differences in algorithm principles between different models, making the predicted probabilities of each sub-model comparable at a unified scale. The integration evaluation module dynamically determines the weights based on the actual distinguishing ability of each sub-model for known distribution point samples and background point samples, allowing sub-models with better predictive performance to contribute more to the integration results. This avoids the problem of weak models dragging down strong models in simple average integration, thereby significantly improving the accuracy, robustness, and spatial reliability of land suitability assessment.
[0022] Thirdly, embodiments of this application also provide an electronic device, which includes a memory and a processor. The memory stores program instructions, and when the processor reads and runs the program instructions, it executes the steps in any of the above implementation methods.
[0023] Fourthly, embodiments of this application also provide a computer storage medium storing computer program instructions, which are read and executed by a processor to perform the steps in any of the above implementations. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a first schematic diagram of a land appraisal method provided in an embodiment of this application; Figure 2 This is a second schematic diagram of a land appraisal method provided in an embodiment of this application; Figure 3 This is a third schematic diagram of the land appraisal method provided in the embodiments of this application; Figure 4 This is a fourth schematic diagram of the land appraisal method provided in the embodiments of this application; Figure 5 The fifth schematic diagram of the land appraisal method provided in the embodiments of this application; Figure 6 The sixth schematic diagram of the land appraisal method provided in the embodiments of this application; Figure 7The seventh schematic diagram of the land appraisal method provided in the embodiments of this application; Figure 8 A schematic diagram of a land appraisal system provided in this application embodiment; Figure 9 The probability density curves and correlation matrices provided in the embodiments of this application; Figure 10 This is a block diagram of an electronic device provided in an embodiment of this application.
[0026] Icons: 010-Data acquisition module; 020-Sub-model construction module; 030-Probability calibration module; 040-Integrated evaluation module; 100-Electronic device; 111-Memory; 112-Memory controller; 113-Processor; 114-Peripheral interface; 115-Input / output unit; 116-Display unit. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.
[0028] In a first aspect, embodiments of this application provide a land assessment method applied to a server, which can be an electronic device with logical computing functions such as a personal computer (PC), tablet computer, smartphone, or personal digital assistant (PDA).
[0029] Please see Figure 1 , Figure 1 This is a first schematic diagram of a land assessment method provided in an embodiment of this application.
[0030] The land assessment method includes: acquiring a spatial dataset of environmental variables for the area to be assessed, as well as environmental variable values from known distribution point samples and background point samples of the target land use type within the area to be assessed, to construct a modeling sample set. Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively, obtaining the land suitability prediction probabilities for each spatial unit by the first, second, and third sub-models; determining the integration weights based on the prediction performance indicators of the first, second, and third sub-models; and weighting and fusing the land suitability prediction probabilities based on the prediction performance indicators to obtain the integrated land suitability assessment result.
[0031] In the above implementation process, spatial data of multi-source environmental variables such as climate, topography, and soil of the area to be evaluated are collected. Distribution point samples are extracted from known distribution records of the target land use type, and background point samples are randomly sampled from the entire area to be evaluated. A modeling sample set containing positive and negative samples is automatically constructed, reducing reliance on manual screening and ensuring the objectivity and reproducibility of the modeling process. Furthermore, by constructing three sub-models with complementary mechanisms in parallel (first, second, and third sub-models), the impact of single-model prediction bias on the evaluation results is reduced, overcoming the dependence of a single model on specific environmental response assumptions. The integration weight is dynamically determined by using the actual distinguishing ability of each sub-model between known distribution point samples and background point samples as the prediction performance index, allowing the better-performing sub-model to make a greater contribution to the fusion result, avoiding the problem of weak models dragging down strong models in simple average integration. The weighted fusion yields the integrated land suitability evaluation result, which combines the prediction advantages of the three heterogeneous sub-models, improving the spatial robustness and accuracy of the evaluation result. The output spatial distribution of land suitability is more consistent with the actual distribution pattern of the target land use type.
[0032] In one embodiment of this application, three sub-models with complementary mechanisms—the first sub-model, the second sub-model, and the third sub-model—characterize the response relationship between environmental variables and land suitability from three different perspectives: discriminative ensemble learning, maximum entropy distribution estimation, and spatial hierarchical Bayesian.
[0033] In one embodiment of this application, the known distribution point samples used by the first sub-model, the second sub-model, and the third sub-model can be kept consistent, while the background point samples can be determined according to the modeling mechanism and background environment coverage requirements of the first sub-model, the second sub-model, and the third sub-model, respectively.
[0034] In one embodiment of this application, a spatial dataset of environmental variables for the area to be evaluated, as well as environmental variable values of known distribution point samples and background point samples of the target land use type within the area to be evaluated, are obtained to construct a modeling sample set. The specific steps are as follows: Obtain spatial distribution samples of the target land use type within the area to be assessed. The target land use type is not limited to a specific crop or vegetation type, but rather refers to any land use pattern or cover type whose spatial distribution can be modeled using known distribution point samples and background point samples, and whose suitability can be assessed based on environmental variables. Accordingly, the environmental variables used in the modeling process can be flexibly adjusted according to the specific ecological needs and growth limiting factors of the target land use type.
[0035] Define the actual existing sample points of the target land use type as known distribution point samples, and assign them the following values: ,in, Indicates the first The category of each sample point is identified. Spatial sparsity and quality control processing are performed on the known distribution point samples to reduce the impact of spatial clustering, duplicate sampling, and outlier records on model training. Spatial sparsity can be achieved using a minimum distance constraint method, i.e., if the spatial distance between two known distribution point samples is less than a preset distance threshold, only one sample point is retained. The preset distance threshold is determined based on the scale of the area to be evaluated, the sample point density, and the spatial distribution characteristics of the target land use type, and is preferably 5 to 10 km.
[0036] The number of known distribution point samples after spatial sparsification and quality control is denoted as . ,in, The criteria are jointly determined by the distribution data of the target land use type, the scope of the area to be evaluated, the spatial sparsity threshold, and the sample point quality control conditions. By performing spatial sparsity and quality control processing on the known distribution point samples with minimum distance constraints, the artifact bias caused by spatial clustering and repeated sampling on model training is effectively eliminated. This allows the training samples to more evenly represent the actual distribution characteristics of the target land use type, avoiding overfitting and prediction bias caused by excessive concentration of local sample points.
[0037] Background point samples are extracted within the area to be evaluated to characterize the environmental background conditions in which the target land use type is not recorded as existing, and are assigned values. .
[0038] The sampling area for background points should be consistent with the suitability evaluation area for the target land use type, excluding marine areas, water bodies, invalid pixels, areas with missing environmental variables, and spatial areas not belonging to the study object. Background point samples can be obtained using spatial random sampling, stratified random sampling, environmental stratified sampling, or spatial uniform sampling. The number of background point samples is denoted as [sample number missing]. The number of background point samples is determined based on the area of the region to be evaluated, environmental heterogeneity, and the number of known distribution point samples. In other words, ,in, The ratio of background point samples to known distribution point samples is preferably set to a value of [value to be filled in]. When the area to be evaluated is large, the environmental gradient is complex, or the background environment is highly heterogeneous, This can be increased to 10 to 20; when the area to be evaluated is small, the number of known distribution point samples is small, or the background environment is relatively uniform, The number of background point samples can be appropriately reduced, and for cases where the number of known distribution point samples is small, a lower limit can be set to avoid insufficient environmental background representation due to insufficient background point samples. The number of background point samples can also be adjusted based on model validation performance, computational resources, and spatial sampling density.
[0039] Background point samples are randomly selected at an adjustable ratio within the area to be evaluated, and invalid areas are excluded, providing the model with comprehensive environmental background reference information. The flexible matching mechanism of the number of background point samples and known distribution point samples enables the modeling sample set to adapt to the evaluation needs under different scales and different environmental heterogeneity conditions, effectively balancing the class distribution of positive and negative samples.
[0040] The known distribution point samples and background point samples are merged to form a unified modeling sample set: ,in, Indicates the first Spatial coordinates of each sample point This represents the environmental variable vector corresponding to the sample point. The sample set is cleaned, including removing samples with missing coordinates, duplicate samples, samples with outdated category identifiers, and samples with missing environmental variables, and standardizing the sample coordinate system. The final modeling sample set is obtained after cleaning. ,in, This represents the number of effective samples after cleaning. Merging known distribution point samples with background point samples forms a unified modeling sample set, which is then cleaned in multiple dimensions. This ensures the integrity and consistency of the sample data, laying a high-quality data foundation for the stable training and reliable prediction of subsequent multi-sub-models.
[0041] Optionally, please refer to Figure 2 , Figure 2 This is a second schematic diagram of the land assessment method provided in the embodiments of this application.
[0042] Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. The process includes: the first sub-model is a discriminative ensemble learning model based on a tree model; taking the modeling sample set as input, and under the condition of training each decision tree, an equal number of samples are randomly drawn from the known distribution point samples and the background point samples respectively; and the first land suitability prediction probability of each spatial unit is output based on the first sub-model.
[0043] In the above implementation process, during the training of each decision tree, equal random sampling is performed on known distribution point samples and background point samples to maintain a balance between the two classes of samples in each tree. This alleviates the class imbalance problem caused by the fact that background point samples far outnumber known distribution point samples, and avoids the model's biased prediction towards the background class. Simultaneously, since each tree is trained independently based on different balancing subsamples, the ensembled results can more comprehensively capture the response patterns of the target land use type to environmental variables, reducing the false negative rate for suitable areas. The output first land suitability prediction probability can more realistically reflect the actual suitability of each spatial unit. Specifically, the output first land suitability prediction result can characterize the land suitability of each spatial unit relative to the background environment.
[0044] In one embodiment of this application, a first sub-model is used to characterize the nonlinear relationship and variable interactions between the existence state of the target land use type and environmental variables. The first sub-model can be a balanced random forest sub-model. The modeling sample set described above is used. The input is the response variable. The corresponding predictor variable is .in, The number of environmental variables involved in the modeling.
[0045] The first sub-model outputs the predicted probability of the first land suitability of the target land use type. In a more specific embodiment, this can be understood as the balanced random forest sub-model outputting the probability of the existence of the target land use type.
[0046] in, This represents the environmental variable vector of the spatial unit to be predicted. The spatial unit to be predicted may include the sample points or spatial pixels to be predicted.
[0047] To reduce the impact of the imbalance between the number of known distribution point samples and background point samples on model training, this application adopts a class-balanced sampling strategy during the construction of each decision tree.
[0048] Let the number of samples from the known distribution points be... The number of background point samples is The number of samples for each class in each tree is:
[0049] in, This represents the actual number of samples taken from each category in each tree. This indicates the maximum number of samples that can be drawn from each category for each tree.
[0050] The actual number of known distribution point samples and background point samples extracted for each tree are as follows:
[0051]
[0052] This ensures that the number of known distribution point samples and background point samples in the training samples of each tree are equal.
[0053] In another preferred embodiment, a maximum sampling number for each category of samples per tree can also be set. Then the number of samples for each category in each tree is:
[0054] in, The determination is based on sample size, number of minority class samples, model performance stability, and computational resources. It is not used as a universal fixed value, but is determined through model performance evaluation under specific land use type, specific area to be evaluated and specific sample size. In this example, it is determined to be 900.
[0055] In a preferred embodiment, it is possible to or The model performance is compared using multiple candidate values, and TSS, AUC, or other evaluation metrics are calculated using an independent validation sample set different from the training samples. When the model performance improvement tends to stabilize as the sampling ratio or sample size increases, while the computational cost continues to increase, the corresponding sampling ratio or sample size is determined as the preferred value in this implementation. This dynamic balancing sampling strategy can avoid the model biasing towards the background class when the number of background point samples is much larger than the number of known distribution point samples, while maintaining the difference in sample perturbation between each tree, thus improving the stability and generalization ability of the random forest model.
[0056] In a preferred embodiment, the first sub-model is a balanced random forest model, which is implemented using ranger, with the following parameter settings:
[0057] When the number of environment variables is 15, mtry is preferably:
[0058] After training, the balanced random forest sub-model outputs the first land suitability probability result:
[0059] in, The outcome variable is the land suitability prediction probability of the first sub-model (balanced random forest model), which is the land suitability prediction probability value output by the first sub-model for each spatial unit after training. The value of this variable ranges from [0,1]. The closer the value is to 1, the more suitable the spatial unit is for the target land use type, and the closer the value is to 0, the less suitable it is.
[0060] in, is the prediction function of the first sub-model, representing the predicted probability of land suitability of the target land use type in the spatial unit, calculated by a trained balanced random forest model with the environmental variable vector x of the spatial unit to be predicted as input.
[0061] in, This means assigning the output of the prediction function of the first sub-model for spatial unit x to the variable. This serves as input data for subsequent probability calibration and integrated evaluation.
[0062] Optionally, please refer to Figure 3 , Figure 3 This is a third schematic diagram of the land appraisal method provided in the embodiments of this application.
[0063] Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. The model also includes: the second sub-model is a generative distribution model based on the maximum entropy principle; it takes known distribution point samples and background point samples from the modeling sample set as input, and uses regularization and cross-validation to control model complexity; based on the maximum entropy principle, it estimates the potential spatial distribution of the target land use type's suitability under the environmental constraints of the known distribution point samples; and it outputs the second land suitability prediction probability for each spatial unit based on the second sub-model.
[0064] In the above implementation process, environmental constraints of known distribution point samples are modeled based on the maximum entropy principle. With only limited distribution information, it can output a probability estimate closest to a uniform distribution, avoiding biases introduced by subjective assumptions and ensuring the objectivity and robustness of the model under data sparsity conditions. Simultaneously, regularization and cross-validation control the model complexity, suppressing overfitting caused by too many environmental variables or sample noise. The output second land suitability prediction probability has better generalization ability and spatial continuity, thus more realistically reflecting the potential spatial distribution pattern of the target land use type's suitability.
[0065] In one embodiment of this application, a second sub-model is used to estimate the potential spatial distribution of suitability for a target land use type based on known distribution point samples and background point samples. This model, using the maximum entropy principle, infers the predicted land suitability probability of the target land use type within the area to be evaluated, while satisfying the environmental constraints of the known distribution point samples. The second sub-model can be a maximum entropy sub-model.
[0066] The input to the maximum entropy submodel includes a known set of distribution point samples:
[0067] And the background point sample set:
[0068] The maximum entropy submodel output is:
[0069] In a preferred embodiment, the maximum entropy model adopts a known distribution point sample-background point sample modeling framework, the regularization factor is set to 1.0, and 10-fold cross-validation is used to reduce the risk of model overfitting.
[0070] The number of background point samples used by the maximum entropy model can be the same as the number of background point samples in the unified modeling sample set, or it can be set separately according to the maximum entropy model's requirements for background environment coverage. Preferably, the number of background point samples for the maximum entropy model is 10,000; when the number of background point samples in the unified modeling sample set is less than 10,000, additional sampling can be performed in the background environment of the area to be evaluated to bring it up to 10,000; when the area to be evaluated is small or the background environment is relatively uniform, the aforementioned proportional rule can be used to determine the number of background point samples.
[0071] The parameters for the maximum entropy submodel are set as follows:
[0072] After training, the maximum entropy submodel outputs the second land suitability probability result:
[0073] Where x is the environmental variable vector corresponding to a certain spatial unit to be predicted, that is, the combination of values of the spatial unit on each environmental variable. Based on the second sub-model and the maximum entropy principle, the output is the predicted probability value of the spatial unit belonging to the suitable area of the target land use type, given the input environmental variable vector x. It is a continuous numerical value with a range of [0,1]. The second land suitability prediction probability result is obtained by predicting all spatial units after the second sub-model has been trained. The spatial distribution data is typically stored in the form of a raster layer, with each cell value representing the second land suitability prediction probability for the corresponding spatial unit.
[0074] Optionally, please refer to Figure 4 , Figure 4 The fourth schematic diagram illustrates the land valuation method provided in this application embodiment.
[0075] Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively to obtain the land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. The third sub-model is a Bayesian spatial hierarchy model that incorporates spatial random effects and a nonlinear response function. Using the modeling sample set as input, a binomial response distribution and link function are adopted, and a nonlinear response function of environmental variables is constructed through a random walk model. The spatial autocorrelation structure is characterized through a spatial random field. Based on the third sub-model, the third land suitability prediction probability for each spatial unit is output.
[0076] In the above implementation process, the random walk model constructs a nonlinear response function for environmental variables, breaking through the pre-defined limitations of traditional parametric models on the shape of response curves. It can flexibly characterize the complex ecological responses of target land use types to different environmental gradients, making the model more ecologically realistic. At the same time, the introduction of spatial random fields to represent spatial autocorrelation structures explicitly incorporates the interdependencies between neighboring spatial units into the model, effectively solving the residual autocorrelation and biased estimation problems caused by neglecting spatial dependencies in traditional non-spatial models, and significantly improving the spatial continuity and rationality of prediction probabilities. In addition, based on the Bayesian inference framework, the model outputs the third land suitability prediction probability for each spatial unit, which can quantify prediction uncertainty while providing point estimates, providing richer information support for decision-making risk management in land suitability assessment.
[0077] In one embodiment of this application, the third sub-model is a Bayesian spatial hierarchy model that incorporates spatial random effects and nonlinear response functions to characterize the spatial autocorrelation structure and nonlinear characteristics of environmental response in the distribution of target land use types.
[0078] Preferably, the third sub-model is constructed based on the INLA-SPDE framework. This model adopts a binomial response structure:
[0079] in, Indicates the first The relative probability that a sample belongs to the target land use type with respect to the background environment and is of a known distribution category.
[0080] The model is built using the logit join function:
[0081] in, For the intercept term; For the first Nonlinear response function of one environmental variable; This is the spatial random effects term; Spatial coordinates of a spatial unit; These are the standardized values of the environmental variables.
[0082] The formula for standardizing environmental variables is:
[0083] in, and The first The mean and standard deviation of each environmental variable in the training samples.
[0084] The preferred method for the nonlinear response function of environmental variables is the second-order random walk model RW2.
[0085] Spatial random effects term The SPDE method is preferred for constructing Matérn spatial random fields to characterize spatial structures beyond environmental variables.
[0086] Before constructing the SPDE spatial grid, it is preferable to process the spatial coordinates according to the extent of the region to be evaluated. For global scales or large areas, it is preferable to convert latitude and longitude coordinates into appropriate equal-area or equidistant projected coordinates, or to use zonal projection, spherical distance approximation, or other methods suitable for large-scale spatial modeling, to reduce spatial distortion caused by directly using Euclidean distances of latitude and longitude. The units of the mesh parameters and spatial prior parameters should be consistent with the units of the spatial coordinates used.
[0087] The parameters for the third sub-model are set as follows:
[0088] The above parameters are the parameter settings in the preferred embodiment, and can be adjusted according to the scale of the area to be evaluated, coordinate system, sample point density, spatial correlation structure, and computing resources.
[0089] After training, the third sub-model's probability of predicting land suitability is as follows:
[0090] in, This represents the third land suitability prediction probability value, which is the final output of the third sub-model. It is the land suitability prediction probability result calculated for each spatial unit after training and inference of the Bayesian spatial hierarchy model, and is used for subsequent model integration and weighted fusion.
[0091] in, This represents the land suitability prediction probability function output by the third sub-model under the given environmental variable vector x. Here, x is the environmental variable vector of the spatial unit to be predicted. This function, based on the INLA-SPDE framework, obtains the posterior average probability value through joint estimation using the logit link function, the RW2 nonlinear response function, and the Matérn spatial random field.
[0092] It's not hard to understand. and In this application, "equivalence" refers to the third land suitability prediction probability output by the third sub-model for each spatial unit, and the two can be used interchangeably.
[0093] In one embodiment of this application, the balanced random forest model of the first sub-model, the maximum entropy model of the second sub-model, and the Bayesian spatial hierarchy model of the third sub-model are applied to the spatial environment data of the study area to obtain three single-model suitability probability maps:
[0094] Before integration, the outputs of the three individual models are spatially consistent to ensure they have a consistent coordinate reference system, spatial extent, spatial resolution, pixel position, and invalid value mask.
[0095] Preferably, one of the model outputs is selected as the reference grid, and the outputs of the other models are resampled to the reference grid so that the prediction probabilities of the three models can correspond to each pixel.
[0096] Optionally, please refer to Figure 5 , Figure 5 The fifth schematic diagram illustrates the land valuation method provided in this application embodiment.
[0097] Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. This includes: determining the optimal classification thresholds for the first sub-model, the second sub-model, and the third sub-model respectively; performing segmented calibration on the prediction probabilities of the first sub-model, the second sub-model, and the third sub-model; and uniformly mapping the optimal classification thresholds of the first sub-model, the second sub-model, and the third sub-model to the same preset benchmark value to obtain the land suitability prediction probability.
[0098] In the above implementation process, the optimal classification threshold for each sub-model was determined and segmented calibration was performed. This overcame the problem that the predicted probabilities of different sub-models could not be directly compared due to differences in algorithm principles and probability output characteristics, making the probability outputs of each sub-model comparable under a unified scale. Mapping the optimal classification threshold of each sub-model to the same preset benchmark value ensured that different sub-models had a consistent discrimination standard when making binary judgments of suitability, eliminating integration bias caused by threshold differences. This resulted in more accurate and reliable integrated land suitability assessment results after subsequent weighted fusion.
[0099] In one embodiment of this application, since the probability output mechanisms of different sub-models are different, their original predicted probabilities do not have completely consistent discriminative meanings. Therefore, it is necessary to determine the optimal classification threshold for each sub-model separately.
[0100] For the Each sub-model calculates sensitivity and specificity under different candidate thresholds based on the predicted probabilities of known distribution point samples and background point samples, and uses the maximum TSS criterion to determine the optimal classification threshold.
[0101] in, For the first The optimal classification threshold for each sub-model.
[0102] The formula for calculating TSS is:
[0103] in:
[0104]
[0105] TP represents the number of true positives, TN represents the number of true negatives, FP represents the number of false positives, and FN represents the number of false negatives.
[0106] In one embodiment of this application, the optimal classification threshold can be dynamically determined based on the target crop, study area, sample composition, and prediction results of each sub-model in the target land use type. Under different crops, different study areas, or different sample sizes, the optimal classification threshold of each sub-model should be recalculated.
[0107] Optionally, please refer to Figure 6 , Figure 6 The sixth schematic diagram illustrates the land valuation method provided in this application embodiment.
[0108] The integration weights are determined based on the prediction performance indicators of the first, second, and third sub-models. The land suitability prediction probabilities are then weighted and fused based on these performance indicators to obtain the integrated land suitability assessment results. This includes: using the ability of each of the first, second, and third sub-models to distinguish between known distribution point samples and background point samples as the prediction performance indicator, determining the integration weights of the first, second, and third sub-models, where the sub-model with stronger distinguishing ability receives a higher weight; and then, based on these integration weights, the land suitability prediction probabilities of the first, second, and third sub-models are weighted and summed to obtain the integrated land suitability assessment results.
[0109] In the above implementation process, the actual distinguishing ability of each sub-model between known distribution point samples and background point samples is used as the basis for weight allocation. This allows the sub-model with better predictive performance to play a leading role in the integrated result, avoiding the drag from weak models on strong models in simple average integration, thereby improving the accuracy of integrated land suitability assessment. Furthermore, the data-driven performance weight determination method does not rely on prior assumptions or expert experience, keeping the entire integration process objective and reproducible. It has good adaptability to different assessment areas and target land use types, and the output integrated land suitability assessment results are superior to those of single sub-models in terms of spatial robustness and ecological rationality.
[0110] In one embodiment of this application, to eliminate probabilistic scale differences between different sub-models, the original predicted probabilities of each sub-model are threshold-calibrated. For the first... There are several sub-models, and their original prediction probabilities are assumed to be... The optimal classification threshold is The probability after calibration is ,but:
[0111] The calibration function maps the optimal classification threshold of each sub-model to 0.5, ensuring that all sub-models have the same optimal classification threshold. These represent the suitable / unsuitable discrimination boundaries. The calibration process maintains the monotonicity of the predicted probabilities of each sub-model and does not change the relative ordering of samples or pixels within the same sub-model.
[0112] In one embodiment of this application, after probability calibration is completed, the TSS value of each sub-model is used as the model performance weight. The formula for calculating the weights of each sub-model is as follows:
[0113] in, The number of sub-models participating in the integration.
[0114] When the TSS of a certain sub-model is less than or equal to 0, its weight can be set to 0, and the weights of the remaining sub-models can be renormalized:
[0115] The weights are calculated based on the TSS values of each sub-model on the corresponding sample set.
[0116] In specific implementations, the TSS weights of each sub-model should be recalculated under different target crops, study areas, sample compositions, or environmental variable systems.
[0117] Optionally, please refer to Figure 7 , Figure 7 The seventh schematic diagram illustrates the land valuation method provided in this application embodiment.
[0118] The land assessment method also includes: calculating the standard deviation of the predicted land suitability probabilities after calibration of the first, second, and third sub-models at the spatial unit scale, as the inter-model divergence degree, which is used to identify the spatial uncertainty of the integrated land suitability assessment results; wherein, if the standard deviation is less than a preset threshold, the assessment results of the spatial unit are determined to have a high degree of consensus; if the standard deviation is greater than the preset threshold, the spatial unit is determined to be a predicted risk area.
[0119] In the above implementation process, the standard deviation of the predicted probabilities after calibration of each sub-model is calculated to quantify the degree of divergence between models. This provides a spatial unit-by-spatial-unit credibility identifier for the integrated evaluation results, overcoming the shortcomings of traditional integration methods that only output a single mean result and cannot reflect the reliability of predictions. Based on a preset threshold, spatial units are divided into high consensus zones and prediction risk zones, enabling the explicit identification of high-uncertainty areas with inconsistent judgments between models. This facilitates focused review or supplementary investigation of risk zones in actual land planning and management, thereby reducing the risk of decision-making errors caused by blindly accepting the integrated results and enhancing the practicality and decision support value of land suitability assessment.
[0120] In one embodiment of this application, to characterize the consistency of different sub-models' suitability judgments for the same spatial unit, the pixel-level standard deviation of the predicted probabilities of the three calibrated sub-models is calculated to obtain the inter-model divergence:
[0121] in, Represents a cell The inter-model divergence is indicated by a smaller value, which means that the different sub-models are more consistent in their assessment of the suitability of the pixel; the larger the value, the stronger the predictive divergence between the different sub-models.
[0122] The final output includes:
[0123] and:
[0124] in, The integrated probability of target crop suitability. This represents the degree of divergence between models.
[0125] In a specific embodiment of this application, the target crop for the above process is oil palm, and the study area is the global land area between 25° north and south latitude. Samples of known distribution points of oil palm are obtained and subjected to spatial sparsification and quality control processing. The spatial sparsification distance threshold is set to 10 km. After cleaning, samples of known distribution points are obtained:
[0126] Background point samples were obtained within the land mask area of the study area using stratified random sampling, and then cleaned to obtain the following background point samples:
[0127] The ratio of background point samples to known distribution point samples is:
[0128] This ratio falls within the preferred range where the number of background point samples is 5–10 times the number of known distribution point samples.
[0129] Construct a balanced random forest model, a maximum entropy model, and a Bayesian spatial hierarchy model. The balanced random forest model is implemented using a ranger, with probability=TRUE, num.trees=1000, mtry=4, min.node.size=5, splittrule="gini", and importance="permutation". The maximum number of samples per class per tree is set as follows:
[0130] The actual sample size for each category in each tree is:
[0131] Substitute the sample size into this implementation method:
[0132] Therefore, the balanced random forest model extracts 900 known distribution point samples and 900 background point samples in each tree.
[0133] The maximum entropy model employs a known distribution point sample-background point sample framework, with a regularization factor of 1.0, 10-fold cross-validation, and an optimal number of background point samples of 10,000. The Bayesian spatial hierarchy model uses the INLA-SPDE framework, with the response variable following a Bernoulli / binomial distribution, a logit link function, an RW2 nonlinear response function for environmental variables, and an SPDE Matérn spatial field for spatial random effects.
[0134] The three sub-models each output a suitability probability:
[0135] After spatially unifying the outputs of the three sub-models, the optimal classification threshold for each sub-model is determined based on the maximum TSS criterion. In this embodiment, the optimal classification thresholds for the three sub-models are as follows:
[0136]
[0137]
[0138] The above thresholds are mapped to a unified calibration benchmark of 0.5, i.e.:
[0139] The calibration formula is:
[0140] Subsequently, the ensemble weights are calculated based on the TSS values of the three sub-models. In this embodiment, the weights of the three sub-models are as follows:
[0141]
[0142]
[0143] Constructing an integrated crop suitability model:
[0144] Substituting the weights in this implementation, we get:
[0145] Further calculation of the divergence between models:
[0146] in, For the suitability integration probability of oil palm, This represents the degree of divergence between models.
[0147] The specific values mentioned above are merely examples of parameters used in this embodiment. For other target crops, other study areas, or other sample sizes, the number of known distribution point samples, the number of background point samples, the upper limit of random forest equilibrium sampling, the optimal classification threshold of the sub-model, and the TSS weights can all be re-determined based on the corresponding data.
[0148] Secondly, this application provides a land appraisal system, please refer to [link to relevant documentation]. Figure 8 , Figure 8 A schematic diagram of a land appraisal system provided for an embodiment of this application.
[0149] The land assessment system includes: a data acquisition module 010, a sub-model construction module 020, a probability calibration module 030, and an integrated assessment module 040. The data acquisition module 010 acquires the spatial dataset of environmental variables for the area to be assessed, as well as the environmental variable values of known distribution point samples and background point samples of the target land use type within the area to be assessed, thus constructing a modeling sample set. The sub-model construction module 020, based on the modeling sample set, constructs a first sub-model, a second sub-model, and a third sub-model, respectively, to obtain the land suitability predictions for each spatial unit from the first, second, and third sub-models. The probability calibration module 030 is used to determine the optimal classification thresholds for the first, second, and third sub-models, respectively, to perform segmented calibration on the predicted probabilities of the first, second, and third sub-models, and to uniformly map the optimal classification thresholds of the first, second, and third sub-models to the same preset benchmark value to obtain the calibrated land suitability probabilities; the integrated evaluation module 040 is used to determine the integration weights based on the prediction performance indicators of the first, second, and third sub-models, and to perform weighted fusion on the calibrated land suitability probabilities to obtain the integrated land suitability evaluation results.
[0150] In the above implementation process, the data acquisition module 010 automatically extracts known distribution point samples and background point samples from multi-source environmental variable spatial data to construct a modeling sample set. This eliminates the need for manual sample selection, ensuring the objectivity and reproducibility of the modeling process. The sub-model construction module 020 constructs three complementary sub-models in parallel, characterizing the response relationship between environmental variables and land suitability from three different perspectives: discriminative ensemble learning, maximum entropy distribution estimation, and spatial hierarchical Bayesian methods. This allows the prediction biases of each sub-model to cancel each other out at different spatial locations. The probability calibration module 030 determines the optimal classification threshold for each sub-model and performs segmented calibration and unified threshold mapping. This solves the problem of incomparable probability outputs caused by differences in algorithm principles among different sub-models, making the prediction probabilities of each sub-model comparable at a unified scale. The integrated evaluation module 040 dynamically determines the weights based on the actual distinguishing ability of each sub-model on known distribution point samples and background point samples, so that the sub-model with better prediction performance can make a greater contribution to the integrated results. This avoids the problem of weak models dragging down strong models in simple average integration, thereby significantly improving the accuracy, robustness and spatial reliability of land suitability assessment.
[0151] In one embodiment of this application, the original importance scores of the random forest, maximum entropy, and Bayesian hierarchical models are extracted separately. Weights are assigned based on the evaluation performance of each sub-model, and an integrated importance index is constructed through weighted averaging. Permutation importance is used to assess the contribution of variables to the model's discriminative accuracy, while gain importance reflects the dominant position of variables in spatial partitioning. A crop maximum potential extraction method is employed to characterize the physiological response of agricultural development to key factors. High-intensity spatial sampling (n=500,000) is conducted within the study area to establish the association between predicted probabilities and environmental gradients, dividing the full environmental range into 120 to 200 intervals, and extracting the 95th percentile of the suitability probability within each interval. This method minimizes the limiting influence of non-target factors, thereby identifying the theoretical maximum habitat potential defined by the variable. For variables with explicit physical constraints (such as altitude), monotonicity constraints are introduced to eliminate statistical noise at extreme values. A generalized additive model (GAM) with cubic contraction splines is used to smooth the response trajectory, ensuring that the response curve is both biologically logical and statistically robust.
[0152] In one embodiment of this application, please refer to Figure 9 , Figure 9 The probability density curves and correlation matrices provided in the embodiments of this application.
[0153] Figure 9The probability density curves on the left intuitively reveal the sensitivity of different models to habitat discrimination. The random forest model exhibits an extreme pinpoint distribution (highly concentrated around 1.0) at known sample distribution points, indicating a significant risk of overfitting; while the Bayesian hierarchical model is too conservative, with a large overlap between the two classes of samples in the 0.2–0.5 range, resulting in insufficient discrimination.
[0154] The ensemble model exhibits the most robust bimodal structure. The peak values of the background sample points and the known sample distribution points are clearly separated around 0.05 and 0.8, respectively, and the curves intersect precisely at 0.5. Compared to a single model, the ensemble strategy effectively smooths out the overconfidence of the random forest model and widens the ambiguity that the Bayesian hierarchical model struggles to distinguish, thus achieving a higher signal-to-noise ratio and discriminant consistency at the statistical level.
[0155] Figure 9 The correlation matrix on the right shows that all models exhibit high consistency in spatial trends, but significant differences exist in algorithmic paradigm-driven approaches. The Random Forest model and the Maximum Entropy model show the highest correlation (r = 0.924), indicating that they share similar logic in capturing dominant environmental factors. The Bayesian hierarchical model shows relatively low correlation with other models (only 0.782 with the Random Forest model). This moderate range of correlation (0.77-0.92) validates the complementarity of various statistical paradigms in geospatial prediction. The ensemble model maintains extremely high correlations with each sub-model, all greater than 0.84, and as high as 0.976 with the Random Forest model, demonstrating that the ensemble model successfully integrates the spatial characteristics of multiple algorithms without producing anomalous predictions that deviate from geographical patterns.
[0156] Optionally, please refer to Figure 10 , Figure 10 This is a block diagram illustrating an electronic device according to an embodiment of this application. The electronic device 100 may include a memory 111, a memory controller 112, a processor 113, a peripheral interface 114, an input / output unit 115, and a display unit 116. Those skilled in the art will understand that... Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device 100. For example, the electronic device 100 may also include components that are more... Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown.
[0157] The aforementioned memory 111, memory controller 112, processor 113, peripheral interface 114, input / output unit 115, and display unit 116 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The aforementioned processor 113 is used to execute executable modules stored in the memory.
[0158] The memory 111 can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory 111 stores programs. After receiving execution instructions, the processor 113 executes the programs. The methods executed by the electronic device 100 as defined in any embodiment of this application can be applied to the processor 113, or implemented by the processor 113.
[0159] The aforementioned processor 113 may be an integrated circuit chip with signal processing capabilities. The processor 113 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a digital signal processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor.
[0160] The peripheral interface 114 described above couples various input / output devices to the processor 113 and the memory 111. In some embodiments, the peripheral interface 114, the processor 113, and the memory controller 112 can be implemented on a single chip. In other instances, they can be implemented on separate chips.
[0161] The input / output unit 115 described above is used to provide user input data. The input / output unit 115 can be, but is not limited to, a mouse and a keyboard.
[0162] The aforementioned display unit 116 provides an interactive interface (e.g., a user interface) between the electronic device 100 and the user, or displays image data for the user's reference. In this embodiment, the display unit can be a liquid crystal display (LCD) or a touch display. If it is a touch display, it can be a capacitive touchscreen or a resistive touchscreen that supports single-point and multi-point touch operations. Supporting single-point and multi-point touch operations means that the touch display can sense touch operations generated simultaneously from one or more locations on the touch display and pass the sensed touch operations to the processor for calculation and processing.
[0163] This application also provides a computer access storage medium storing computer program instructions. When the computer program instructions are read and executed by a processor, the steps in the land assessment method are performed.
[0164] In summary, this application provides a land assessment method, system, electronic device, and computer storage medium, relating to the field of land assessment technology. It involves acquiring a spatial dataset of environmental variables for the area to be assessed, along with environmental variable values from known distribution point samples and background point samples of the target land use type within the area, to construct a modeling sample set. Based on this sample set, a first sub-model, a second sub-model, and a third sub-model are constructed, yielding the land suitability prediction probabilities for each spatial unit. Integration weights are determined based on the prediction performance indicators of the first, second, and third sub-models. The land suitability prediction probabilities are then weighted and fused based on the prediction performance indicators to obtain the integrated land suitability assessment result. Through weighted integration, a single algorithm with heterogeneous mathematical logic is deeply integrated, overcoming the prediction bias of individual models in complex habitats, thus characterizing the distribution of suitability for agricultural development in saline-alkali land.
[0165] In the several embodiments provided in this application, it should be understood that the disclosed device can also be implemented in other ways. The device embodiments described above are merely illustrative; for example, the block diagrams in the accompanying drawings illustrate the possible architecture, functions, and operations of the device according to various embodiments of this application. In this regard, each block in the block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram, and combinations of block diagrams, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0166] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0167] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0168] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0169] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
[0170] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
Claims
1. A land valuation method, characterized in that, The land valuation methods include: Obtain the spatial dataset of environmental variables for the area to be evaluated, as well as the environmental variable values of known distribution point samples and background point samples of the target land use type within the area to be evaluated, and construct a modeling sample set; Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively, and the land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model for each spatial unit are obtained; The integration weights are determined based on the prediction performance indicators of the first sub-model, the second sub-model, and the third sub-model; the land suitability prediction probabilities are weighted and fused based on the prediction performance indicators to obtain the integrated land suitability assessment results.
2. The method according to claim 1, characterized in that, Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively, and the land suitability prediction probabilities for each spatial unit by the first sub-model, the second sub-model, and the third sub-model are obtained. include: The first sub-model is a discriminative ensemble learning model based on a tree model; Using the modeling sample set as input, in the case of training each decision tree, an equal number of samples are randomly drawn from the known distribution point samples and the background point samples respectively; The first land suitability prediction probability for each spatial unit is output based on the first sub-model.
3. The method according to claim 1, characterized in that, Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively. The land suitability prediction probabilities for each spatial unit are obtained using the first sub-model, the second sub-model, and the third sub-model. include: The second sub-model is a generative distribution model based on the maximum entropy principle; Using known distribution point samples and background point samples from the modeling sample set as input, regularization and cross-validation are used to control model complexity. Based on the maximum entropy principle, the potential suitability spatial distribution of the target land use type is estimated under the condition of satisfying the environmental constraints of the known distribution point samples. The second sub-model outputs the second land suitability prediction probability for each spatial unit.
4. The method according to claim 1, characterized in that, Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively. The land suitability prediction probabilities for each spatial unit are obtained using the first sub-model, the second sub-model, and the third sub-model. include: The third sub-model is a Bayesian spatial hierarchy model that incorporates spatial random effects and nonlinear response functions. Using the modeling sample set as input, a nonlinear response function of the environmental variable is constructed through a random walk model by employing a binomial response distribution and a link function, and the spatial autocorrelation structure is characterized by a spatial random field. Based on the third sub-model, the third land suitability prediction probability of each spatial unit is output.
5. The method according to claim 1, characterized in that, Based on the modeling sample set, a first sub-model, a second sub-model, and a third sub-model are constructed respectively, and the land suitability prediction probabilities for each spatial unit by the first sub-model, the second sub-model, and the third sub-model are obtained, including: Determine the optimal classification thresholds for the first sub-model, the second sub-model, and the third sub-model, respectively; The predicted probabilities of the first sub-model, the second sub-model, and the third sub-model are calibrated in segments. The optimal classification thresholds of the first sub-model, the second sub-model, and the third sub-model are uniformly mapped to the same preset benchmark value to obtain the land suitability prediction probability.
6. The method according to claim 1, characterized in that, The integration weights are determined based on the prediction performance metrics of the first sub-model, the second sub-model, and the third sub-model. The land suitability prediction probabilities are weighted and fused based on the predicted performance indicators to obtain integrated land suitability assessment results, including: The ability of each of the first sub-model, the second sub-model, and the third sub-model to distinguish between known distribution point samples and background point samples is used as the prediction performance index to determine the ensemble weight of the first sub-model, the second sub-model, and the third sub-model, wherein the sub-model with stronger distinguishing ability receives a higher weight. Based on the integrated weights, the land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model are weighted and summed to obtain the integrated land suitability assessment results.
7. The method according to claim 1, characterized in that, The land valuation method also includes: The standard deviation of the calibrated land suitability prediction probabilities of the first sub-model, the second sub-model, and the third sub-model at the spatial unit scale is calculated as the inter-model divergence degree, which is used to identify the spatial uncertainty of the integrated land suitability assessment results. Specifically, if the standard deviation is less than a preset threshold, the evaluation results of the spatial unit are considered to have a high degree of consensus; if the standard deviation is greater than the preset threshold, the spatial unit is considered to be a predicted risk area.
8. A land valuation system, characterized in that, The land assessment system includes: a data acquisition module, a sub-model construction module, a probability calibration module, and an integrated assessment module; The data acquisition module is used to acquire the spatial dataset of environmental variables of the area to be evaluated, as well as the environmental variable values of known distribution point samples and background point samples of the target land use type in the area to be evaluated, and to construct a modeling sample set. The sub-model construction module is used to construct a first sub-model, a second sub-model, and a third sub-model based on the modeling sample set, and to obtain the land suitability prediction probability of the first sub-model, the second sub-model, and the third sub-model for each spatial unit. The probability calibration module is used to determine the optimal classification thresholds of the first sub-model, the second sub-model, and the third sub-model respectively, to perform segmented calibration on the predicted probabilities of the first sub-model, the second sub-model, and the third sub-model, and to uniformly map the optimal classification thresholds of the first sub-model, the second sub-model, and the third sub-model to the same preset benchmark value to obtain the calibrated land suitability probability. The integrated evaluation module is used to determine the integration weights based on the prediction performance indicators of the first sub-model, the second sub-model, and the third sub-model, and to perform weighted fusion on the calibrated land suitability probabilities to obtain the integrated land suitability evaluation results.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores program instructions, and when the processor executes the program instructions, it performs the steps of the method according to any one of claims 1-7.
10. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, perform the steps of the method according to any one of claims 1-7.