Method and system for analyzing influencing factors of soil organic matter in cane field based on multi-source data
By using a multi-source data-based method to analyze the influencing factors of soil organic matter in sugarcane areas, and combining random forest regression and SHAP models with the PLS-SEM algorithm, nonlinear driving factors of soil organic matter are identified. This solves the problem of identifying nonlinear factors in existing technologies, achieves transparent interpretation and high-precision prediction, and supports the scientific management of agricultural ecosystems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUGARCANE RES INST OF YUNNAN ACADEMY OF AGRI SCI
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies rely mainly on linear regression and machine learning models to identify factors influencing soil organic matter. However, these models cannot effectively identify nonlinear driving factors, and the decision-making process of machine learning models is opaque and lacks a reasonable explanation.
We adopted a multi-source data-based method to analyze the influencing factors of soil organic matter in sugarcane areas. By obtaining the spatial distribution of soil sampling points, we calculated the spatial autocorrelation index, constructed a random forest regression model, used SHAP values for interpretation, drew a SHAP dependency graph, and combined the PLS-SEM algorithm to quantify direct and indirect effects and verify the importance ranking of SHAP.
The nonlinear driving factors and their interaction mechanisms of soil organic matter have been clarified, providing a transparent explanation mechanism, improving prediction accuracy and the credibility of conclusions, supporting differentiated management and precision fertilization, and applicable to the diagnosis and regulation of complex agricultural ecosystems.
Smart Images

Figure CN122117149A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of soil analysis technology, and in particular to a method and system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data. Background Technology
[0002] Yunnan Province is one of the most important sugar production bases, and its sugarcane industry holds a pivotal position nationwide. Compared with other regions, sugarcane-growing areas have developed distinct characteristics due to their unique geographical and climatic conditions. In this region, sugarcane is mainly grown on arid slopes in hilly areas, with a small amount grown on plains. In recent years, due to the scarcity of land resources, sugarcane cultivation has been continuously shifting to higher altitudes, making planting sugarcane "in the clouds" and "on rocks" a development trend in this sugarcane-growing area. However, the increase in altitude has also brought problems such as poor soil conditions and drought, especially the inability to retain soil organic matter for long periods, which seriously restricts sugarcane yield. Therefore, clarifying the driving mechanism of soil organic matter changes in this sugarcane-growing area is crucial for understanding how to increase the organic matter content and promote sugarcane yield.
[0003] Soil organic matter is the cornerstone of soil health; it serves as an energy source and food for soil microorganisms and animals, while maintaining soil biodiversity and ensuring the normal functioning of the soil ecosystem. As a cementing agent for soil aggregates, soil organic matter binds small soil particles into loose, porous aggregates, enhancing soil aeration and water retention capacity. Its internal pores can store essential nutrients such as nitrogen, phosphorus, and potassium for plant growth, slowly releasing them for plant absorption, forming a slow-release reservoir and preventing their loss with water. Furthermore, soil organic matter also influences climate change.
[0004] Investigating the mechanisms by which various driving factors influence soil organic matter is crucial. Traditional methods primarily employ linear regression, such as studies on the relationship between soil organic carbon and particulate matter (POC) or mineral-associated organic carbon (MAOC). Correlation heatmap analysis is also used, for example, to examine the correlation between soil organic matter components and enzyme activity. However, these correlation coefficients, calculated based on linear fitting, can only evaluate cases where the dependent variable monotonically increases or decreases with the independent variable; for non-monotonic cases, they show poor correlation. This often overlooks these seemingly poor factors because these methods do not consider nonlinear driving forces. Remote sensing prediction has also become a commonly used method in recent years, offering advantages such as short data acquisition time, wide coverage, rich information, and low cost. For example, ZY1-02D hyperspectral satellite imagery data is used to map soil organic matter content in agricultural areas. With the in-depth development of machine learning, researchers have begun to utilize various machine learning models for predicting soil organic matter content, such as artificial neural networks, convolutional neural networks, support vector machines, random forests (RF), and extreme gradient boosting (XGBoost). While machine learning models possess excellent predictive capabilities, as "black box" models, their decision-making processes are opaque, and most cannot provide a reasonable explanation for their predictions. Therefore, researchers have introduced the Shapley Additive Explanation Model (SHAP). Based on game theory, SHAP can be used to explain the predictions of machine learning models, making these "black box" models transparent. SHAP analysis can explain how each factor in the trained model affects the response variable and to what extent it contributes. For example, Wu et al. used the XGBoost-SHAP model to explain the nonlinear impact of environmental regulation on the ecological resilience of the Yangtze River Delta region. Summary of the Invention
[0005] The main objective of this invention is to provide a method and system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, in order to solve the problem that existing technologies only perform partial dependency plots to identify dominant factors.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, the method comprising: Obtain the spatial distribution of all sampling points, using different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete, or random; divide the sugarcane areas and analyze the soil nutrients in each sugarcane area for comparative analysis. Based on the analysis results, a random forest regression model was constructed; characteristic data of each sugarcane area were extracted and interpreted using SHAP values; a scatter plot of the relationship between the characteristic data of each sugarcane area and the corresponding SHAP value was drawn to obtain a SHAP dependency plot; the relationship between factors and organic matter was described based on the SHAP dependency plot. The PLS-SEM algorithm is used to estimate the random forest model to obtain direct or indirect effects; the sum of direct and indirect effects is used to obtain the total effect; the importance ranking of the SHAP is verified by comparing the magnitude of the total effect of different driving factors.
[0007] As a further improvement of the present invention, the process of determining whether the distribution of sampling points is clustered, discrete, or random includes the following steps: Obtain the geographic coordinates of all soil sampling points to form a spatial distribution map. Visually distinguish the sampling points based on the different color tones of the sampling year and the sugarcane area to which they belong. Based on the coordinates of the sampling points and their corresponding soil property values, the global spatial autocorrelation index is calculated. Based on the range of the global spatial autocorrelation index values, it is determined whether the spatial distribution of all sampling points exhibits a clustered pattern, a discrete pattern, or a random pattern. Based on the established sugarcane area management boundaries, the entire study area was divided into several spatial units. Soil nutrient data from sampling points were integrated within each spatial unit, and comparative analysis of multiple nutrient indicators between units was conducted to obtain results characterizing the spatial differentiation of soil nutrients in each sugarcane area.
[0008] As a further improvement of the present invention, the process of describing the relationship between factors and organic matter based on SHAP dependency graphs includes the following steps: Soil organic matter was collected from each sugarcane growing area. Soil organic matter was used as the target variable, and soil properties, environmental factors, and management measures were used as feature variables to construct a random forest regression model. The random forest regression model was trained using data from each sugarcane growing area to obtain the optimal random forest regression model. The data from each sugarcane region is input into a random forest regression model to obtain a preliminary global importance ranking; based on the SHAP value of each feature of the data samples from each sugarcane region, the global importance of SHAP is obtained. For the key driving factors that rank in the preset position, draw SHAP dependency graphs respectively; identify the curves of SHAP value changes with the value of key driving factors; analyze nonlinear interactions, calculate the interaction SHAP value between pairs of features, and obtain the relationship between key driving factors and organic matter.
[0009] As a further improvement of the present invention, the process of verifying the importance ranking of SHAPs specifically includes the following steps: The feature variables in the random forest model are used as observed variables, and the associated features among the feature variables are summarized as latent variables; based on historical domain knowledge, causal relationships between latent variables are assumed to form a model path graph; Iteratively estimate the scores of latent variables. The initial scores of latent variables are calculated by weighting the variables to perform external approximation. Based on the latent variable associations in the model path graph, the scores of adjacent latent variables are reweighted to estimate the current latent variable score to perform internal approximation. The external approximation and internal approximation are repeated until the change in the latent variable score is less than a preset threshold to obtain a stable estimate. The path coefficient directly represents the immediate impact of a latent variable on organic matter, yielding the direct effect; the effect transmitted through mediating variables is calculated by multiplying the path coefficients, yielding the indirect effect; the sum of the direct and indirect effects yields the overall intensity of the factor's impact on organic matter. Assess the explanatory power of endogenous latent variables and predict their correlations; determine the reasonableness of the fit index; rank the total effects of each driving factor by size and compare them with the importance ranking of the SHAP values; if the two are consistent, the random forest regression model is considered qualified; if there are differences, readjust the parameters of the random forest regression model.
[0010] As a further improvement of the present invention, the process of forming a model path diagram includes the following steps: Based on the key feature variables identified by the SHAP dependency graph that have nonlinear or significant marginal impacts on soil organic matter, the key features are used as explicit variable indicators to construct latent variables that reflect different attribute dimensions, and initial path relationship assumptions are established between the latent variables to form a conceptual model framework for structural equations. The partial least squares estimation method is used to calculate the complete model containing latent and manifest variables. The solution is iteratively solved until the latent variable scores and path coefficients are stable, and the standardized path coefficient estimates between each latent variable after data fitting are obtained. Based on the obtained stable path coefficients, in the conceptual model framework of structural equation modeling, the direct effect value of the driving factor on organic matter is defined as the coefficients of all directly connected paths from a driving factor latent variable to the soil organic matter latent variable; the indirect effect value of the driving factor on organic matter is defined as the sum of the products of the coefficients of all indirect paths from a driving factor latent variable to the soil organic matter latent variable through other mediating latent variables. The direct effect value of each driving factor on soil organic matter is integrated with all indirect effect values to obtain the total effect value of the driving factor on soil organic matter. The set of total effect values will be used for subsequent validation.
[0011] As a further improvement of the present invention, the process of calculating the established complete model containing latent and manifest variables includes the following steps: Using the manifest variable measurement indicators and their observation data defined in the formed structural equation conceptual model framework, an initial external approximate score is generated for each latent variable through a linear weighting method based on external weights. Using the obtained external approximation scores of each latent variable, combined with the initial path relationship assumptions established between the latent variables, the correlation weights between the scores of each latent variable are calculated through an internal weighting strategy, and the internal approximation scores of each latent variable are updated accordingly. Repeat the external approximation and internal weighting process until the score sequence of all latent variables and the path coefficient estimates in the model reach a preset stable state. Then, based on the stable latent variable scores generated in the final iteration, calculate and output the standardized path coefficient estimates between each latent variable.
[0012] As a further improvement of the present invention, the process of calculating and outputting standardized path coefficient estimates among the latent variables includes the following steps: Based on the established initial path relationships, all latent variable pairs with direct causal orientations are identified in the conceptual model framework, and the final stable score sequences of the causal latent variables and the outcome latent variables in each pair are extracted and paired. For each pair of latent variable score sequences, the stable score sequence of the causal latent variable is used as the input source data, and the stable score sequence of the outcome latent variable is used as the target reference data. The optimization criterion of minimizing the sum of squared differences between the predicted sequence and the target reference sequence is applied to calculate the optimal weight of the contribution of the causal latent variable to the outcome latent variable. All locally optimal weights are standardized to eliminate the influence of different dimensions caused by the different dispersion of the latent variables. The standardized weights are then output as the path coefficient estimates between the corresponding latent variable pairs.
[0013] As a further improvement of the present invention, the process of processing the optimization criterion that minimizes the sum of squared differences between the predicted sequence and the target reference sequence includes the following steps: The stable score sequence of the latent causal variables is used as input, and scaled using an initial trial-and-error weight parameter to form a preliminary set of prediction sequences. The obtained preliminary prediction sequence is compared point by point with the target reference sequence of the extracted latent variables. The difference at each corresponding point is calculated to form a difference sequence. Then, the squares of each value in the difference sequence are summed to obtain an overall difference measure. By systematically adjusting the values of the trial weight parameters, repeatedly performing scaling and overall difference measurement calculations until the weight parameter value that minimizes the overall difference measurement is found, the weight parameter value is output as the optimal weight of the contribution of the causal latent variable to the outcome latent variable.
[0014] As a further improvement of the present invention, the process of systematically adjusting the numerical values of the trial weight parameters includes the following steps: Based on the obtained initial overall difference metric, under a preset initial exploration step size, the current trial weight parameters are positively and negatively fine-tuned respectively to obtain two new weight values; Substitute the two newly generated weight values into the system, scale them independently, and calculate their corresponding overall difference measure values to obtain three sequences of difference measure values corresponding to different weight parameters. Compare the magnitudes of the three obtained difference measures, select the weight parameter corresponding to the minimum value as the new current weight; update the exploration step size for the next step based on the local change pattern formed by the three values; repeat the process until a small adjustment to the current weight no longer causes a meaningful decrease in the overall difference measure, and output the current weight parameter value as the optimal weight.
[0015] To achieve the above objectives, the present invention also provides the following technical solution: A system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, applied to the aforementioned method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, wherein the system comprises: The sugarcane area data acquisition module is used to acquire the spatial distribution of all sampling points, using different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete, or random; divide the sugarcane area and analyze the soil nutrients of each sugarcane area for comparative analysis. The module for identifying dominant factors is used to construct a random forest regression model based on the analysis results; extract characteristic data of each sugarcane area and interpret them using SHAP values; draw scatter plots of the relationship between individual characteristic data of each sugarcane area and the corresponding SHAP values to obtain SHAP dependency plots; and describe the relationship between factors and organic matter based on SHAP dependency plots. The Validation Dominant Factors module is used to estimate the random forest model using the PLS-SEM algorithm to obtain direct or indirect effects; the sum of direct and indirect effects is used to obtain the total effect; and the importance ranking of the validation SHAP is performed by comparing the magnitude of the total effects of different driving factors.
[0016] To achieve the above objectives, the present invention also provides the following technical solution: An electronic device includes a processor and a memory coupled to the processor, the memory storing program instructions executable by the processor; when the processor executes the program instructions stored in the memory, it implements the above-described method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data.
[0017] To achieve the above objectives, the present invention also provides the following technical solution: A storage medium storing program instructions, which, when executed by a processor, implement the method described above for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data.
[0018] This invention first analyzes the geographical environment of sugarcane-growing areas in Province E and systematically evaluates soil nutrients. Then, it selects three models—RF, XGBoost, and Adaboost—to predict soil organic matter in these areas. The optimal random forest model, combined with Shapley additive interpretation, is chosen to explore the driving mechanism of soil organic matter changes in the sugarcane-growing areas of Province E. The aim is to identify the response thresholds of key driving factors through the random forest-SHAP model and elucidate the nonlinear interaction mechanism among these factors influencing organic matter. Finally, by constructing a partial least squares structural equation model (PLS-SEM), the effects of each factor are clarified, and their influence is quantified, distinguishing between direct and indirect effects. Applying machine learning models combined with Shapley additive interpretation to the sugarcane-growing areas of Province E provides a new method for predicting and analyzing soil organic matter in sugarcane-growing regions, offering new insights for subsequent sugarcane production. Attached Figure Description
[0019] Figure 1 This is a schematic flowchart of an embodiment of the method for analyzing the influencing factors of soil organic matter in sugarcane areas based on multi-source data according to the present invention. Figure 2 This is a schematic diagram illustrating the steps of determining whether the distribution of sampling points is clustered, discrete, or random, in an embodiment of the method for analyzing the influencing factors of soil organic matter in sugarcane areas based on multi-source data according to the present invention. Figure 3 This is a schematic diagram of the steps in an embodiment of the method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data of the present invention, which describes the relationship between factors and organic matter based on SHAP dependency graphs. Figure 4 This is a schematic flowchart illustrating the steps of verifying the importance ranking of SHAP in an embodiment of the method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data of the present invention. Figure 5 This is a schematic diagram of the functional modules of an embodiment of the sugarcane growing area soil organic matter influencing factor analysis system based on multi-source data of the present invention; Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device of the present invention; Figure 7 This is a schematic diagram of the structure of a storage medium according to an embodiment of the present invention; Figure 8 This is a spatial distribution diagram of the sampling points in this invention; Figure 9 This is a map showing the current soil nutrient status in the sugarcane growing area of City E. Figure 10 This is a SHAP abstract diagram of the present invention; Figure 11 This is a partial dependency graph of the driving factors of this invention; Figure 12 This is a heatmap showing the interaction of driving factors in this invention. Figure 13 This is a diagram showing the nonlinear interaction of the driving factors in this invention. Figure 14 This invention relates to the PLS-SEM model of the sugarcane area in City E. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] The terms "first," "second," and "third" used in this invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this invention are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the accompanying drawings). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0023] like Figure 1As shown, this embodiment provides an example of a method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data. In this embodiment, the method specifically includes the following steps: Step S1: Obtain the spatial distribution of all sampling points, and use different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete or random; divide the sugarcane areas and analyze the soil nutrients in each sugarcane area for comparative analysis. Step S2: Based on the analysis results, construct a random forest regression model; extract characteristic data of each sugarcane area and interpret them using SHAP values; draw a scatter plot of the relationship between the characteristic data of each sugarcane area and the corresponding SHAP values to obtain a SHAP dependency plot; describe the relationship between factors and organic matter based on the SHAP dependency plot. Step S3: Use the PLS-SEM algorithm to estimate the random forest model and obtain the direct or indirect effects; sum the direct and indirect effects to obtain the total effect; verify the importance ranking of SHAP by comparing the total effects of different driving factors.
[0024] Preferably, this embodiment quantifies the distribution pattern of sampling points by calculating the spatial autocorrelation index, providing a basis for regional division; it uses a random forest model to predict soil organic matter and visualizes the feature contribution with the help of SHAP values to generate a dependency graph to reveal the nonlinear relationship between factors and organic matter; it decomposes direct and indirect effects through a partial least squares structural equation model, comprehensively evaluates the total effect of driving factors, and verifies it with the importance ranking of SHAP.
[0025] In summary, this embodiment clarifies the spatial differentiation patterns of soil nutrients in sugarcane growing areas, supporting differentiated management and precision fertilization. The SHAP dependency graph visually displays the influence patterns of key factors, while PLS-SEM verifies the importance of features from a causal path perspective. The combination of these two methods enhances the robustness and credibility of the conclusions. It overcomes the limitations of single methods, providing a more comprehensive and scientific basis for agricultural decision-making. This embodiment achieves a closed-loop analysis from spatial pattern identification to mechanistic explanation, combining predictive accuracy with logical rigor, and is suitable for the precise diagnosis and regulation of complex agricultural ecosystems.
[0026] This embodiment first analyzes the geographical environment of a sugarcane growing area and systematically evaluates soil nutrients. Then, three models—RF, XGBoost, and Adaboost—were selected to predict soil organic matter in the sugarcane growing area. The optimal Random Forest model was chosen and combined with Shapley additive interpretation to explore the driving mechanism of soil organic matter changes in the sugarcane growing area. The aim is to identify the response thresholds of key driving factors through the Random Forest-SHAP model and elucidate the nonlinear interaction mechanism among the driving factors affecting organic matter. Finally, partial least squares structural equation modeling (PLS-SEM) is constructed to clarify the action paths of each factor, quantify the impact, and distinguish between direct and indirect effects. This is the first time that a machine learning model combined with Shapley additive interpretation has been applied to a sugarcane growing area, providing a new method for the prediction and analysis of soil organic matter in sugarcane growing areas and offering new insights for subsequent sugarcane production.
[0027] Furthermore, such as Figure 2 As shown, step S1, which determines whether the distribution of sampling points is clustered, discrete, or random, specifically includes the following steps: Step S11: Obtain the geographic coordinates of all soil sampling points to form a spatial distribution map. In the spatial distribution map, visual distinctions are made based on the different color tones of the sampling year and the sugarcane area to which they belong. Step S12: Based on the coordinates of the sampling points and their corresponding soil property values, calculate the global spatial autocorrelation index. Based on the range of the global spatial autocorrelation index values, determine whether the spatial distribution of all sampling points exhibits a clustered pattern, a discrete pattern, or a random pattern. Step S13: Based on the established sugarcane area management boundaries, the entire study area is divided into several spatial units. Soil nutrient data from sampling points are integrated within each spatial unit, and comparative analysis of multiple nutrient indicators between units is conducted to obtain results characterizing the spatial differentiation of soil nutrients in each sugarcane area.
[0028] Preferably, this embodiment obtains the geographic coordinates of all soil sampling points and presents them on a spatial distribution map with different color tones, allowing for a direct observation of the spatial distribution of sampling points in different years and sugarcane areas. Calculating the global spatial autocorrelation index using the coordinates of the sampling points and soil attribute values helps determine whether the spatial distribution of the sampling points is clustered, discrete, or random, which is crucial for understanding the distribution patterns of soil nutrients. Dividing the study area into several spatial units according to the sugarcane area management boundaries and integrating and comparing the soil nutrient data within each unit allows for detailed analysis of the spatial differentiation of soil nutrients in each sugarcane area. This enables a systematic analysis and evaluation of the spatial distribution characteristics of soil nutrients, providing a scientific basis for soil management and sugarcane area planning.
[0029] Furthermore, the process of calculating the global spatial autocorrelation index in step S12 specifically includes the following steps: Step S121: Based on the obtained geographic coordinates of all soil sampling points, construct a spatial proximity definition for each pair of sampling points, and calculate the difference between the soil attribute value of each point and the average value of the attribute in the entire study area; using the difference and the proximity definition, generate a standardized spatial covariance factor for each pair of sampling points to characterize the synergistic change of space and attribute. Step S122: Integrate and normalize the standardized spatial covariance factors of all sampling point pairs to obtain a single value that characterizes the overall spatial correlation strength, namely the global spatial autocorrelation index value. Step S123: Compare the global spatial autocorrelation index value calculated in the second step with the judgment threshold range determined in advance through theoretical distribution. Based on whether the global spatial autocorrelation index value falls into the high positive value range, the negative value range, or the range close to zero, determine whether the spatial distribution of all sampling points presents a clustered pattern, a discrete pattern, or a random pattern.
[0030] Preferably, this embodiment calculates the standardized spatial covariance factor for each pair of sampling points to identify the spatial synergistic changes in soil property values, thereby revealing the spatial correlation between sampling points. Integrating and normalizing the spatial covariance factors of all sampling point pairs yields a single value, the global spatial autocorrelation index, which quantifies the spatial correlation strength of soil properties throughout the study area. By comparing the calculated global spatial autocorrelation index with a preset judgment threshold range, the spatial distribution pattern of the sampling points can be classified and determined. If the index value is positive and high, it indicates a clustered pattern; if it is negative, it indicates a discrete pattern; and if it is close to zero, it indicates a random pattern.
[0031] Furthermore, such as Figure 3 As shown, step S2, which describes the relationship between factors and organic matter based on the SHAP dependency graph, specifically includes the following steps: Step S21: Obtain soil organic matter from each sugarcane area, use soil organic matter as the target variable, and use treated soil properties, environmental factors, management measures, etc. as feature variables to construct a random forest regression model; train the random forest regression model using data from each sugarcane area to obtain the optimal random forest regression model. Step S22: Input the data from each sugarcane region into the random forest regression model to obtain a preliminary global importance ranking; based on the SHAP value of each feature of the data samples from each sugarcane region, obtain the SHAP global importance. Step S23: For the key driving factors that rank in the preset position, draw SHAP dependency graphs respectively; identify the curve of SHAP value change of the factor value; analyze the nonlinear interaction, calculate the interaction SHAP value between pairs of features, and obtain the relationship between the factor and organic matter.
[0032] Preferably, this embodiment uses soil properties, environmental factors, and management measures as feature variables to construct a random forest regression model for soil organic matter, achieving a systematic fusion of multi-dimensional influencing factors; combining the feature importance ranking inherent in random forests with SHAP global importance analysis, key driving factors are cross-validated from the perspectives of model intrinsic weight and contribution; SHAP dependency graphs are drawn for highly important features to intuitively show their nonlinear relationship with organic matter, and further synergistic or antagonistic effects between features are quantified through interactive SHAP values.
[0033] In summary, this embodiment, while ensuring prediction accuracy, overcomes the limitations of traditional machine learning's black box approach by leveraging the SHAP framework, enabling the model output to have clear attributional explanations. Through dependency graphs and interaction analysis, it identifies the nonlinear influence paths of key factors on organic matter and the coupling effects of multiple factors, deepening the understanding of the formation and regulation mechanisms of soil organic matter. It clarifies the dominant driving factors and their modes of action in different sugarcane growing areas, providing a quantitative basis for regionally customized soil improvement measures and optimized field management strategies. This embodiment achieves a closed loop from predictive modeling to mechanistic explanation, combining algorithmic advancement with application orientation, providing a typical practical example of interpretable artificial intelligence for agricultural ecosystem research.
[0034] Furthermore, the process of deriving the relationship between factors and organic matter in step S23 specifically includes the following steps: Step S231: For the key driving factors that rank in the preset position, draw SHAP dependency diagrams respectively; identify the curve of SHAP value change of the factor value; the point where the slope of the curve changes is the potential threshold or inflection point of the factor's influence on soil organic matter; identify the curve of SHAP value change with the factor value. Step S232: Based on the recognition results, determine whether the curve is monotonically increasing / decreasing, single-peak / single-valley, step-shaped, or saturated; map the third important variable on the dependency graph using color mapping. Step S233: Calculate the SHAP value of the interaction between each pair of features, calculate the average interaction strength between all feature pairs that meet the important features, and draw a heat map; for the strongest preset logarithmic interaction, draw the interaction dependency graph to obtain the influence surface of soil organic matter under dual constraints.
[0035] Preferably, this embodiment analyzes the points where the slope of the SHAP value changes with the feature to accurately locate the potential threshold or turning point affecting soil organic matter, revealing the critical state of the factor's effect; it summarizes the dependency relationship into typical patterns such as monotonic, unimodal, stepwise, or saturated, and introduces color mapping to present the interactive influence of the third variable, realizing a three-dimensional expression of multi-dimensional relationships; it calculates the average interaction SHAP value between feature pairs, displays the global interaction intensity with a heat map, and draws a three-dimensional dependency graph for strong interaction features, presenting the influence surface under dual constraints.
[0036] In summary, this embodiment breaks through the traditional assumption of linear or monotonic relationships, identifies nonlinear thresholds and multimodal responses, and is more in line with the complex behavior of agricultural ecosystems; through interactive heatmaps and three-dimensional surfaces, it intuitively reveals the synergistic or antagonistic effects between features, and clarifies the complex pathways through which multiple driving factors jointly affect organic matter; the identified impact thresholds can directly guide the critical adjustment points of field management measures, helping to achieve precise regulation and risk warning.
[0037] Furthermore, such as Figure 4 As shown, the process of verifying the importance ranking of SHAPs in step S3 specifically includes the following steps: Step S31: Use the feature variables in the random forest model as observed variables, and summarize the associated features among the feature variables as latent variables; based on historical domain knowledge, assume causal relationships between latent variables to form a model path diagram; Step S32: Iteratively estimate the latent variable scores. Calculate the initial scores of the latent variables by weighting the variables to perform external approximation. Based on the latent variable associations in the model path graph, reweight the scores of adjacent latent variables to estimate the current latent variable score to perform internal approximation. Repeat external and internal approximation until the change in latent variable scores is less than a preset threshold to obtain a stable estimate. Step S33: The path coefficient directly represents the immediate impact of a latent variable on organic matter, yielding the direct effect; the effect transmitted through the mediating variable is calculated by multiplying the path coefficients to obtain the indirect effect; the sum of the direct and indirect effects yields the overall impact strength of the factor on organic matter. Step S34: Evaluate the explanatory power of endogenous latent variables and predict their correlations; assess the reasonableness of the fit index; rank the total effects of each driving factor by size and compare them with the importance ranking of the SHAP values; if they are consistent, the random forest regression model is qualified; if there are differences, readjust the parameters of the random forest regression model.
[0038] Preferably, this embodiment summarizes the feature variables of random forest into latent variables, and establishes a causal relationship path graph between latent variables based on domain knowledge to achieve the integration of observational data and theoretical framework; through alternating iterations of external approximation and internal approximation until the score stabilizes, it effectively handles the problem of latent variable estimation in high-dimensional complex data; it quantifies direct effects, indirect effects and total effects to reveal the influence path of factors on organic matter; and through comparison of total effect ranking and SHAP importance ranking, it achieves dual verification of machine learning and statistical models.
[0039] In summary, this embodiment overcomes the limitations of traditional machine learning correlation analysis by analyzing the direct and indirect effects of features on organic matter from a causal path perspective, thus enhancing the mechanistic credibility of the conclusions. Through consistency testing of the total effect ranking of PLS-SEM and the importance ranking of SHAP, it provides statistical validation for the random forest model, enhancing the robustness and reliability of the overall analytical framework. Clarifying the direct contributions and indirect transmission paths of different driving factors helps identify key intervention nodes and provides a scientific basis for developing multi-level, interconnected soil management plans.
[0040] Furthermore, the process of forming the model path diagram in step S31 specifically includes the following steps: Step S311: Based on the key feature variables that have nonlinear or significant marginal impact on soil organic matter identified by the SHAP dependency graph, the key features are used as manifest variable indicators to construct latent variables that reflect different attribute dimensions, and initial path relationship assumptions are established between latent variables to form a conceptual model framework for structural equations. Step S312: Use the partial least squares estimation method to calculate the established complete model containing latent and manifest variables, and iterate until the latent variable scores and path coefficients are stable, to obtain the standardized path coefficient estimates between each latent variable after data fitting. Step S313: Based on the obtained stable path coefficient results, in the conceptual model framework of structural equation modeling, the direct effect value of the driving factor on organic matter is defined as the coefficient of all directly connected paths from a driving factor latent variable to the soil organic matter latent variable; the indirect effect value of the driving factor on organic matter is defined as the sum of the products of the coefficients of all indirect paths from a driving factor latent variable to the soil organic matter latent variable through other mediating latent variables. Step S314: Integrate the direct effect value of each driving factor on soil organic matter with all indirect effect values to obtain the total effect value of the driving factor on soil organic matter. This total effect value set will be used for subsequent verification.
[0041] Preferably, this embodiment achieves a comprehensive analysis and assessment of the factors influencing soil organic matter. First, by identifying key characteristic variables and constructing latent variables reflecting different attribute dimensions, a comprehensive conceptual model framework is formed. Next, partial least squares estimation is used to perform calculations, enabling a quantitative analysis of the relationship between latent and manifest variables in the model, ensuring the model's stability and accuracy. By calculating direct and indirect effects, the specific impact of each driving factor on soil organic matter can be accurately assessed, including both direct and indirect effects. Finally, by integrating the direct and indirect effect values, a comprehensive set of total effect values is obtained, providing a scientific basis for subsequent verification and policy formulation. This embodiment not only improves the understanding of the factors influencing soil organic matter but also provides more effective decision support for soil management and protection.
[0042] Furthermore, step S312, which involves calculating the complete model containing latent and manifest variables, specifically includes the following steps: Step S3121: Using the manifest variable measurement indicators and their observation data defined in the formed structural equation conceptual model framework, generate an initial external approximate score for each latent variable through a linear weighting method based on external weights; Step S3122: Using the obtained external approximation scores of each latent variable, combined with the initial path relationship assumptions established between the latent variables, calculate the correlation weights between the scores of each latent variable through an internal weighting strategy, and update the internal approximation scores of each latent variable accordingly. Step S3123: Repeat the external approximation and internal weighting process until the score sequence of all latent variables and the path coefficient estimates in the model reach a preset stable state. Then, based on the stable latent variable scores generated in the final iteration, calculate and output the standardized path coefficient estimates between each latent variable.
[0043] The principle of the external weighting linear weighting method is to construct a unique linear combination expression for each latent variable. Specifically, the observed data of all manifest variables belonging to the same latent variable are used as input, and an iterative algorithm is used to determine a set of optimal weight coefficients. The goal of calculating this set of weights is to ensure that the score of the latent variable generated by this linear combination not only summarizes the information carried by its corresponding manifest variable to the greatest extent, but also maintains the strongest correlation with the scores of other latent variables associated with it in the model structure.
[0044] The principle of the internal weighting strategy is based on information transmission and feedback through a pre-defined structural path network. Its core lies in identifying all other latent variables directly related to any given latent variable, according to the established initial path relationship assumptions. The strategy calculates a set of internal weights based on the external approximation scores of these related latent variables in the current iteration step, and the pre-defined relationship direction between them and the target latent variable (such as positive or negative influence). This set of internal weights is used to re-weight and summarize the scores of all related latent variables, thereby updating the new internal approximation score of the target latent variable. This process essentially enforces a constraint and adjustment of the theoretical path relationship based on the latent variable scores of the current iteration, making the updated latent variable scores more closely match the assumed causal structure, and through repeated iterations, gradually converges the computational results of the entire model to a stable and self-consistent state.
[0045] Preferably, this embodiment uses a linear weighting method combined with observed data of manifest variables to generate initial external approximation scores for latent variables, which helps to preliminarily estimate the role of latent variables in the structural equation model. Utilizing the external approximation scores and the initial path relationship assumptions between latent variables, an internal weighting strategy is used to calculate the association weights between latent variables, updating the latent variable scores and enhancing the accuracy of the model's internal structure. The external approximation and internal weighting process is repeated until the latent variable score sequence and path coefficient estimates in the model tend to stabilize, ensuring the convergence and stability of the model estimation. Based on the finally stable latent variable scores, standardized path coefficient estimates between each latent variable are calculated and output, providing a basis for further analysis and interpretation of causal relationships and effect sizes in the structural equation model. This enhances the accuracy and reliability of the model estimation, provides an effective quantitative analysis tool for exploring the complex relationships between latent and manifest variables, thereby improving the explanatory and predictive power of the structural equation model.
[0046] Furthermore, the process of calculating and outputting the standardized path coefficient estimates among the latent variables in step S3123 specifically includes the following steps: Step S31231: Based on the established initial path relationship, identify all latent variable pairs with direct causal orientation in the conceptual model framework, and extract and pair the final stable score sequences of the causal latent variable and the outcome latent variable in each pair. Step S31232: For each pair of latent variable score sequences, the stable score sequence of the causal latent variable is used as the input source data, and the stable score sequence of the outcome latent variable is used as the target reference data. The optimization criterion of minimizing the sum of squared differences between the predicted sequence and the target reference sequence is applied to calculate the optimal weight of the contribution of the causal latent variable to the outcome latent variable. Step S31233: Standardize all the obtained local optimal weights to eliminate the influence of the dimension caused by the different dispersion of the latent variables themselves, and output the standardized weights as the path coefficient estimates between the corresponding latent variable pairs.
[0047] Preferably, this embodiment achieves accuracy and standardization in the estimation of path coefficients between latent variables, thereby achieving a more objective and accurate assessment of the strength of relationships between various latent variables. Specifically: by identifying pairs of latent variables with direct causal orientation, the corresponding stable score sequences are extracted; by using optimization criteria to calculate the contribution weights of causal latent variables to outcome latent variables, the direct interaction between latent variables can be quantified; the obtained local optimal weights are standardized to eliminate the influence of the dispersion differences of the latent variables themselves, making the obtained path coefficient estimates comparable and able to more accurately reflect the relative strength of relationships between latent variables. This embodiment achieves accurate estimation and standardization of the strength of relationships between latent variables.
[0048] Furthermore, the process of minimizing the sum of squared differences between the predicted sequence and the target reference sequence in step S31232 specifically includes the following steps: Step S312321: Take the stable score sequence of the latent causal variables as input, scale it using an initial trial weight parameter, and form a preliminary prediction sequence; Step S312322: Compare the obtained preliminary prediction sequence with the target reference sequence of the extracted latent variables point by point, calculate the difference at each corresponding point and form a difference sequence, and then sum the squares of each value in the difference sequence to obtain an overall difference measure. Step S312323: By systematically adjusting the values of the trial weight parameters, repeatedly perform scaling and overall difference measurement calculations until the weight parameter value that minimizes the overall difference measurement is found, and output the weight parameter value as the optimal weight of the contribution of the causal latent variable to the outcome latent variable.
[0049] Preferably, this embodiment minimizes the difference between the predicted sequence and the target reference sequence by continuously adjusting the weight parameters, thereby improving the accuracy of the prediction results. By systematically adjusting the weight parameters and calculating the overall difference metric, the optimal weight parameter values can be found, maximizing the contribution of the causal latent variable to the outcome latent variable. Through point-by-point comparison, difference calculation, and summation, the impact of outliers can be reduced, improving the model's robustness. Through preliminary prediction, difference calculation, and parameter optimization, the optimal parameters can be quickly found, reducing computational load and improving computational efficiency. Continuous iterative optimization allows the model to better generalize to new data, improving its applicability. Dynamically adjusting the weight parameters adapts to data changes, enhancing the model's dynamic adjustment capability. This approach improves prediction accuracy, optimizes parameters, enhances robustness and generalization ability, increases computational efficiency, and strengthens the model's dynamic adjustment capability.
[0050] Furthermore, the process of systematically adjusting the numerical values of the trial weight parameters in step S312323 specifically includes the following steps: Step S3123231: Based on the obtained initial overall difference metric, under a preset initial exploration step size, the current trial weight parameters are positively and negatively fine-tuned respectively to obtain two new weight values; Step S3123232: Substitute the two generated new weight values into the system, scale them independently, and calculate their corresponding overall difference measure values to obtain three sequences of difference measure values corresponding to different weight parameters. Step S3123233: Compare the magnitudes of the three obtained difference measures, select the weight parameter corresponding to the minimum value as the new current weight; update the exploration step size for the next step based on the local change pattern formed by the three values; repeat until a small adjustment of the current weight no longer causes a meaningful decrease in the overall difference measure value, and output the current weight parameter value as the optimal weight.
[0051] Preferably, this embodiment can effectively find the optimal weight parameters by systematically adjusting and optimizing the weight parameters, thereby improving the overall performance of the difference metric.
[0052] like Figure 5 As shown, this embodiment also provides an embodiment of a system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data. In this embodiment, the system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data is applied to the method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data as described in the above embodiment. The system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data includes: Sugarcane area data acquisition module 1 is used to acquire the spatial distribution of all sampling points, using different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete or random; divide the sugarcane area and analyze the soil nutrients of each sugarcane area for comparative analysis. Module 2, which identifies dominant factors, is used to construct a random forest regression model based on the analysis results; extract characteristic data of each sugarcane region and interpret them using SHAP values; draw a scatter plot of the relationship between individual characteristic data of each sugarcane region and the corresponding SHAP value to obtain a SHAP dependency plot; and describe the relationship between factors and organic matter based on the SHAP dependency plot. The Validation Dominant Factor Module 3 is used to estimate the random forest model using the PLS-SEM algorithm to obtain the direct or indirect effects; the sum of the direct and indirect effects is used to obtain the total effect; and the importance ranking of the validation SHAP is performed by comparing the magnitude of the total effects of different driving factors.
[0053] Preferably, this embodiment uses spatial autocorrelation index to determine the distribution pattern of sampling points and combines it with regional soil nutrient comparison to achieve synergistic analysis of geospatial attributes and soil attributes; it uses random forest regression as the basic prediction model, reveals the nonlinear influence of single factors through SHAP dependency graph, and then uses PLS-SEM to verify the causal path; by comparing the consistency between SHAP importance ranking and PLS-SEM total effect ranking, the dominant factor is double-verified to improve the reliability of the conclusion.
[0054] In summary, this embodiment not only clarifies the spatial differentiation pattern of soil nutrients in sugarcane growing areas, but also reveals in depth the key driving factors affecting organic matter and their mechanisms of action. The SHAP dependency diagram visually demonstrates the complex relationship between factors and organic matter, while PLS-SEM further distinguishes between direct and indirect effects, making the mechanism explanation more hierarchical and convincing. By spatially zoning to identify key management areas and by analyzing dominant factors to clarify regulatory targets, it provides systematic support for precision soil management in sugarcane growing areas, from macro-level planning to micro-level intervention.
[0055] like Figure 6 As shown, this embodiment provides an embodiment of an electronic device 4, which includes a processor 41 and a memory 42 coupled to the processor 41.
[0056] The memory 42 stores program instructions for implementing the method for analyzing the influencing factors of soil organic matter in sugarcane areas based on multi-source data in any of the above embodiments.
[0057] The processor 41 is used to execute program instructions stored in the memory 42 to lay out a method for analyzing the influencing factors of soil organic matter in sugarcane areas based on multi-source data.
[0058] The processor 41 can also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.
[0059] Furthermore, Figure 7 This is a schematic diagram of the structure of a storage medium according to an embodiment of this application. The storage medium 5 of this embodiment stores program instructions 51 capable of implementing all the methods described above. These program instructions 51 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.
[0060] This embodiment is a specific application of the above embodiment's method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data. The details are as follows: This invention collects and integrates data from multiple sources to investigate the impact of driving factors on soil organic matter in sugarcane fields in Region A. Data on sampling point latitude and longitude coordinates, altitude, soil nutrients, fertilization management, and mulching conditions were obtained from Yingmao Sugar Industry Group Co., Ltd. in Region A. This company is a leading sugar producer in Region A, owning 1.2 million mu of sugarcane fields, and soil nutrient indicators were tested in the company's soil analysis laboratory. Climate data, soil type data, elevation data for Region A, and soil texture data were all obtained from the National Earth System Science Data Center. The climate data has a spatial resolution of 1 km, precipitation units of 0.1 mm, and temperature units of 0.1℃. The data were generated by downscaling the region using the Delta spatial downscaling scheme based on global 0.5° climate data released by CRU and global high-resolution climate data released by WorldClim, and were validated using data from 496 independent meteorological observation points. The validation results are reliable. The elevation data for Area A is a 30m resolution Copernicus DEM elevation, slope, and aspect dataset (2015). The data was created by mosaicking and cropping the Copernicus Digital Elevation Model data, sourced from the Copernicus Global DEM product released by the European Space Agency in 2022. Soil texture data includes 250m resolution data on soil silt, sand, and clay content in Area A, comprising 18 layers at different depths (0-5 cm, 5-15 cm, 15-30 cm, 30-60 cm, 60-100 cm, and 100-200 cm). This invention uses the 0-30cm data for analysis.
[0061] The research framework includes data collection and processing, spatial distribution analysis of sampling points, soil nutrient data analysis, and analysis of nonlinear interaction mechanisms (see appendix for details). Figure 8 First, the collected data were integrated, resampled, and outlier-handled. Then, non-numerical data (such as soil type and whether the area was covered with plastic film) were coded. Next, spatial distribution and soil nutrient analysis of the sampling points were performed, with comparative analysis conducted on both the overall and local aspects of the sugarcane growing area in region A. Finally, based on the random forest-SHAP model and a partial least squares-structural equation model, the nonlinear interaction mechanism of each driving factor on organic matter was explored, including the ranking of driving factor importance, threshold analysis of driving factors, nonlinear interaction analysis, and the direct and indirect pathways of their effects.
[0062] To investigate the mechanisms driving changes in soil organic matter in sugarcane growing areas, this invention selected 23 indicators, including both numerical and non-numerical indicators, and categorized them into six latent variables: climate, soil nutrients, agronomic management, topography, soil type, and soil texture. Indicator codes were assigned to simplify terminology in subsequent plotting. Label codes were also assigned to the non-numerical indicators to facilitate calculation and analysis. Table 1 Classification of Driving Factors
[0063] All soil samples were collected in November 2022. Before sampling, the latitude, longitude and altitude of the sampling points were recorded. Following the five-point sampling method, soil samples were collected from 0-30 cm depths at each point, mixed together, and placed in cloth bags as soil samples for that point. The samples were then transported back to the laboratory and placed in a cool, ventilated place to dry. After passing through a 2 mm standard sieve, they were ready for use. Soil organic matter was determined using the potassium dichromate oxidation method, in accordance with agricultural standard NY / T1121.6-2006; soil pH was determined using the potentiometric method, in accordance with agricultural standard NY / T 1377-2007; soil available nitrogen was determined using the alkaline diffusion method, in accordance with local standard DB51 / T 1875-2014; available phosphorus was determined using the sodium bicarbonate extraction-molybdenum antimony spectrophotometric method, in accordance with agricultural standard NY / T 1121.7-2014; available potassium was determined using the ammonium acetate extraction-flame photometer method, in accordance with agricultural standard NY / T 889-2004.
[0064] The data from the National Earth System Science Data Center were resampled. The specific process was as follows: the downloaded grid data map was imported into ArcGIS Pro software, and then the sampling points were imported as layers. The coordinate system used was WGS 1984; for projected coordinate systems, coordinate system transformation was performed first. The "Extract Values to Points" tool was used to extract the point data. For the extracted monthly temperature and precipitation data, the annual average temperature and precipitation were calculated. For soil texture data, the average value of data extracted from different depths (0-5 cm, 5-15 cm, and 15-30 cm) was taken to represent the content at the 0-30 cm depth.
[0065] ArcGIS Pro software was used to display the location and elevation distribution of sampling points. Origin Pro software was used to create 3D pie charts of the overall and partial (four states) elevation distribution of sugarcane growing areas in region A. The elevation distribution of each elevation segment represents the elevation distribution of sugarcane planting in the sugarcane growing areas.
[0066] Based on the nutrient grading standards of the Second National Soil Survey, soil nutrients in sugarcane growing areas were graded, evaluated, and analyzed. Soil nutrients were classified into six levels according to content: extremely abundant, abundant, relatively abundant, adequate, scarce, and extremely scarce. Based on soil pH, they were classified into six levels: strongly acidic, acidic, weakly acidic, neutral, weakly alkaline, and alkaline. Using Origin Pro software, 3D pie charts of soil nutrients in the sugarcane growing area A, both overall and in specific areas (four states), were created. The proportion of each level represents the soil nutrient status of the sugarcane growing area.
[0067] Table 2 Evaluation Criteria for Soil Nutrient Indicators
[0068] This invention employs machine learning methods to study the complex relationships and interactions between soil organic matter and other explanatory variables in sugarcane growing areas. Models such as RF, XGBoost, and Adaboost were used in Python (3.12) to train and predict data from sugarcane growing areas in 2022. A fixed random seed of 42 was set to ensure reproducibility. Each model was trained 100 times, using 5-fold cross-validation. The average coefficient of determination (R²) after 100 training iterations was used. 2 The predictive performance of the model was evaluated using mean absolute error (MAE) and root mean square error (RMSE). In model training, soil organic matter in sugarcane fields of region A was used as the response variable, while explanatory variables included altitude, mulch application, nitrogen application, phosphorus application, potassium application, available nitrogen, available phosphorus, available potassium, pH, average annual temperature, average annual precipitation, soil type, and soil clay, sand, and silt content. To investigate the driving mechanisms behind changes in soil organic matter in sugarcane fields of region A, a machine learning model and the SHAP algorithm were integrated. Based on this model, a beehive diagram was plotted between the SHAP values of each driving factor and the corresponding soil organic matter, identifying the dominant driving factors, clarifying the direction of influence, and analyzing dependencies and interaction effects.
[0069] Partial Least Squares (PLS) structural equation modeling is a variance-based structural equation modeling technique. It is primarily used to analyze complex causal relationships between variables, with its core objective being prediction and theoretical development, rather than theoretical testing. Compared to other models (such as CB-SEM), PLS-SEM has lower sample size requirements and does not require data to follow a normal distribution. This study utilizes SmartPLS 4 to build models for numerical variables. Before modeling, the sample data was Z-score standardized using Python to meet the requirements of the modeling software. The coefficient of determination R was calculated for the model. 2 The path coefficients were then calculated again using PLSpredict in the software, with both the number of folds and repetitions set to 10. A fixed seed was selected for the random number generator to ensure the reproducibility of the model calculations. After the calculations were completed, the predicted correlation Q was recorded. 2 Using R 2 and Q 2 The model fit was evaluated using the criteria described in the literature by Du et al.
[0070] Table 3 Evaluation Criteria for PLS-SEM Models
[0071] The sampling points of this invention are mainly distributed in areas A, B, C, and D of E (for details on the principle, please refer to the appendix). Figure 8 ). Figure 8The data shows the elevation distribution of sugarcane sampling points in State E and the four states. Overall, the sugarcane growing areas in Region A have a large elevation range, mainly distributed between 800-1000 m. Furthermore, 48.30% of the sugarcane growing areas are above 1100 m, indicating that high-altitude cultivation is a key characteristic of sugarcane growing in Region A, where sugarcane cultivation faces challenges such as frost and drought. Locally, the sugarcane growing areas in Region A are mainly distributed between 800-1100 m because sugarcane in Region A is primarily grown on large, flat plains. Sugarcane growing areas in Regions B and D are mainly distributed between 1100-1400 m. In Region C, besides the 1100-1400 m elevation range, the proportion of sugarcane grown above 1400 m is the highest among the four states, followed by Region D. This is related to the prevalence of mountains and limited flatlands in the southeastern part of Region A. like Figure 9 As shown, the proportions of soil nutrient content at various sampling points in the overall and local sugarcane-growing areas of City A are shown. The organic matter content in the four states is concentrated in... and This indicates that sugarcane growing areas in region E have relatively abundant and adequate organic matter. Furthermore, regions A and B, compared to the other two states, are at the level of extremely abundant organic matter. The proportion of alkaline soils (7.5-8.5) has significantly increased. The pH of sugarcane in Zone E is mainly 4.5-5.5, a phenomenon also evident in Zones A and B, indicating that the soils in these two zones are predominantly acidic. This is attributed to the high temperature and humidity conditions in these areas, resulting in a predominance of lateritic soils. In contrast, Zones C and D show a certain degree of pH increase, with a higher proportion of neutral soils. Particularly in Zone C, the proportion of alkaline soil samples (7.5-8.5) reached 6.56%, attributed to lower rainfall in these areas, retaining more cations. Regarding available nitrogen, the sugarcane in Zone E is generally between moderate and relatively abundant. Compared to Zone A, the available nitrogen content in the other three zones shows a significant increase because... and The proportion of available phosphorus in the sampling points increased. The graph shows that the available phosphorus in sugarcane area E is mainly concentrated in… The difference is that Area A has more areas at the extremely rich level ( Sampling points for available phosphorus. The trend of available potassium in the four states is similar to that of alkaline nitrogen, that is, the available potassium content in area A is lower than that in the other three states, and the available potassium content in area B is the highest.
[0072] To systematically identify the key driving factors behind changes in soil organic matter in sugarcane growing areas and improve the accuracy of driving factor analysis, this invention employs three machine learning algorithms: Random Forest, XGBoost, and Adaboost. The performance of the models is rigorously evaluated using three commonly used metrics: R², MAE, and RMSE. Through comparison of the evaluation results, Random Forest demonstrates the best R² performance. 2Based on a comprehensive evaluation of the results (0.5934), MAE (4.4088), and RMSE (5.8246), RF was determined to be the optimal model for this invention. The RF model is a comprehensive learning method for classification and regression. It constructs a large number of decision trees during the training phase and then clusters these trees to achieve predictive effects. Compared with other models, the RF model has the advantages of high robustness and reduced overfitting. These results demonstrate the RF model's ability to capture the complex nonlinear functional relationships between driving factors such as climate, soil, and agronomic practices that influence changes in soil organic matter in sugarcane growing areas, while also providing a foundation for subsequent SHAP analysis of the contributions and interaction effects of each factor.
[0073] Table 4 Performance Comparison of the Three Models
[0074] This invention uses the SHAP method to interpret the prediction results of soil organic matter in sugarcane fields in region E in 2022 using a random forest model. The SHAP summary diagram shows the direction and intensity of the dynamic influence of each driving factor on soil organic matter (see appendix for details). Figure 10 This result allows for a precise assessment of the driving mechanisms of each factor. The driving factors are arranged from top to bottom according to their contribution to changes in organic matter. Among all influencing factors, available nitrogen (SNI) ranks first, being the most important factor affecting changes in soil organic matter in sugarcane growing areas. Further analysis shows that low values of SNI appear in the negative SHAP region, while high values correspond to the positive SHAP region, indicating a significant positive correlation between SNI and organic matter. Available phosphorus shows a similar pattern to SNI, also positively correlated with organic matter. SHAP analysis reveals that, among climatic factors, average annual temperature contributes more to changes in organic matter than average annual precipitation and is negatively correlated with organic matter. Among soil texture factors, organic matter is more easily affected by soil sand and silt content, with sand content positively correlated with organic matter and silt content negatively correlated. Among topographic factors, altitude and nitrogen fertilizer application in agronomic practices both show weak positive correlations with organic matter. The remaining driving factors have relatively small impacts on organic matter.
[0075] Based on the ranking of SHAP analysis, this invention constructs a dependency graph of driving factors and organic matter in sequence. Figure 11 This study aimed to investigate the nonlinear relationship and threshold effect between driving factors and organic matter. Among soil nutrient factors, available nitrogen and organic matter showed an approximately linear positive correlation (…). Figure 11 (a) indicates that carbon and nitrogen always vary co-evolved in soil. Available phosphorus exhibits a distinct threshold ( Figure 11When the available phosphorus content is below 20 mg / kg, soil organic matter increases with increasing available phosphorus. Once it exceeds 20 mg / kg, the soil organic matter content gradually stabilizes. Available potassium exhibits numerous positive and negative values throughout the process, indicating that the effect of available potassium on organic matter may be influenced by other factors. Soil pH has two threshold values (pH=4.5 and pH=5.5). Figure 11 When pH < 4.5, organic matter increases with increasing pH; when pH is between 4.5 and 4.5, organic matter content decreases; and when pH > 4.5, the effect of pH on organic matter becomes complex. Among soil texture factors, soil sand and soil silt exhibit diametrically opposed behavior; soil sand content increases significantly when it approaches 350 g. Figure 11 (b) The organic matter content changed from a continuous increase to a stable level, but when the soil silt content approached 350 ( Figure 11 (c) The organic matter content changed from stable to continuously declining. Soil clay particles exhibit two thresholds: 350 and 370. When <350, a negative correlation exists; between 350 and 370, a positive correlation emerges; and when >370, a negative correlation returns. Among climatic factors, the threshold for average annual temperature appears at 21℃. When <21℃, soil organic matter content is relatively stable; above 21℃, average annual temperature shows a positive correlation with organic matter. Average annual precipitation initially shows a positive correlation with organic matter, but when precipitation exceeds 1300 mm, organic matter begins to decline. For agronomical measures, nitrogen application exceeding 300 kg / ha and potassium application below 200 kg / ha showed a positive correlation. Phosphorus application exhibited three thresholds: 60 kg / ha, 96 kg / ha, and 144 kg / ha, displaying an "M"-shaped variation. A weak positive correlation was observed when phosphorus application was below 60 kg / ha; a weak negative correlation was observed between 60 kg / ha and 96 kg / ha; a weak positive correlation was observed between 96 kg / ha and 144 kg / ha; and a weak negative correlation was observed again when phosphorus application exceeded 144 kg / ha. Based on this analysis, the eigenvalues of mulched plants were predominantly negative, while those without mulch were predominantly positive, indicating a potential reduction in organic matter after mulching. 1000 m is the altitude threshold; a negative correlation was observed below 1000 m, and a positive correlation was observed above 1000 m. Paddy soil, red soil, and lateritic red soil are the most common soil types in sugarcane-growing areas. From the scatter distribution of characteristic values, they have different effects on organic matter. Paddy soil mostly has a SAP value above 0, indicating that it has a positive effect on increasing organic matter, while red soil and lateritic red soil mostly have negative values, indicating that they have a negative effect on increasing organic matter.
[0076] Based on the Random Forest-SHAP algorithm, the interaction analysis of various driving factors shows that ( Figure 12The top six most significant interaction effects were all related to available nitrogen, indicating that other factors such as soil sand content, available phosphorus, soil silt content, annual average temperature, pH, and altitude (SHAP interaction values of 0.4436, 0.3076, 0.2857, 0.1816, 0.1488, and 0.1322, respectively) all regulate organic matter content by influencing available nitrogen in sugarcane fields. Interestingly, some factors, such as available phosphorus and pH, contribute weakly when acting alone, but their impact on organic matter is enhanced when interacting with available nitrogen. In particular, available phosphorus shows the largest increase, indicating that available phosphorus is more dependent on available nitrogen in the soil when it exerts its effect.
[0077] To investigate the dynamic effects of driving factors on organic matter in sugarcane fields, this invention selected the top six interactions with the strongest interaction for further analysis and plotted their interaction distribution diagram. Figure 13The SHAP interaction value between available nitrogen and soil sand generally increases with increasing available nitrogen content. When available nitrogen is <75 mg / kg, the interaction is mainly limited by available nitrogen; when available nitrogen is >125 mg / kg, the interaction is mainly limited by soil sand content. The interaction between available nitrogen and available phosphorus has a threshold of 75 mg / kg. Below this threshold, medium to high levels of available phosphorus promote a positive interaction, while low levels of available phosphorus contribute more to a negative interaction. Overall, the SHAP interaction value between available nitrogen and available phosphorus gradually approaches 0, indicating that the main contribution of the interaction between available nitrogen and available phosphorus to organic matter is when available nitrogen is <75 mg / kg. The positive contribution of the interaction between available nitrogen and soil silt content to model predictions gradually increases with increasing available nitrogen content. Similar to soil sand, when available nitrogen is below 100 mg / kg, higher soil silt content contributes positively to the interaction, while lower soil silt content enhances the negative effect. The opposite is true when available nitrogen is above 100 mg / kg. The interaction between available nitrogen and average annual temperature increases with increasing available nitrogen. Simultaneously, average annual temperature and available nitrogen exhibit a synergistic complementary effect; at low available nitrogen content, an increase in average annual temperature can even turn the negative interaction into a positive one. The SHAP interaction values between available nitrogen and pH were mostly around 0, indicating a weak impact of their interaction on organic matter. However, when the pH was around 4.5, i.e., when the sugarcane soil was acidic, the impact of this interaction on organic matter at some sampling points depended on the content of available nitrogen in the soil. For example, some points showed a strong positive contribution when available nitrogen was <75 mg / kg, while others showed a strong negative contribution when available nitrogen was between 75-125 mg / kg. With increasing available nitrogen, the SHAP interaction values between available nitrogen and altitude generally showed a decreasing trend at low altitudes (800-1000 m). As altitude continued to increase, the interaction values showed an increasing trend, especially when the available nitrogen content was >125 mg / kg, indicating that high altitude and high available nitrogen content would push the model toward positive predictions.
[0078] Figure 14The results of PLS-SEM analysis are presented, describing the relationship between various indicators and organic matter changes in sugarcane growing areas. The R² values for soil texture, climate, and organic matter are all >0.33, indicating moderate fit for these latent variables. While the fit for soil nutrients is weaker, the overall model fit is within an acceptable range. The Q² values for all observed variables are greater than 0, indicating good predictive relevance. The PLS-SEM model explains 53.7% of the organic matter variance. The path coefficients of each latent variable on organic matter show that soil nutrients, climate, and topography have a positive impact on organic matter, especially soil nutrients (reaching 0.777). Soil texture and fertilization management, on the other hand, have a negative impact on organic matter. Furthermore, the model can also illustrate that changes in organic matter can simultaneously affect both the direct and indirect effects of latent variables. For example, soil nutrients have a direct impact on organic matter, while indirect effects can be represented by the product of path coefficients. More significant indirect effects include soil texture-soil nutrients-organic matter (0.421×0.777=0.327) and topography-soil nutrients-organic matter (0.393×0.777=0.306), indicating that the mechanism by which other factors regulate organic matter is mainly through influencing soil nutrients and thus altering organic matter. Figure 14 Solid lines represent positive effects, dashed lines represent negative effects, red lines represent path coefficients with an absolute value greater than 0.67, green lines represent path coefficients with an absolute value between 0.33 and 0.67, and purple lines represent path coefficients with an absolute value less than 0.33.
[0079] Table 5 Evaluation of PLS-SEM Model
[0080] In a certain region, mountainous terrain dominates the sugarcane-growing area in Zone E, with significant elevation differences, mainly concentrated above 800 m. This is significantly higher than the Guangxi sugarcane-growing region, which has the largest sugarcane-growing area in my country, with an elevation range of 0-217 m. Therefore, the sugarcane-growing areas in Zone E face more complex problems, including drought, frost damage, soil erosion, and difficulties in mechanized operations. These factors severely restrict the improvement of sugarcane yield and quality. The planting area in flatlands is relatively small.
[0081] Based on the classification and evaluation of soil nutrients in sugarcane growing areas of Region E, the overall soil level is at a medium level. Locally, areas B and A in western Region E perform better than areas C and two states in the east. This is particularly evident in organic matter reserves, as the relatively flat terrain of areas A and B makes organic matter less susceptible to water erosion. Furthermore, the greater number of plots suitable for combined mechanized harvesting in these areas, coupled with the in-situ crushing and returning of sugarcane leaves during harvest, is a significant reason for the higher organic matter reserves. For example, Hao et al.'s 15-year study of a corn-soybean field system found that straw return increased SOM content by 14.2%, with a significant increase in the proportion of microbial biomass carbon. Simultaneously, sugarcane leaf return may also contribute to the higher available phosphorus content in these areas, as sugarcane leaves release phosphorus during microbial decomposition. A study in the sugarcane growing area of City F supports this view, suggesting that sugarcane leaves are an important source of inorganic phosphorus components in the soil, and that crushing sugarcane leaves at a rate of 5 t / ha... -1 When returning fertilizer to the field, additional phosphate fertilizer is not necessary. The content of available nitrogen and potassium may be affected by both fertilization management and climate conditions. Although more nitrogen and potassium fertilizers were applied in areas A and B, especially with increased urea application, the excellent hydrothermal conditions accelerated the mineralization of fertilizers by microorganisms, resulting in higher nutrient utilization by sugarcane and lower residual amounts in the soil. In areas C and D, due to severe drought, nutrients were not efficiently utilized and gradually accumulated in the soil. At the same time, as mentioned above, the low rainfall in areas C and D led to an increase in soil pH in the sugarcane growing areas, even making them slightly alkaline. This limited the growth of sugarcane, as sugarcane is suitable for growing in neutral and slightly acidic soil environments, and alkaline soil reduces the availability of trace elements such as iron, manganese, zinc, and boron.
[0082] This study utilizes a random forest model combined with SHAP analysis to understand the complex mechanisms of soil organic matter changes in sugarcane growing areas in region E, considering the threshold effects of various driving factors and their nonlinear interactions. Furthermore, a PLS-SEM equation is constructed to demonstrate the direct and indirect effects between driving factors and the target variable, organic matter, as well as among the driving factors themselves. In complex soil environments, previous studies have explored the effects of driving factors on organic matter through linear regression or correlation analysis, failing to fully explain the overall impact. Alternatively, they have only used black-box models to predict soil organic matter, without further analysis. Therefore, this study is the first to apply this method to sugarcane growing areas in region E, which is of significant importance for enriching sugarcane cultivation theory.
[0083] When driving factors affect soil organic matter, except for available nitrogen (NG), which shows an approximately linear correlation (ranked first), the remaining factors exhibit non-linear correlations and threshold effects. The close synergistic changes observed between NG and organic matter are attributed to the simultaneous absorption and utilization of C and N elements by crops and microorganisms. In sugarcane production, more nitrogen fertilizer is typically required, sometimes necessitating additional application of fertilizers such as urea. After microbial mineralization, more NG is released and used for sugarcane growth, and sugarcane litter accumulates in the soil. The threshold for available phosphorus may be related to the Olsen-P agronomic threshold in the soil; above this value, the crop response to increased soil phosphorus content is small or zero. This means that the average Olsen-P agronomic threshold in sugarcane fields of Zone E is 20 mg / kg. Below this threshold, the growth and metabolic processes of sugarcane accelerate with increasing available phosphorus, leading to a rapid increase in sugarcane litter, the main carbon source in sugarcane fields. Above 20 mg / kg, the growth and metabolic processes of sugarcane stabilize, and the rate of newly converted organic matter from litter reaches equilibrium with the rate of organic matter consumption. Furthermore, the application of phosphate fertilizer did not have a significant impact on organic matter. Interestingly, this invention found that when the content of soil sand and silt is less than 350, as the content increases, organic matter tends to be stored more in sand rather than silt and clay. Wu et al. attributed this phenomenon to the low primary productivity in arid and semi-arid regions, limiting the ability of soil microorganisms to decompose plant-derived organic matter, leading to the accumulation of organic matter in coarser soil particles. The influence of average annual temperature on organic matter has a starting point of 21°C; after reaching this point, the positive contribution of average annual temperature to organic matter accumulation gradually increases. This differs from the findings of other scholars, who generally agree that soil organic matter content decreases with rising temperatures because higher temperatures increase microbial activity and accelerate organic matter decomposition. Therefore, this invention suggests that the observed phenomenon may be attributed to the combined effects of average annual temperature and precipitation. Locations with average annual temperatures >21℃ are primarily located in sugarcane-growing areas of Zone A, where heavy and concentrated rainfall frequently causes waterlogging. The lack of oxygen leads to a reduction in aerobic microorganisms, resulting in the accumulation of organic matter. The occurrence of the 1000m altitude threshold is also related to precipitation, especially in Zone A. High-intensity precipitation carries unstable carbon components (such as DOC) from high altitudes to lower altitudes through scouring and leaching, where they accumulate. However, precipitation decreases sharply above 1000m, preserving most of the organic matter. Although red soil and lateritic soil are the most widely distributed soil types in Zone E sugarcane-growing areas, making the optimal contribution point for organic matter accumulation occur in an acidic environment with a pH of 4.5, they are not conducive to further organic matter storage. The Al content in red soil and lateritic soil... 3+It can even be toxic to sugarcane and soil microorganisms. The importance of mulching in this study seems negligible, as full mulching was widely adopted and applied throughout the sugarcane-growing areas of Zone E in 2022 to combat drought, and it has had a significant effect on improving soil conditions and increasing sugarcane yield.
[0084] This invention further elucidates how multi-factor interactions affect organic matter. When driving factors influence organic matter, the main interactions are all related to available nitrogen (MPN), and the main limiting factors differ at different stages. This indicates that C-N coupling is a crucial process in the change of soil organic matter in sugarcane fields in region E, and other driving factors primarily regulate organic matter by influencing this process. When MPN content is below a threshold, the interaction is mainly limited by MPN, while when MPN content is above this threshold, the limiting effect of other driving factors on the interaction gradually increases. Through analysis, this invention suggests that when MPN content is above the threshold, increasing the content of soil sand and silt can positively regulate their interaction with MPN, because large-particle soils are more susceptible to external disturbances such as tillage and fertilization, and organic matter is more likely to accumulate there. The interaction between available phosphorus and MPN shows that if MPN is maintained above the threshold, even a low level of available phosphorus in the soil can make a positive contribution, consistent with the aforementioned Olsen-P agronomic threshold of 20 mg / kg. The interaction between temperature, altitude, and available nitrogen is influenced by spatial heterogeneity, with precipitation playing a significant role. Although the importance of average annual precipitation in the interaction is relatively low, it remains a non-negligible factor. Altitude affects temperature and precipitation, further influencing soil pH and nitrogen availability, resulting in the diversity of sugarcane growing areas in region E. The PLS-SEM equations also indicate that changes in organic matter are primarily influenced by direct soil nutrient levels. Other latent variables indirectly regulate soil organic matter by affecting soil nutrients; for example, soil texture-soil nutrient-organic matter and topography-soil nutrient-organic matter are the two most important indirect effects, consistent with the results obtained from the Random Forest-SHAP model analysis.
[0085] In this study, the invention integrates multi-source data and utilizes a random forest model for prediction, while introducing SHAP to interpret the black-box model's prediction results. However, the fitted R... 2 With an explanatory power of only 0.5934, the figure is moderate and cannot provide a comprehensive explanation of the mechanisms underlying changes in soil organic matter in the sugarcane growing areas of Zone E. Firstly, agricultural soils are inherently highly heterogeneous, and the five-point sampling method can only minimize errors, but it is still insufficient to accurately represent a specific plot. Secondly, the complex geographical environment of the sugarcane growing areas in Zone E amplifies these limitations. In practice, the 70° sugarcane planting slope results in significant elevation variations within the same plot, a situation very common in Zone E.
[0086] Regardless, the Random Forest-SHAP algorithm employed in this invention is unprecedented in the sugarcane region of City E and even in the field of sugarcane research worldwide. Although it still has certain shortcomings, it provides a new solution for predicting and explaining various variables of interest in the sugarcane field. Furthermore, future research should collect and integrate data from a wider geographical area and with more comprehensive dimensions, assigning appropriate geographical weights to gradually improve the accuracy of the model's predictions. By adjusting certain mechanisms within the model to alter the target variable, this will help farmers more accurately manage sugarcane cultivation.
[0087] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0088] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
[0089] The specific embodiments of the invention have been described in detail above, but these are merely examples, and the invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this invention. Therefore, all equivalent transformations, modifications, and improvements made without departing from the spirit and principles of this invention should be included within the scope of this invention.
Claims
1. A method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, characterized in that, The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data includes: Obtain the spatial distribution of all sampling points, using different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete, or random; divide the sugarcane areas and analyze the soil nutrients in each sugarcane area for comparative analysis. Based on the analysis results, a random forest regression model was constructed; characteristic data of each sugarcane area were extracted and interpreted using SHAP values; a scatter plot of the relationship between the characteristic data of each sugarcane area and the corresponding SHAP value was drawn to obtain a SHAP dependency plot; the relationship between factors and organic matter was described based on the SHAP dependency plot. The PLS-SEM algorithm is used to estimate the random forest model to obtain direct or indirect effects; the sum of direct and indirect effects is used to obtain the total effect; the importance ranking of the SHAP is verified by comparing the magnitude of the total effect of different driving factors.
2. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 1, characterized in that, The process of determining whether the distribution of sampling points is clustered, discrete, or random includes the following steps: Obtain the geographic coordinates of all soil sampling points to form a spatial distribution map. Visually distinguish the sampling points based on the different color tones of the sampling year and the sugarcane area to which they belong. Based on the coordinates of the sampling points and their corresponding soil property values, the global spatial autocorrelation index is calculated. Based on the range of the global spatial autocorrelation index values, it is determined whether the spatial distribution of all sampling points exhibits a clustered pattern, a discrete pattern, or a random pattern. Based on the established sugarcane area management boundaries, the entire study area was divided into several spatial units. Soil nutrient data from sampling points were integrated within each spatial unit, and comparative analysis of multiple nutrient indicators between units was conducted to obtain results characterizing the spatial differentiation of soil nutrients in each sugarcane area.
3. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 1, characterized in that, The process of describing the relationship between factors and organic matter based on SHAP dependency graphs includes the following steps: Soil organic matter was collected from each sugarcane growing area. Soil organic matter was used as the target variable, and soil properties, environmental factors, and management measures were used as feature variables to construct a random forest regression model. The random forest regression model was trained using data from each sugarcane growing area to obtain the optimal random forest regression model. The data from each sugarcane region is input into a random forest regression model to obtain a preliminary global importance ranking; based on the SHAP value of each feature of the data samples from each sugarcane region, the global importance of SHAP is obtained. For the key driving factors that rank in the preset position, draw SHAP dependency graphs respectively; identify the curves of SHAP value changes with the value of key driving factors; analyze nonlinear interactions, calculate the interaction SHAP value between pairs of features, and obtain the relationship between key driving factors and organic matter.
4. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 1, characterized in that, The process of ranking the importance of SHAPs includes the following steps: The feature variables in the random forest model are used as observed variables, and the associated features among the feature variables are summarized as latent variables; based on historical domain knowledge, causal relationships between latent variables are assumed to form a model path graph; Iteratively estimate the scores of latent variables. The initial scores of latent variables are calculated by weighting the variables to perform external approximation. Based on the latent variable associations in the model path graph, the scores of adjacent latent variables are reweighted to estimate the current latent variable score to perform internal approximation. The external approximation and internal approximation are repeated until the change in the latent variable score is less than a preset threshold to obtain a stable estimate. The path coefficient directly represents the immediate impact of a latent variable on organic matter, yielding the direct effect; the effect transmitted through mediating variables is calculated by multiplying the path coefficients, yielding the indirect effect; the sum of the direct and indirect effects yields the overall intensity of the factor's impact on organic matter. Assess the explanatory power of endogenous latent variables and predict their correlations; determine the reasonableness of the fit index; rank the total effects of each driving factor by size and compare them with the importance ranking of the SHAP values; if the two are consistent, the random forest regression model is considered qualified; if there are differences, readjust the parameters of the random forest regression model.
5. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 4, characterized in that, The process of creating a model path diagram includes the following steps: Based on the key feature variables identified by the SHAP dependency graph that have nonlinear or significant marginal impacts on soil organic matter, the key features are used as explicit variable indicators to construct latent variables that reflect different attribute dimensions, and initial path relationship assumptions are established between the latent variables to form a conceptual model framework for structural equations. The partial least squares estimation method is used to calculate the complete model containing latent and manifest variables. The solution is iteratively solved until the latent variable scores and path coefficients are stable, and the standardized path coefficient estimates between each latent variable after data fitting are obtained. Based on the obtained stable path coefficients, in the conceptual model framework of structural equation modeling, the direct effect value of the driving factor on organic matter is defined as the coefficients of all directly connected paths from a driving factor latent variable to the soil organic matter latent variable; the indirect effect value of the driving factor on organic matter is defined as the sum of the products of the coefficients of all indirect paths from a driving factor latent variable to the soil organic matter latent variable through other mediating latent variables. The direct effect value of each driving factor on soil organic matter is integrated with all indirect effect values to obtain the total effect value of the driving factor on soil organic matter. The set of total effect values will be used for subsequent validation.
6. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 5, characterized in that, The process of calculating the complete model containing latent and manifest variables includes the following steps: Using the manifest variable measurement indicators and their observation data defined in the formed structural equation conceptual model framework, an initial external approximate score is generated for each latent variable through a linear weighting method based on external weights. Using the obtained external approximation scores of each latent variable, combined with the initial path relationship assumptions established between the latent variables, the correlation weights between the scores of each latent variable are calculated through an internal weighting strategy, and the internal approximation scores of each latent variable are updated accordingly. Repeat the external approximation and internal weighting process until the score sequence of all latent variables and the path coefficient estimates in the model reach a preset stable state. Then, based on the stable latent variable scores generated in the final iteration, calculate and output the standardized path coefficient estimates between each latent variable.
7. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 6, characterized in that, The process of calculating and outputting standardized path coefficient estimates among the latent variables includes the following steps: Based on the established initial path relationships, all latent variable pairs with direct causal orientations are identified in the conceptual model framework, and the final stable score sequences of the causal latent variables and the outcome latent variables in each pair are extracted and paired. For each pair of latent variable score sequences, the stable score sequence of the causal latent variable is used as the input source data, and the stable score sequence of the outcome latent variable is used as the target reference data. The optimization criterion of minimizing the sum of squared differences between the predicted sequence and the target reference sequence is applied to calculate the optimal weight of the contribution of the causal latent variable to the outcome latent variable. All locally optimal weights are standardized to eliminate the influence of different dimensions caused by the different dispersion of the latent variables. The standardized weights are then output as the path coefficient estimates between the corresponding latent variable pairs.
8. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 7, characterized in that, The process of applying the optimization criterion that minimizes the sum of squared differences between the predicted sequence and the target reference sequence includes the following steps: The stable score sequence of the latent causal variables is used as input, and scaled using an initial trial-and-error weight parameter to form a preliminary set of prediction sequences. The obtained preliminary prediction sequence is compared point by point with the target reference sequence of the extracted latent variables. The difference at each corresponding point is calculated to form a difference sequence. Then, the squares of each value in the difference sequence are summed to obtain an overall difference measure. By systematically adjusting the values of the trial weight parameters, repeatedly performing scaling and overall difference measurement calculations until the weight parameter value that minimizes the overall difference measurement is found, the weight parameter value is output as the optimal weight of the contribution of the causal latent variable to the outcome latent variable.
9. The method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data according to claim 8, characterized in that, The process of systematically adjusting the numerical values of the trial weight parameters includes the following steps: Based on the obtained initial overall difference metric, under a preset initial exploration step size, the current trial weight parameters are positively and negatively fine-tuned respectively to obtain two new weight values; Substitute the two newly generated weight values into the system, scale them independently, and calculate their corresponding overall difference measure values to obtain three sequences of difference measure values corresponding to different weight parameters. Compare the magnitudes of the three obtained difference measures, select the weight parameter corresponding to the minimum value as the new current weight; update the exploration step size for the next step based on the local change pattern formed by the three values; repeat the process until a small adjustment to the current weight no longer causes a meaningful decrease in the overall difference measure, and output the current weight parameter value as the optimal weight.
10. A system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data, which is applied to the method for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data as described in any one of claims 1 to 9, characterized in that, The system for analyzing the influencing factors of soil organic matter in sugarcane growing areas based on multi-source data includes: The sugarcane area data acquisition module is used to acquire the spatial distribution of all sampling points, using different colors to distinguish different years and different sugarcane areas; calculate the spatial autocorrelation index to determine whether the distribution of sampling points is clustered, discrete, or random; divide the sugarcane area and analyze the soil nutrients of each sugarcane area for comparative analysis. The module for identifying dominant factors is used to construct a random forest regression model based on the analysis results; extract characteristic data of each sugarcane area and interpret them using SHAP values; draw scatter plots of the relationship between individual characteristic data of each sugarcane area and the corresponding SHAP values to obtain SHAP dependency plots; and describe the relationship between factors and organic matter based on SHAP dependency plots. The Validation Dominant Factors module is used to estimate the random forest model using the PLS-SEM algorithm to obtain direct or indirect effects; the sum of direct and indirect effects is used to obtain the total effect; and the importance ranking of the validation SHAP is performed by comparing the magnitude of the total effects of different driving factors.