A method for screening soil geochemical long-term monitoring indicators

By combining principal component analysis, KMO test, Bartlett's sphericity significance test, and random forest regression model, the technical gap in the screening of long-term soil geochemical monitoring indicators was filled, enabling efficient and scientific screening of soil monitoring indicators and improving the scientific nature and automation level of monitoring.

CN121093311BActive Publication Date: 2026-03-24GEOLOGICAL PROSPECTING TECH INST BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies lack a systematic approach for screening long-term soil geochemical monitoring indicators, leading to redundant investment of human resources and inefficient use of funds during the monitoring process. Furthermore, the monitoring indicators are not specific enough to meet the differentiated needs of different regions.

Method used

Principal component analysis was used to reduce the dimensionality of the data. The KMO test and Bartlett's sphericity significance test were combined to identify indicators that did not change. The repeatability test was used to initially screen indicators that were suitable for long-term monitoring. A random forest regression model was constructed to evaluate the predictive performance and finally determine the optimal combination.

Benefits of technology

A monitoring system with a reasonable indicator structure and high sensitivity to changes has been constructed, which has improved the scientific nature and practicality of monitoring work, realized the representativeness and predictability of monitoring indicators, and improved the accuracy and automation level of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093311B_ABST
    Figure CN121093311B_ABST
Patent Text Reader

Abstract

The application relates to a method for screening soil geochemical long-term monitoring indexes, belonging to the technical field of geochemistry, and solving the problems of technical blank, weak pertinence of monitoring indexes, waste of manpower and material resources caused by repeated monitoring process and the like in the prior art. The application adopts a principal component analysis method to perform data dimension reduction on monitoring data of at least two periods, to determine the number N of screening factors of various indexes; in combination with a repeatability test method, no-change indexes are identified and removed, and determined indexes M meeting long-term monitoring are preliminarily screened; if M is less than or equal to N, final monitoring indexes are screened; when N is greater than M, a random forest regression model is constructed to perform prediction performance evaluation on the screening factors, the screening factor with the best expected effect is removed, until N=M, and the final monitoring indexes are screened. The application significantly improves the scientificity, practicability and automation level of a soil geochemical long-term monitoring system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geochemistry, specifically to a method for screening long-term monitoring indicators of soil geochemistry. Background Technology

[0002] Soil is an indispensable and vital component of Earth's ecosystem, serving as the material basis for plant growth and supporting the habitat and reproduction of terrestrial organisms. To scientifically understand soil quality evolution trends and promptly identify problems such as soil degradation, acidification, salinization, and heavy metal pollution, soil geochemical monitoring has become a crucial tool for environmental protection and land resource management. Its basic principle involves systematically collecting surface cover samples, analyzing their elemental content and geochemical characteristics, and revealing the evolutionary pathways and potential risks of the soil environment. In particular, long-term, continuous geochemical monitoring data can accurately grasp the migration, enrichment, and transformation processes of elements within the soil ecosystem, providing fundamental support for predicting environmental change trends and formulating appropriate management policies.

[0003] However, due to the slow changes in soil systems and the complexity of influencing factors, single or short-term monitoring is insufficient to fully reflect their true evolution. Therefore, establishing a comprehensive, long-term, and continuous monitoring system is crucial. Currently, in my country's long-term soil monitoring practice, the existing monitoring system based on the "Multi-Objective Regional Geochemical Survey Plan (1:250000)" (DZ / T 0258-2014) generally adopts a unified set of 54 elemental indicators without distinguishing regional differences or indicator variability. However, the aforementioned standards and methods do not fully consider the reality that some elements change extremely slowly on long-term scales and do not require high-frequency monitoring. This leads to problems such as redundant investment of human resources and inefficient use of funds, weakening the efficiency and accuracy of long-term monitoring. Furthermore, current technologies for selecting key geochemical indicators suitable for long-term monitoring still rely on empirical judgment, lacking systematic screening methods, clear technical pathways, and adaptive algorithm frameworks, thus failing to support the differentiated monitoring needs for different soil types in different regions.

[0004] Therefore, it is necessary to construct a scientific, reasonable, regionally adaptable, and data-efficiency-oriented method for screening long-term soil geochemical monitoring indicators to provide technical support for improving the scientific rigor and standardization of soil monitoring work in my country. Summary of the Invention

[0005] To address the shortcomings of existing methods for screening long-term soil geochemical monitoring indicators, such as technological gaps, weak indicator specificity, and waste of human and material resources due to repetitive monitoring, this invention proposes a method for screening long-term soil geochemical monitoring indicators. The technical solution of this invention is as follows:

[0006] A method for screening long-term monitoring indicators of soil geochemistry includes the following steps:

[0007] S1: Principal component analysis is used to reduce the dimensionality of at least two periods of monitoring data to determine the number N of screening factors for each type of indicator;

[0008] S2: Based on the repeatability test method, identify and eliminate indicators that do not change, and preliminarily screen out the definitive indicators M that meet the requirements for long-term monitoring;

[0009] S3: If M is less than or equal to N, the final monitoring indicators are obtained through screening; when N is greater than M, a random forest regression model is constructed to evaluate the predictive performance of the screened factors, and the screened factors with the best expected effect are eliminated until N=M, and the final monitoring indicators are obtained through screening.

[0010] Furthermore, before the principal component analysis described in S1, the KMO test and Bartlett's sphericity significance calculation are performed. When the KMO value is >0.5 and the Bartlett's sphericity significance P-value is <0.05, there is a certain correlation between the indicators, and principal component data factor analysis can be performed. If the correlation between the indicators does not meet the conditions of KMO value >0.5 and Bartlett's sphericity significance P-value <0.05, the indicators are classified, and principal component analysis is performed under the premise of indicators in the same category.

[0011] Furthermore, in S1, the monitoring data for each period are analyzed separately using principal component analysis, and the number of screening factors for each period is confirmed. The number of screening factors in each period is compared, and the maximum value is selected as the final number of screening factor indicators. The data dimensionality reduction is achieved by selecting the number of principal components with a cumulative contribution rate of more than 85% as the number of screening factor indicators.

[0012] Furthermore, the total monitoring period should be at least 5 years.

[0013] Furthermore, the calculation formula for the KMO test is shown in Equation I below:

[0014] Formula I;

[0015] In the formula, r ij Let α be the simple correlation coefficient between variable i and variable j; ij KMO is the partial correlation coefficient between variable i and variable j, and its value ranges from 0 to 1.

[0016] Furthermore, the Bartlett sphericity significance calculation step includes Bartlett sphericity test calculation and Bartlett test degrees of freedom calculation;

[0017] The Bartlett sphericity test calculation formula is shown in Equation II below:

[0018] Formula II;

[0019] In the formula, n is the sample size, p is the number of variables, and |R| is the determinant of the sample covariance matrix;

[0020] The formula for calculating the degrees of freedom of the Bartlett test is shown in Equation III below:

[0021] Formula III;

[0022] In the formula, p is the number of variables;

[0023] x calculated using equation II 2 The significance p-value is calculated by comparing the chi-square distribution corresponding to df obtained from Equation III.

[0024] Furthermore, the repeatability test method described in S2 is achieved by calculating the relative double difference and the rate of no change between every two periods of monitoring data for the same indicator and the same location. The calculation formulas for the relative double difference and the rate of no change are shown in Equations IV and V below:

[0025] Formula IV;

[0026] In the formula, A1 represents the result of the first monitoring; A2 represents the result of the second monitoring.

[0027] Formula V.

[0028] Furthermore, when the relative double difference allowable limit RD ≤ 40%, the monitoring data is within the acceptable range of parallel samples and is considered to have no change; when the rate of no change calculated from two adjacent monitoring data periods is ≥ 90%, the indicator can be considered to have no change and the indicator can be determined to be a deductible indicator.

[0029] Furthermore, the steps in S3 to construct a random forest regression model to evaluate the predictive performance of the selected factors are as follows:

[0030] P1: Integrate data from at least two periods into one dataset, use N screening factor indicators as dependent variables in sequence, and other indicators as independent variables; divide the dataset into training and test sets in a 4:1 ratio and perform random forest regression simulation to obtain the predicted value of each indicator;

[0031] P2: Calculate the mean absolute error Mae, root mean square error RMSE, and fit index R for each indicator's predicted value. 2 ;

[0032] P3: Remove the screening factor with the best expected effect calculated in P2. If N is still greater than M, return to step P2 until N=M and the process ends.

[0033] P4: Integrate the various categories of indicators obtained from P3 to obtain the final long-term monitoring indicators for soil geochemistry.

[0034] Furthermore, the mean absolute error Mae, root mean square error RMSE, and fit exponent R... 2 The calculation formulas are shown in equations VI-VIII below:

[0035] Formula VI;

[0036] Equation VII;

[0037] Formula VIII; where, ;

[0038] In the formula, n is the total number of samples; y1 i y2 represents the true value. i These are predicted values.

[0039] Compared with existing technologies, this invention solves the problems of existing methods for screening long-term soil geochemical monitoring indicators, such as technological gaps, weak targeting of monitoring indicators, and waste of human and material resources due to repeated monitoring processes. Specifically, the beneficial effects are as follows:

[0040] 1. This invention integrates principal component analysis, repeatability testing, and random forest regression models, breaking away from the conventional thinking of linear or modular combination strategies of existing technologies. Furthermore, it breaks away from the traditional framework where factor analysis, error assessment, and model prediction are independent and executed step by step, forming a closed-loop index selection path of data-driven—stability verification—intelligent screening. This makes the monitoring index system selected by this invention representative, stable, and predictive, significantly improving the scientific nature and operability of long-term soil geochemical monitoring work, and fundamentally enhancing the scientific decision-making basis for index selection.

[0041] 2. This invention utilizes principal component analysis supplemented by the KMO test and Bartlett's sphericity significance test to scientifically evaluate the correlation and representativeness among soil monitoring indicators. Simultaneously, by combining regional land evolution characteristics and historical multi-period data, this invention employs a repeatability test to eliminate stable indicators lacking temporal representativeness, thereby constructing a monitoring system with a reasonable indicator structure and high sensitivity to change. This invention achieves a dual screening logic from correlation analysis to time sensitivity identification, significantly improving the scientific rigor and practicality of the monitoring system.

[0042] 3. This invention constructs a random forest regression model to train monitoring indicators through multiple rounds of modeling. Using prediction accuracy as the objective function, variables with minimal contribution to model prediction are gradually eliminated, ultimately determining the optimal combination of long-term monitoring indicators. This method not only improves the model's responsiveness to environmental change trends but also enhances the reliability and interpretability of data in subsequent decision-making processes such as land quality assessment and pollution evolution analysis, truly achieving the goal of "few but precise monitoring indicators." It embodies a data-driven intelligent analysis path and significantly improves the automation level of the method. Attached Figure Description

[0043] Figure 1 A process flow diagram for screening long-term monitoring indicators for soil geochemistry. Detailed Implementation

[0044] To make the technical solutions of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that the following embodiments are only used to better understand the technical solutions of the present invention and should not be construed as limiting the present invention.

[0045] Example 1.

[0046] Within a certain study area, there were 48 sampling points. Based on soil monitoring data from 2015 and 2021, the monitoring indicators included heavy metals As, Cd, Cr, Cu, Hg, Ni, Pb, and Zn; nutrient indicators N, P, and S; beneficial trace elements B, Br, Cl, Co, Ge, Mn, Mo, and V; healthy elements F, I, and Se; rare earth elements Ce, La, Sc, and Y; and other metal elements Ag, Ba, Be, Bi, Ga, Li, Nb, Rb, Sb, Sn, Sr, Th, Ti, Tl, U, W, and Zr, totaling 43 items.

[0047] From the above monitoring indicators, indicators that meet the requirements for long-term monitoring are selected. In this embodiment, principal component analysis is used to reduce the dimensionality of the two monitoring periods to determine the final number of screening factors. Before performing principal component analysis on the monitoring indicators, the KMO test and Bartlett's sphericity significance calculation are performed on the two periods of data, respectively. The formulas for calculating the KMO test, Bartlett's sphericity test, and their degrees of freedom are shown in Equations I-III below:

[0048] Formula I;

[0049] In the formula, r ij Let α be the simple correlation coefficient between variable i and variable j; ij KMO is the partial correlation coefficient between variable i and variable j, and its value ranges from 0 to 1.

[0050] Formula II;

[0051] In the formula, n is the sample size, p is the number of variables, and |R| is the determinant of the sample covariance matrix;

[0052] Formula III;

[0053] In the formula, p is the number of variables;

[0054] x calculated using equation II 2 The significance p-value is calculated by comparing the chi-square distribution corresponding to df obtained from Equation III.

[0055] Table 1 shows the correlation results of various indicators of the two periods of data provided in this embodiment. As can be seen from the results in the table, the test results of the two periods of data are within the range of KMO value > 0.5 and Bartlett's sphericity test significance P value < 0.05, indicating that there is a correlation between the indicators. The number of principal components with a cumulative contribution rate of more than 85% can be selected as the number of screening factors. Finally, the number of screening factors for various indicators in the two periods is determined to be 23.

[0056] Table 1. Correlation of various indicators between the two periods

[0057]

[0058] Example 2.

[0059] Based on the number of screening factors for various indicators determined in Example 1, this embodiment conducts repeatability tests on 48 soil monitoring samples from 2015 and 2021 to initially identify and eliminate indicators without change, and preliminarily screen out indicators with change that meet the requirements for long-term monitoring. The repeatability test is completed by calculating the relative double difference and the rate of no change between monitoring data of the same indicator and the same location in each two periods. The formulas for calculating the relative double difference and the rate of no change are shown in Equations IV and V below:

[0060] Formula IV;

[0061] In the formula, A1 represents the result of the first monitoring; A2 represents the result of the second monitoring.

[0062] Formula V;

[0063] When the relative double difference allowable limit RD ≤ 40%, the monitoring data is within the acceptable range of parallel samples and is considered to have no change. The statistical results of the no-change rate for each indicator are shown in Table 2 below. As can be seen from the table, the no-change rates for the detection indicators Li, Tl, Cu, Ga, Ge, Th, V, Y, Be, Co, Rb, Sc, and Ti are all 90% or higher. The higher the no-change rate, the smaller the degree of change. The change rates of the above 13 indicators are all within the acceptable range of parallel samples, and their changes can be ignored and considered as no change. These indicators are excluded from long-term high-frequency soil monitoring indicators. Therefore, excluding the above-mentioned excluded indicators, the long-term monitoring indicators for beneficial trace elements can be preliminarily determined to be B, Br, Cl, Mn, and Mo (5 items, the same number of beneficial element screening factors determined in Example 1), and the long-term monitoring indicators for rare earth elements are Ce and La (2 items, the same number of rare earth element screening factors determined in Example 1), for a total of 7 preliminarily determined long-term monitoring indicators. The remaining change indicators were further screened and eliminated to ensure that the number of long-term monitoring indicators was consistent with the number of screening factors determined in Example 1 (4 heavy metal indicators, 2 nutrient indicators, 2 health elements, and 8 other metal elements).

[0064] Table 2 Results of No Change Rate for Each Indicator

[0065]

[0066] Example 3.

[0067] Since the number of screening factors determined in Example 1 (4 heavy metal indicators, 2 nutrient indicators, 2 health elements, and 8 other metal elements) is less than the number of long-term monitoring indicators determined in Example 2 (7 heavy metal indicators, 3 nutrient indicators, 3 health elements, and 10 other metal elements), this example constructs a random forest regression model to evaluate the predictive performance of the remaining change indicators in Example 2 that were not determined to be long-term monitorable. The screening factors with the best performance evaluation are removed until the number of long-term monitoring indicators equals the number of screening factors (23). The steps for constructing the random forest regression model to evaluate the predictive performance of the screening factors in this example are as follows:

[0068] P1: The data from the two periods are integrated into one dataset. The 23 screening factor indicators are used as dependent variables according to their categories, and other indicators are used as independent variables. The dataset is divided into training and test sets in a 4:1 ratio, and random forest regression simulation is performed to obtain the predicted value of each indicator.

[0069] P2: Calculate the mean absolute error Mae, root mean square error RMSE, and fit index R for each indicator's predicted value. 2Among them, the mean absolute error Mae, the root mean square error RMSE, and the fit index R0 2 The calculation formulas are shown in equations VI-VIII below:

[0070] Formula VI;

[0071] Equation VII;

[0072] Formula VIII; where, ;

[0073] In the formula, n is the total number of samples; y1 i y2 represents the true value. i These are predicted values.

[0074] P3: When the number of screening factors for each type of indicator is greater than the number of long-term monitoring indicators for each type initially determined, the screening factor with the best expected effect calculated in P2 is removed. If the number of screening factors is still greater than the number of long-term monitoring indicators initially determined, return to step P2 until the number of long-term monitoring indicators is equal to the number of screening factors for each type of indicator in Example 1, and the process ends.

[0075] (1) According to the results in Table 2 of Example 2, the indicators of change of heavy metal elements include seven indicators to be screened: As, Cd, Cr, Hg, Ni, Pb, and Zn. The number of indicators for screening heavy metal elements determined in Example 1 is four. Based on the above examples, Table 3 below shows the results of further screening of heavy metal elements that meet the requirements for long-term monitoring using the random forest regression model in this example.

[0076] Table 3 Statistical index values ​​under different dependent variables

[0077]

[0078] Table 3 shows that Ni has the best fitting and prediction effect when it is used as the dependent variable, so it is removed. At this time, there are 6 targets to be screened, which is still greater than the number of determined indicators (4). The remaining indicators need to be recalculated using the random forest regression model, and the results are shown in Table 4 below.

[0079] Table 4 Statistical index values ​​under different dependent variables

[0080]

[0081] Table 4 shows that Pb performed best in terms of prediction when used as the dependent variable, so it was removed. At this point, there were 5 targets to be screened, which is still greater than the number of indicators determined (4). The remaining indicators need to be predicted using a random forest regression model, and the results are shown in Table 5. The results also show that As performed best in terms of prediction when used as the dependent variable, so it was removed. At this point, there were 4 targets to be screened, the same as the number of indicators determined (4). The final long-term monitoring indicators for heavy metal elements are Cd, Cr, Hg, and Zn.

[0082] Table 5 Statistical index values ​​under different dependent variables

[0083]

[0084] (2) According to the results in Table 2 of Example 2, the nutrient index changes include three indicators to be screened: N, P, and S. Based on the number of screening factors determined in Example 1, the final number of indicators for heavy metal element categories is two. Based on the above examples, Table 6 below shows the results of further screening of nutrient indicators suitable for long-term monitoring using the random forest regression model in this example. The results show that S has the best fitting and prediction effect when used as the dependent variable, so it is removed. At this time, there are two targets to be screened, which is the same as the number of indicators determined (2). The final long-term monitoring indicators for nutrient elements are N and P.

[0085] Table 6 Statistical index values ​​under different dependent variables

[0086]

[0087] (3) According to the results in Table 2 of Example 2, the indicators of change in health elements include three indicators to be screened: F, I, and Se. Based on the number of screening factors determined in Example 1, the final number of indicators for the health element category is two. Based on the above examples, Table 7 below shows the results of further screening of nutrient indicators suitable for long-term monitoring using the random forest regression model in this example. The results show that I has the best fitting and prediction effect when used as the dependent variable, so it is removed. At this time, there are two targets to be screened, which is the same as the number of indicators determined (two). The final long-term monitoring indicators for health elements are F and Se.

[0088] Table 7 Statistical index values ​​under different dependent variables

[0089]

[0090] (4) According to the results in Table 2 of Example 2, the change indicators of other metal elements include 10 indicators to be screened: Ag, Ba, Bi, Nb, Sb, Sn, Sr, U, W, and Zr. Based on the number of screening factors determined in Example 1, the final number of indicators for other metal element categories is 8. Based on the above examples, Table 8 below shows the results of further screening of nutrient indicators that meet long-term monitoring requirements using the random forest regression model in this example.

[0091] Table 8 Statistical index values ​​under different dependent variables

[0092]

[0093] Table 8 shows that Nb has the best fit and prediction effect when used as the dependent variable, so it is removed. At this point, there are 9 targets to be screened, which is still greater than the number of determined indicators (8). The remaining indicators need to be predicted and calculated, and the results are shown in Table 9. Table 9 shows that Sr has the best fit and prediction effect when used as the dependent variable, so it is removed. At this point, there are 7 targets to be screened, the same as the number of determined indicators (7). The 7 long-term monitoring indicators for other metal elements are Ag, Ba, Bi, Sb, Sn, U, W, and Zr.

[0094] Table 9 Statistical index values ​​under different dependent variables

[0095]

[0096] P4: Integrate the various categories of indicators obtained from P3 to obtain the final long-term monitoring indicators for soil geochemistry. The long-term monitoring indicators for soil geochemistry in this area include Cd, Cr, Hg, Zn, N, P, F, Se, B, Br, Cl, Mn, Mo, Ce, La, Ag, Ba, Bi, Sb, Sn, U, W, and Zr, totaling 23 items.

[0097] In summary, this invention integrates principal component analysis, repeatability testing, and random forest regression models to form a closed-loop indicator selection path of data-driven, stability verification, and intelligent screening. Principal component analysis, supplemented by KMO and Bartlett's sphericity significance tests, enables a scientific evaluation of the correlation and representativeness among soil monitoring indicators. Simultaneously, by combining regional land evolution characteristics and historical multi-period data, the invention uses repeatability testing to eliminate stable indicators lacking temporal representativeness, thereby constructing a monitoring system with a reasonable indicator structure and high sensitivity to change. Further integration with the random forest regression model gradually eliminates variables with minimal contribution to model prediction, ultimately determining the optimal combination of long-term monitoring indicators. This dual screening logic, from correlation analysis to time sensitivity identification, significantly improves the scientific rigor, practicality, and automation level of the monitoring system. The above descriptions of the embodiments are merely for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims of this invention.

[0098] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for screening long-term monitoring indicators of soil geochemistry, characterized in that, Includes the following steps: S1: Principal component analysis is used to reduce the dimensionality of at least two periods of monitoring data to determine the number N of screening factors for each type of indicator; S2: Based on the repeatability test method, identify and eliminate indicators with no change, and initially screen out the change indicators M that meet the requirements for long-term monitoring; S3: If M is less than or equal to N, the final monitoring indicators are obtained through screening; when N is greater than M, a random forest regression model is constructed to evaluate the predictive performance of the screened factors, and the screened factors with the best expected effect are eliminated until N=M, and the final monitoring indicators are obtained through screening. The repeatability test method described in S2 is achieved by calculating the relative double difference and the rate of no change between two monitoring periods for the same indicator and the same location. The calculation formulas for the relative double difference and the rate of no change are shown in Equations IV and V below: Formula IV; In the formula, A1 represents the result of the first monitoring; A2 represents the result of the second monitoring. Formula V.

2. The method for screening long-term monitoring indicators of soil geochemistry according to claim 1, characterized in that, Before the principal component analysis described in S1, the KMO test and Bartlett's sphericity significance calculation are performed. When the KMO value is >0.5 and the Bartlett's sphericity significance p-value is <0.05, there is a certain correlation between the indicators, and principal component data factor analysis can be performed. If the correlation between the indicators does not meet the conditions of KMO value >0.5 and Bartlett's sphericity significance p-value <0.05, the indicators are classified, and principal component analysis is performed under the premise of indicators in the same category.

3. The method for screening long-term monitoring indicators of soil geochemistry according to claim 1, characterized in that, In S1, principal component analysis is performed on the monitoring data for each period separately, and the number of screening factors for each period is confirmed. The number of screening factors in each period is compared, and the maximum value is selected as the final number of screening factor indicators. The data dimensionality reduction is to select the number of principal components with a cumulative contribution rate of more than 85% as the number of screening factor indicators.

4. The method for screening long-term monitoring indicators of soil geochemistry according to claim 1, characterized in that, The total monitoring period should be at least 5 years.

5. The method for screening long-term monitoring indicators of soil geochemistry according to claim 2, characterized in that, The calculation formula for the KMO test is shown in Equation I below: Formula I; In the formula, r ij Let α be the simple correlation coefficient between variable i and variable j; ij KMO is the partial correlation coefficient between variable i and variable j, and its value ranges from 0 to 1.

6. The method for screening long-term monitoring indicators of soil geochemistry according to claim 2, characterized in that, The Bartlett sphericity significance calculation steps include Bartlett sphericity test calculation and Bartlett test degrees of freedom calculation; The Bartlett sphericity test calculation formula is shown in Equation II below: Formula II; In the formula, n is the sample size, p is the number of variables, and |R| is the determinant of the sample covariance matrix; The formula for calculating the degrees of freedom of the Bartlett test is shown in Equation III below: Formula III; In the formula, p is the number of variables; x calculated using equation II 2 The significance p-value is calculated by comparing the chi-square distribution corresponding to df obtained from Equation III.

7. The method for screening long-term monitoring indicators of soil geochemistry according to claim 1, characterized in that, When the relative double difference allowable limit RD ≤ 40%, the monitoring data is within the acceptable range of parallel samples and is considered to have no change; when the rate of no change calculated from two adjacent monitoring data periods is ≥ 90%, the indicator can be considered to have no change and the indicator can be determined to be a deductible indicator.

8. The method for screening long-term monitoring indicators of soil geochemistry according to claim 1, characterized in that, The steps for constructing a random forest regression model in S3 to evaluate the predictive performance of the selected factors are as follows: P1: Integrate data from at least two periods into one dataset, use N screening factor indicators as dependent variables in sequence, and other indicators as independent variables; divide the dataset into training and test sets in a 4:1 ratio and perform random forest regression simulation to obtain the predicted value of each indicator; P2: Calculate the mean absolute error Mae, root mean square error RMSE, and fit index R for each indicator's predicted value. 2 ; P3: Remove the screening factor with the best expected effect calculated in P2. If N is still greater than M, return to step P2 until N=M and the process ends. P4: Integrate the various categories of indicators obtained from P3 to obtain the final long-term monitoring indicators for soil geochemistry.

9. The method for screening long-term monitoring indicators of soil geochemistry according to claim 8, characterized in that, The mean absolute error Mae, root mean square error RMSE, and fit index R 2 The calculation formulas are shown in equations VI-VIII below: Formula VI; Equation VII; Formula VIII; where, ; In the formula, n is the total number of samples; y1 i y2 represents the true value. i These are predicted values.

Citation Information

Patent Citations

  • Water eutrophication prediction method based on improved MIMO depth dual 3Q learning

    CN116307255A

  • Method for screening water quality risk factors of wetland in cold region

    CN119646496A