Soil heavy metal pollution ecological risk assessment method based on machine learning and application
Through machine learning-based methods, integrating multiple indexes and combining PCA dimensionality reduction and machine learning algorithms, a multi-classification ecological risk assessment model for heavy metals in soil is constructed, solving the problems of low efficiency and poor accuracy in the existing technology, and achieving more efficient and accurate assessment of heavy metals in soil pollution.
Patent Information
- Application Number
- CN202510115705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is inefficient and poorly accurate in the ecological risk assessment of soil heavy metal pollution, and cannot meet the demand of modern environmental assessment for accurate and comprehensive grasp of pollution conditions.
Using a machine learning-based method, a multi-classified ecological risk assessment model for soil heavy metals is constructed. By integrating and improving the biomarker index, pollution load index, Nemero index, potential ecological risk index and geological accumulation index, combined with PCA dimensionality reduction and machine learning algorithms, a multi-index data set is constructed and evaluated.
The accuracy and efficiency of the assessment of ecological risk of soil heavy metal pollution has been improved. The comprehensive index R2 reaches 0.94, which is higher than the single index, which can more accurately reflect the soil pollution status.
Smart Images

Figure CN120146552A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ecological risk assessment, and particularly to a method and application for ecological risk assessment of soil heavy metal pollution based on machine learning. Background Art
[0002] In recent decades, with the rapid development of urbanization and industrialization, a series of environmental problems have emerged, such as factory emissions, mining activities, and increasing traffic. Among them, soil heavy metal pollution has become a key focus of observation. Soil is an important part of the entire terrestrial ecosystem of the earth and is also the most basic natural resource on which humans rely for survival. Since heavy metals are non-degradable in the environment, they will accumulate in organisms until they exceed the critical value, thereby causing biological toxicity and ultimately affecting human health. Therefore, the ecological assessment of soil heavy metal pollution has attracted increasing public attention.
[0003] In the field of ecological risk assessment, traditional assessment methods usually adopt a single index or a simple combination of multiple indexes. The single index is limited to a specific dimension and has obvious defects; while the simple combination of multiple indexes lacks a rigorous logic, and in essence, it is still similar to a single index, that is, although it seems to be multi-dimensional, it is not effectively integrated, and still cannot meet the needs of modern environmental assessment for accurately and comprehensively grasping the pollution situation. Summary of the Invention
[0004] The main object of the present invention is to provide a method and application for ecological risk assessment of soil heavy metal pollution based on machine learning, aiming to solve the problems of low efficiency and poor accuracy in the existing technology for ecological risk assessment of soil heavy metal pollution.
[0005] To achieve the above object, the present invention provides a method for ecological risk assessment of soil heavy metal pollution based on machine learning, including the steps of:
[0006] S1, constructing a metal data set; the metal data set includes 0.5, 1, 2, 5, and 10 times the allowable content values of the heavy metals to be measured in the environmental quality standards; the heavy metals to be measured include copper, zinc, lead, cadmium, mercury, and arsenic.
[0007] S2, obtaining an improved biomarker index.
[0008] S3, integrating the improved biomarker index in step S2 with the pollution load index, Nemerow index, potential ecological risk index, and improved geoaccumulation index into a multi-index data set.
[0009] S4, using PCA to reduce the dimension of the multi-index data set in step S3, taking the principal component with the largest variance in the data as the dimensionality reduction result and classifying according to the dimensionality reduction result to obtain a PCA classification result.
[0010] S5. Integrate the metal data set with the PCA classification results in step S4 to form an ecological risk assessment data set; and randomly divide the ecological risk assessment data set into a training set and a test set; where the training set accounts for 70% and the test set accounts for 30%.
[0011] S6. Provide a machine learning algorithm and use the machine learning algorithm to screen the training set and the test set in step S5 to construct a multi-class ecological risk assessment model for soil heavy metals; where the machine learning algorithm includes random forest, support vector machine, decision tree, k-nearest neighbor, and naive Bayes model.
[0012] S7. Use the multi-class ecological risk assessment model for soil heavy metals in step S6 to conduct an ecological risk assessment of heavy metal pollution in the target soil.
[0013] Further, in step S2, the method for obtaining the improved biomarker index is as follows: Construct an initial biomarker data matrix X based on the initial data of earthworm biomarkers and soil metals obtained through literature search.
[0014] Standardize the data in the initial biomarker data matrix and calculate the correlation coefficient r of the biomarkers ij , conflict information A j , the change level AL and change index S of the biomarkers j , and obtain the overall information index C according to the calculation results j .
[0015] Calculate the fine-tuned weight W of the biomarker according to the basic weight W of the biomarker n j .
[0016] Among them, i refers to a certain day; j refers to a certain biomarker;
[0017] Establish an evaluation index set U = {U1 U2 U3 U4} and a fuzzy comment set V = {V1 V2 V3 V4}; and establish a fuzzy comprehensive matrix through the established evaluation index set and the fuzzy comment set
[0018] Among them, U refers to different biomarkers; V refers to the change level of the biomarker compared with the control group; V1 means that when |AL| ≤ 20%, its score is 4 points; V2 means that when 20% < |AL| ≤ 50%, its score is 3 points; V3 means that when 50% < |AL| ≤ 100%, its score is 2 points; V4 means that when |AL| > 100%, its score is 4 points;
[0019] The biomarker index CFBRI is calculated by using the weight vector W after fine-tuning and the fuzzy comprehensive matrix; wherein,
[0020]
[0021] Provide the standardized metal concentration, and fit it with the biomarker index to generate a fitting curve; standardize the metal data set in the standardization step S1, and substitute the metal data set into the fitting curve to obtain the improved biomarker index.
[0022] Furthermore, in step S2, the objective function of the initial biomarker data matrix X is
[0023]
[0024] The standardized initial biomarker data matrix X ij ′s objective function is
[0025]
[0026] Furthermore, in step S2, the correlation coefficient r ij of the biomarker, the conflict information A j , the change level AL of the biomarker, the change index S j and the overall information index C j 's objective functions are respectively:
[0027]
[0028] X ij ″ = |AL|;
[0029]
[0030] C j = A j ·S j (j = 1, 2, 3, …, n).
[0031] Furthermore, in step S2, the objective function of the weight W j after fine-tuning of the biomarker is
[0032]
[0033] The objective function of the standardized metal concentration is
[0034]
[0035] Further, in step S3, the other traditional indices are the pollution load index, the Nemerow index, the potential ecological risk index, and the modified geo-accumulation index.
[0036] Further, in step S4, after the step of taking the principal component with the largest variance in the data as the dimensionality reduction result, it further includes classifying the dimensionality reduction result according to the numerical quartiles to obtain the PCA classification result; wherein the PCA classification results are successively mild, moderate, severe, and extremely severe.
[0037] Further, in step S5, it further includes performing balanced dataset processing and cross-validation processing on the results of the training set.
[0038] Further, in step S6, the machine learning algorithm is obtained in the following way:
[0039] Using the target index to evaluate the model performance and screening the optimal solution to obtain the corresponding machine learning algorithm; wherein the target indices include the area under the curve, accuracy, sensitivity, specificity, precision, F1 score, macro F1 score, Kappa coefficient, and Hamming loss.
[0040] Further, in step S6, the machine learning algorithm is the random forest.
[0041] The present invention also provides an application of the machine learning-based soil heavy metal pollution ecological risk assessment method described in any one of the above in the classification and grading planning of soil heavy metal pollution.
[0042] The beneficial effects achieved by the present invention:
[0043] The machine learning-based soil heavy metal pollution ecological risk assessment method of the present invention constructs a multi-index dataset containing five indices - a multi-index dataset based on six metal contents of 0.5, 1, 2, 5, and 10 times of the "Environmental Quality Standards for Agricultural Land" and the corresponding pollution load index, Nemerow index, potential ecological risk index, geo-accumulation index, and modified biomarker index. For this multi-index dataset, PCA is used to achieve the determination of coupled weights and dimensionality reduction processing based on the eigen-decomposition of the data covariance matrix, and then combined with machine learning modeling to construct a soil heavy metal multi-classification ecological risk assessment model for the ecological risk assessment of heavy metal pollution in the target soil. The result shows that the comprehensive index R 2 is 0.94, higher than other single indices. This method combines the advantages of multiple biological and non-biological indices, solves the limitations of single indices in dealing with complex pollution situations, and has high accuracy.
[0044] Applying the ecological risk assessment method for soil heavy metal pollution based on machine learning provided by the present invention to the classification planning of soil heavy metal pollution areas can efficiently and accurately control and delimit soil heavy metal pollution areas; it has strong applicability and is suitable for large-scale popularization and use. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0046] Figure 1 It is a comparison chart of the evaluation proportion of each index in the multi-index dataset during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention;
[0047] Figure 2 It is a comparison chart of the weight proportion of each index in the multi-index dataset during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention;
[0048] Figure 3 It is a PCA classification result chart during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention;
[0049] Figure 4 It is a comparison chart of the data distribution in the training set and the test set of the ecological risk assessment dataset during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention;
[0050] Figure 5 It is a cross-validation performance chart of the soil heavy metal multi-classification ecological risk assessment model using different machine learning algorithms during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention;
[0051] Figure 6 It is a cross-validation performance radar chart of the soil heavy metal multi-classification ecological risk assessment model using different machine learning algorithms during the construction process of the soil heavy metal multi-classification ecological risk assessment model of the present invention; among them, (a) is the cross-validation performance radar chart of the soil heavy metal multi-classification ecological risk assessment model using different machine learning algorithms for sensitivity; (b) is the cross-validation performance radar chart of the soil heavy metal multi-classification ecological risk assessment model using different machine learning algorithms for specificity; (c) is the cross-validation performance radar chart of the soil heavy metal multi-classification ecological risk assessment model using different machine learning algorithms for accuracy.
[0052] The realization, functional features, and advantages of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific Embodiments
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] It should be noted that, without conflict, the following embodiments and the features in the embodiments may be combined with each other. It should also be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific embodiments, rather than limiting the scope of protection of the present invention.
[0055] Unless otherwise defined, all technical and scientific terms used in the present invention are the same as those understood by those of ordinary skill in the technical field of the present invention and the description of the present invention. Any methods, devices, and materials similar or equivalent to the methods, devices, and materials described in the embodiments of the present invention can also be used to implement the present invention.
[0056] When the embodiments give a numerical range, it should be understood that, unless otherwise specified in the present invention, any value between the two endpoints of each numerical range and any one of the two endpoints can be selected. The test methods without specific conditions noted in the following embodiments are usually carried out under conventional conditions or according to the conditions recommended by each manufacturer. The materials or reagents required in the following embodiments are commercially available unless otherwise specified.
[0057] In order to solve the problems of low efficiency and poor accuracy in the existing technology for ecological risk assessment of soil heavy metal pollution, the present invention provides a method for ecological risk assessment of soil heavy metal pollution based on machine learning, including the steps of:
[0058] S1, constructing a metal data set; the metal data set includes 0.5, 1, 2, 5, and 10 times the allowable content values of the heavy metals to be measured in the environmental quality standards; the heavy metals to be measured include copper, zinc, lead, cadmium, mercury, and arsenic. Specifically, the metal data set is the data of six metals (copper, zinc, lead, cadmium, mercury, and arsenic) at 0.5, 1, 2, 5, and 10 times the "Environmental Quality Standards for Agricultural Land" (GB 15618-2008).
[0059] S2, obtaining an improved biomarker index.
[0060] S3. Integrate the improved biomarker index in step S2 with the pollution load index, Nemerow index, potential ecological risk index, and improved geo-accumulation index into a multi-index dataset.
[0061] S4. Use PCA to reduce the dimension of the multi-index dataset in step S3, take the principal component with the largest variance in the data as the dimensionality reduction result, and classify according to the dimensionality reduction result to obtain the PCA classification result. Specifically, use PCA with the "tidymodels" package in R language to reduce the dimension of the multi-index dataset in step S3, and take the principal component with the largest variance in the data as the dimensionality reduction result to obtain the PCA classification result.
[0062] S5. Integrate the metal dataset with the PCA classification result in step S4 into an ecological risk assessment dataset; and randomly divide the ecological risk assessment dataset into a training set and a test set; among them, the training set accounts for 70% and the test set accounts for 30%.
[0063] S6. Provide machine learning algorithms, and use the machine learning algorithms to screen the training set and test set in step S5 to construct a multi-class ecological risk assessment model for soil heavy metals; among them, the machine learning algorithms include random forest, support vector machine, decision tree, k-nearest neighbor, and naive Bayes model.
[0064] S7. Use the multi-class ecological risk assessment model for soil heavy metals in step S6 to conduct an ecological risk assessment of heavy metal pollution in the target soil.
[0065] The ecological risk assessment method for soil heavy metal pollution based on machine learning according to the present invention constructs a dataset containing five indices - a multi-index dataset - according to six metal contents of 0.5, 1, 2, 5, and 10 times of the "Environmental Quality Standards for Agricultural Land" and the corresponding pollution load index, Nemerow index, potential ecological risk index, geo-accumulation index, and improved biomarker index. For this multi-index dataset, PCA is used to achieve coupled weight determination and dimensionality reduction processing based on the eigen-decomposition of the data covariance matrix, and then combined with machine learning modeling to construct a multi-class ecological risk assessment model for soil heavy metals to conduct an ecological risk assessment of heavy metal pollution in the target soil. The result shows that the comprehensive index R 2 is 0.94, higher than other single indices. This method combines the advantages of multiple biological and non-biological indices, solves the limitations of single indices in dealing with complex pollution situations, and has high accuracy.
[0066] Specifically, during the research process, it was found that the ecological risk assessment of heavy metal pollution in soil mainly covers two categories: non-biological methods (such as potential ecological, Nemerow and other index methods) and biological methods (involving animals, plants, microorganisms, and soil enzymes). The non-biological method detects the metal content in the soil and judges the degree of pollution in combination with standards and background values, which is easy to operate. Among them, the Nemerow index and the pollution load index can intuitively reflect the overall intensity of pollution. The geological accumulation index considers the natural diagenesis process, making the assessment close to reality. The potential ecological risk index considers metal toxicity and has a certain warning effect on the degree of pollution damage. However, they also have their own shortcomings. For example, the Nemerow and pollution load indices cannot reflect metal toxicity and accumulation. The geological accumulation index does not consider metal toxicity, and the metal toxicity parameters of the potential ecological risk index are uncertain.
[0067] Biological methods have unique advantages. They are not limited to soil metal content. They can not only accurately reflect the immediate impact of pollution by capturing subtle dynamic changes in biological indicators under multi-metal pollution through high-precision and continuous monitoring of the physiological characteristics of plants, animals and microorganisms and soil enzyme activity, but also comprehensively consider the synergistic effects of multiple biological pollution, outline the overall picture of soil multi-metal pollution, and make assessments more accurate. However, except for a few biological methods such as earthworm method, which are relatively mature in species selection, experimental process specification, and evaluation standard setting, and have formed a relatively complete system, most biological methods still have many shortcomings that need to be improved. Take the biomarker response index (BRI) as an example. It scores and weights the comprehensive index based on the changes in biomarkers in the pollution treatment group relative to the control group. Although the system is relatively complete, it is not perfect and has obvious defects: focusing only on the performance of biomarkers on a certain day, the BRI values measured on different dates for the same soil sample may be very different; and the data of biomarkers are naturally irregular, and it is easy to have abnormal situations where the corresponding biomarker scores for high pollution levels are higher than those for low pollution levels. To solve this problem, the present invention introduces fuzzy comprehensive evaluation and CRITIC weight method to improve BRI and create a new CFBRI model (see step S2 for details). The improved biomarker index obtained in step S2 of the present invention solves the problem of complex reactions of biomarkers (such as enzyme activity). Even if a complex dose-effect relationship occurs, it can be reasonably utilized by fine-tuning the weight. Compared with the traditional BRI, the earthworm biomarker assessment of polymetallic contaminated soil has more accurate ecological risk assessment results under medium and high pollution, which improves the accuracy of ecological risk assessment of polymetallic pollution in soil.
[0068] Moreover, both the abiotic method and the biotic method have their own disadvantages and advantages, and they both conduct ecological risk assessments of soil heavy metal pollution from different perspectives. Therefore, it is impossible to accurately conduct an ecological risk assessment of soil metal pollution by using any one method alone. In addition, most studies on the ecological risk assessment of soil heavy metals only simply combine multiple index methods. Due to the lack of a rigorous integration logic in the simple combination, its essence is still a single index, and it is difficult to break through the limitation of the one-dimensional perspective. The traditional multi-index combination method is usually the weighted comprehensive method, but the traditional weighted comprehensive method exposes many drawbacks. Analyzing from the level of data characteristics, the high-dimensional data contained in multiple indexes is filled with a large amount of redundant information and noise interference. These invalid information not only greatly increases the computational complexity and consumes a large amount of computing resources, but also makes the key data features hidden in the complex data ocean and difficult to be accurately extracted. As a result, the evaluation results deviate greatly and cannot truly reflect the actual situation of soil pollution.
[0069] In view of this situation, after obtaining the multi-index data set (step S3), the present invention uses the principal component analysis (PCA) technology (see step S4 in detail) to perform dimensionality reduction processing on the multi-index data set (based on the precise eigenvalue decomposition operation of the data covariance matrix). Through this unique mathematical transformation, PCA can skillfully project high-dimensional data into a low-dimensional space. In this process, like a precise "data filter", while retaining the main variance information of the data, it effectively eliminates redundancy and noise, greatly reduces the computational complexity, and enables the key features of the data to be clearly presented. This dimensionality reduction method based on the internal structure of the data can, compared with other conventional dimensionality reduction methods, maximize the maintenance of the inherent information of the data and ensure the accuracy of subsequent evaluations. This is the core reason for preferentially selecting PCA for dimensionality reduction.
[0070] Delving deeper into the weight determination mechanism, this is a key link in the multi-index combination evaluation. During the PCA dimensionality reduction process, according to the statistical characteristics of the data itself, each index is automatically assigned an initial weight adapted to it. This process makes full use of the internal correlation of the data, making the weight distribution have a solid objective basis. Although PCA demonstrates excellent dimensionality reduction ability in the process of ecological risk assessment of soil heavy metal pollution, however, the PCA technology has inherent shortcomings. Its model architecture and principle determine that it is unable to perform prediction tasks and cannot prospectively predict the pollution status and development trend of unknown soil samples based on existing data.
[0071] In view of this, referring to the operation in step S6 of the present invention, a machine learning algorithm is introduced for modeling. With its powerful self-learning and adaptive capabilities, the machine learning algorithm can deeply explore the complex patterns and rules hidden behind massive data. By constructing a model architecture containing multiple algorithms, such as random forest, decision tree, etc., the data and metal concentration after PCA dimensionality reduction and classification processing are used for learning and training. The model gradually masters the non-linear relationship between the comprehensive index classification result of the combined multi-index of soil heavy metal pollution and the metal concentration, and then has the function of accurate prediction, effectively making up for the deficiency of PCA in the prediction dimension, and jointly promoting the ecological risk assessment of soil heavy metal pollution towards a more accurate and efficient direction with PCA.
[0072] Furthermore, in step S2, the acquisition method of the biomarker index is improved as follows: the initial biomarker data matrix X is constructed according to the initial data of earthworm biomarkers and soil metals obtained through literature search.
[0073] The data in the initial biomarker data matrix are standardized, and the correlation coefficient r of the biomarker is calculated ij , conflict information A j , the change level AL and change index S of the biomarker j , and the overall information index C is obtained according to the calculation result j . Preferably, the change level AL of the biomarker is ≥ 100%, and the change level AL of the biomarker is always 1.
[0074] According to the basic weight W of the biomarker n the fine-tuned weight W of the biomarker is calculated j .
[0075] where i refers to a certain day; j refers to a certain biomarker;
[0076] The evaluation index set U = {U1 U2 U3 U4} and the fuzzy comment set V = {V1 V2 V3 V4} are established; and the fuzzy comprehensive matrix is established through the established evaluation index set and fuzzy comment set
[0077] where U refers to different biomarkers; V refers to the change level situation of the biomarker and the control group; V1 means that when |AL| ≤ 20%, its score is 4 points; V2 means that when 20% < |AL| ≤ 50%, its score is 3 points; V3 means that when 50% < |AL| ≤ 100%, its score is 2 points; V4 means that when |AL| > 100%, its score is 4 points;
[0078] The biomarker index CFBRI is calculated through the fine-tuned weight vector W of the biomarker and the fuzzy comprehensive matrix; where,
[0079]
[0080] Provide a standardized metal concentration and fit it with a biomarker index to generate a fitting curve; standardize the metal data set in step S1 and substitute the metal data set into the fitting curve to obtain an improved biomarker index.
[0081] Further, in step S2, the objective function of the initial biomarker data matrix X is
[0082]
[0083] The standardized initial biomarker data matrix X ij ′ has an objective function of
[0084]
[0085] Further, in step S2, the correlation coefficient r ij , conflict information A j , the change level AL, change index S j and overall information index C j have objective functions respectively as follows:
[0086]
[0087] X ij ″ = |AL|;
[0088]
[0089] C j = A j · S j (j = 1, 2, 3, …, n).
[0090] Further, in step S2, the objective function of the weight W j after fine-tuning of the biomarker is
[0091]
[0092] The objective function of the standardized metal concentration is
[0093]
[0094] Further, in step S4, after the step of taking the principal component with the largest variance in the data as the dimensionality reduction result, it further includes classifying the dimensionality reduction result according to the numerical quartiles to obtain a PCA classification result; wherein, the PCA classification results are mild, moderate, severe, and critical in sequence.
[0095] Further, in step S5, it also includes processing the results of the training set for balanced dataset and cross-validation. Specifically, the results of the training set use the "themis" package in R language to solve the problem of unbalanced datasets. The cross-validation model is constructed using the "caret" package in R language to ensure a reliable evaluation of its performance.
[0096] Further, in step S6, the machine learning algorithm is obtained as follows:
[0097] The performance of the model is evaluated using target indices, and the optimal solution is screened to obtain the corresponding machine learning algorithm; the target indices include area under the curve, accuracy, sensitivity, specificity, precision, F1-score, macro F1-score, Kappa coefficient, and Hamming loss. Specifically, the objective functions of area under the curve (AUC), accuracy, sensitivity, specificity, precision, F1-score, macro F1-score, Kappa coefficient, and Hamming loss are as follows:
[0098]
[0099] F1marco = ∑ i F1 i / i
[0100]
[0101] Further, in step S6, the machine learning algorithm is random forest.
[0102] The present invention also provides an application of the ecological risk assessment method for soil heavy metal pollution based on machine learning as described in any one of the above in the classification and zoning planning of soil heavy metal pollution.
[0103] Applying the ecological risk assessment method for soil heavy metal pollution based on machine learning provided by the present invention to the classification and zoning planning of soil heavy metal pollution can efficiently and accurately control and delimit the soil heavy metal pollution areas; it has strong applicability and is suitable for large-scale popularization and use.
[0104] To further understand the present invention, the following is an example for illustration:
[0105] Example 1
[0106] Step 1 Verification of CFBRI
[0107] 1. Literature collection and data extraction: Literature was searched through the online databases Web of Science and CNKI. Peer-reviewed journal papers published between 1974 and 2024 (as of May 2024) were selected to verify the applicability of the biomarker index (CFBRI) model. To obtain studies on the effects of earthworm biomarkers on metal pollutants in soil, the following search strings were used: "earthworm" and "metal" and "biomarker". The literature studies considered in this CFBRI model must meet the following inclusion criteria: (1) simultaneously include earthworm biomarkers and metal pollutants; (2) the study must have an uncontaminated soil control group; (3) the total weight of the biomarkers must not be less than 6.5. The initial biomarker data in the literature studies were obtained through GetDataGraph Digitizer 2.25.
[0108] 2. Calculation of CFBRI and BRI values:
[0109] (1) Calculate the CFBRI value through the content of the technical solution of the present invention. Specifically:
[0110] Construct the initial biomarker data matrix X, and then standardize the initial biomarker data. First, calculate the correlation coefficient r of the biomarkers ij and the conflict information A j , then calculate the change level AL and change index S of the biomarkers j , and obtain the overall information index C j , and finally, according to the basic weight W of the biomarkers n , calculate the fine-tuned weight W of the biomarkers j . The objective functions of each item are as follows:
[0111] Initial biomarker data matrix X and standardization of initial biomarker data:
[0112]
[0113] Correlation coefficient r of biomarkers ij and conflict information A j :
[0114]
[0115] Change level AL and change index AL of biomarkers:
[0116]
[0117] X ij ″ = |AL|
[0118]
[0119] Overall information index C j :
[0120] C j = A j ·S j (j = 1, 2, 3, …, n)
[0121] Weight W of biomarker after fine-tuning j :
[0122]
[0123] Establish the evaluation index set U = {U1 U2 U3 U4} and the fuzzy comment set V = {V1 V2 V3 V4}. Through the index set and the comment set, establish the fuzzy comprehensive matrix. Through the weight vector W and the matrix R obtained in the previous step, the CFBRI value can be calculated. The objective function is as follows:
[0124] Fuzzy comprehensive matrix R:
[0125]
[0126] (2) The objective function for calculating the BRI value is as follows:
[0127]
[0128] 3. Standardized metal concentration: Standardize the metal concentration through the technical solution content of the present invention. The metal weights and agricultural soil standards are shown in Table 1. Specifically:
[0129] Constructed a data set (15,625 rows of data) of six metals (copper, zinc, lead, cadmium, mercury, and arsenic) at 0.5, 1, 2, 5, and 10 times the "Environmental Quality Standards for Agricultural Land" (GB 15618-2008). The collected soil metal data is either single-metal pollution or multi-metal pollution. For the convenience of subsequent use, standardize the soil metal concentration according to the national standard (GB36600-2018). The objective function is as follows:
[0130]
[0131] Table 1 Standardized metal concentration values, weights, and Chinese soil element background values of six metals
[0132]
[0133] 4. Comparison of CFBRI and BRI: Fit the standardized metal concentration (Metal) with the CFBRI value and the BRI value, as shown in Table 2.
[0134] Performance indicators of the fitting curves of each index in Table 2
[0135]
[0136] As can be seen from Table 2, the R of CFBRI 2 is 0.1 - 0.25 higher than the R of BRI 2 , and the R 2 only increases slightly. The following explanations are as follows: (1) The responses of earthworm biomarkers to metal toxicity are extremely irregular and complex; (2) The differences between different earthworm species; (3) The defect of metal normalization: the toxicities of different metals to earthworms are not the same. Therefore, the slight increase in R 2 can prove that CFBRI has more advantages in the field of using earthworm biomarkers to evaluate multi-metal contaminated soils.
[0137] Step 2 Comparative analysis of the sources of comprehensive indices and single indices
[0138] 1. Construction of multi-index dataset: The multi-index data is constructed through the technical solution content of the present invention. The standard values and background values in its function formula are based on Table 1.
[0139] Substitute the "normalized metal concentration value" in Step 1 into the "fitting curve formula of normalized metal concentration (Metal) and CFBRI value" to obtain the improved biomarker index for standby.
[0140] Provide other traditional indices, the pollution load index (PLI), the Nemerow index (PN), the potential ecological risk index (RI), and the improved geo-accumulation index (Igeo), which are calculated according to the data in the metal dataset, and their objective functions are as follows:
[0141] P i = C i / B 0
[0142]
[0143] P max = C max / KB
[0144]
[0145] Integrate the improved biomarker index with other traditional indices to form a multi-index dataset containing five indices in total, and the construction is completed.
[0146] 2. Descriptive statistical analysis of each index: Through the calculation of the maximum value, minimum value and mean value of the multi-index dataset and the evaluation criteria of each index according to Table 3, Table 4 (Descriptive statistical analysis table of each index) can be obtained and Figure 1Analysis results of (Comparison chart of the proportion of each index evaluation).
[0147] Table 3 Evaluation criteria for each index
[0148]
[0149] Table 4 Descriptive statistical analysis of each index
[0150]
[0151] The results show that there are significant differences between different indices in terms of both numerical values and evaluation analysis results. This may be due to the fact that different indices analyze pollution from different perspectives. Generally speaking, such significant differences between different indices may lead to differences in risk assessment, thus affecting the scientific and reasonable division of soil pollution areas. Therefore, the birth of a comprehensive index combining biological and non-biological multi-indices is reasonable and scientific.
[0152] 3. Dimensionality reduction and classification of PCA: The multi-index data is reduced in dimension and classified through the technical solution content of the present invention. Specifically, PCA is used to reduce the dimension of the multi-index data set using the "tidymodels" package in R language, and the principal component with the largest variance in the data is taken as the dimensionality reduction result to obtain the PCA classification result; the dimensionality reduction result is classified according to the numerical quartiles, and the classification results are successively mild, moderate, severe, and extremely severe.
[0153] 4. Comparative analysis of the comprehensive index and single indices: The comprehensive index is obtained by fitting the data of the largest principal component of PCA with the corresponding standardized metal concentration Metal. The other indices are also analyzed as above, and the fitting formula for all of them is ExpDec.2. However, for Igeo, the sum of all corresponding Igeo of metals is used to obtain a comprehensive Igeo, and the fitting analysis is carried out with this comprehensive Igeo. The results are shown in Table 2 in Step 1. The R 2 , MSE, and RMSE of the comprehensive index are basically higher than those of all other single indices except for the MSE and RMSE of PLI.
[0154] The purpose of the fitting analysis comparison in the present invention is to explore which index can better represent the information of metal data, rather than considering the prediction accuracy. Therefore, although the MSE and RMSE of PLI are slightly higher than those of the comprehensive index, the R 2 = 0.94 of the comprehensive index can explain 94% of the data variance, while the effect of PLI is slightly worse. Generally speaking, the comprehensive index has more advantages than other single indices.
[0155] Specifically, Figure 2is the weight ratio of each index in the comprehensive index, which indicates that the comprehensive index incorporates five indices. Among them, CFBRI, PLI, and PN are relatively important indices, while Igeo and RI are of secondary importance. The results of PCA classification are as Figure 3 shown. According to Figure 3 the results, regarding the general pollution situation of all points, moderate pollution is the most, severe pollution is the second most, light pollution is the third most, and serious pollution is the least.
[0156] Step 3: Construct a multi-class ecological risk assessment model for soil heavy metals (combining biological and abiotic indices)
[0157] 1. Construction and preprocessing of the ecological risk assessment dataset:
[0158] Integrate the aforementioned PCA classification results and the metal dataset into a dataset - the ecological risk assessment dataset.
[0159] Randomly divide the ecological risk assessment dataset into two groups: 70% as the training set and 30% as the test set. The "themis" package in R language is used to solve the problem of imbalanced datasets for the training set results. The cross-validation model is constructed using the "caret" package in R language to ensure a reliable assessment of its performance. To screen for a suitable model, five machine learning algorithms are used, namely random forest (RF), support vector machine (SVM), decision tree (DT), k-nearest neighbor (KNN), and naive Bayes model (NBM).
[0160] Through the SMOTE oversampling technique, the purpose of balancing the training set data is achieved. The specific comparison results of the data distribution in the training set and test set are as Figure 4 shown.
[0161] 2. Screening of the optimal multi-class ecological risk assessment model for soil heavy metals:
[0162] Select many target index pairs to evaluate the performance of models using five different machine learning algorithms (random forest (RF), support vector machine (SVM), decision tree (DT), k-nearest neighbor (KNN), and naive Bayes model (NBM)), specifically including: area under the curve (AUC), accuracy, sensitivity, specificity, precision, F1-score, macro F1-score, Kappa coefficient, and Hamming loss. The objective function is as follows:
[0163]
[0164] F1marco = ∑ iF1 i / i
[0165]
[0166] The cross - validation performance diagrams of the soil heavy metal multi - classification ecological risk assessment models with different machine learning algorithms are as Figure 5 shown; the radar diagrams of the cross - validation performance of the soil heavy metal multi - classification ecological risk assessment models with different machine learning algorithms are as Figure 6 shown.
[0167] According to Figure 5 and Figure 6 display, among the performances of the finally trained soil heavy metal multi - classification ecological risk assessment models with five different machine learning algorithms in terms of area under the curve (AUC), accuracy, sensitivity, specificity, precision, F1 - score, macro F1 - score, Kappa coefficient, and Hamming loss, higher AUC, accuracy, sensitivity, specificity, precision, F1 - score, macro F1 - score, and Kappa coefficient or smaller Hamming loss mean better model performance.
[0168] The results show that RF obtained higher AUC, sensitivity, and precision, as well as the highest accuracy, specificity, F1 - score, macro F1 - score, and Kappa coefficient, and the smallest Hamming loss. Only the sensitivity of NBM is better than that of RF.
[0169] In addition, the AUC and precision of SVM are better than those of RF, but RF is superior to SVM in other performance metrics. Therefore, considering all evaluation metrics comprehensively, the performance of RF is better. The AUC value of RF is between 0.967 and 0.992, the accuracy is 0.888, the sensitivity is between 0.536 and 0.909, the specificity is between 0.868 and 0.999, the precision is between 0.850 and 0.956, the F1 - score is between 0.687 and 0.909, the macro F1 - score is 0.829, the Kappa coefficient is 0.799, and the Hamming loss is 0.111.
[0170] In summary, in the soil heavy metal multi - classification ecological risk assessment model, the RF model can more effectively solve the problems of data non - linearity and inherent complex structure. The soil heavy metal multi - classification ecological risk assessment model using RF as the machine learning algorithm is the optimal model.
[0171] In conclusion, in the above - mentioned technical solutions of the present invention, the above are only the preferred embodiments of the present invention. Therefore, it does not limit the patent scope of the present invention. Any equivalent structural transformation made under the technical concept of the present invention by using the content of the specification and drawings of the present invention, or directly / indirectly applied in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A soil heavy metal pollution ecological risk assessment method based on machine learning, characterized in that: Includes steps: S1, constructing a metal data set; the metal data set includes 0.5, 1, 2, 5 and 10 times the allowable content of the heavy metal to be measured in the environmental quality standard; the heavy metals to be measured include copper, zinc, lead, cadmium, mercury and arsenic; S2, obtaining improved biomarker index; S3, integrating the improved biomarker index in step S2 with the pollution load index, the Nemerow index, the potential ecological risk index and the improved geoaccumulation index into a multi-index data set; S4, using PCA to reduce the dimension of the multi-index data set in step S3, taking the principal component with the largest variance in the data as the dimension reduction result and classifying according to the dimension reduction result to obtain a PCA classification result; S5, integrating the metal data set and the PCA classification result in step S4 into an ecological risk assessment data set; The ecological risk assessment data set is randomly divided into a training set and a test set; wherein the training set accounts for 70% and the test set accounts for 30%; S6, providing a machine learning algorithm, and using the machine learning algorithm to screen the training set and the test set in step S5 to construct a soil heavy metal multi-classification ecological risk assessment model; wherein the machine learning algorithm includes a random forest, a support vector machine, a decision tree, a k-nearest neighbor, and a naive Bayes model; S7, using the soil heavy metal multi-classification ecological risk assessment model in step S6 to conduct heavy metal pollution ecological risk assessment on the target soil.
2. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 1 is characterized in that: In step S2, the improved biomarker index is obtained by: obtaining initial data of earthworm biomarkers and soil metals through literature search to construct an initial biomarker data matrix X; The data in the initial biomarker data matrix are standardized and the correlation coefficient r of the biomarkers is calculated. ij , conflict information A j , biomarker change level AL and change index S j , and obtain the overall information index C based on the calculation results j ; According to the biomarker basic weight W n Calculate the biomarker fine-tuning weight W j ; Here, i refers to a certain day; j refers to a certain biomarker; Establish a set of evaluation indicators U = {U1 U2 U3 U4} and a set of fuzzy comments V = {V1 V2 V3 V4}; and establish a fuzzy comprehensive matrix through the evaluation indicator set and the fuzzy comment set. Among them, U refers to different biomarkers; V refers to the change level of biomarkers and the control group; V1 refers to when |AL|≤20%, and the score is 4 points; V2 refers to when 20%<|AL|≤50%, and the score is 3 points; V3 refers to when 50%<|AL|≤100%, and the score is 2 points; V4 refers to when |AL|>100%, and the score is 4 points; The biomarker index CFBRI is calculated by the biomarker fine-tuned weight vector W and the fuzzy comprehensive matrix; wherein, Providing a standardized metal concentration and fitting it with the biomarker index to generate a fitting curve; standardizing the metal data set in step S1, and substituting the metal data set into the fitting curve to obtain the improved biomarker index.
3. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 2 is characterized in that: In step S2, the objective function of the initial biomarker data matrix X is The standardized initial biomarker data matrix X ij ′ The objective function is 4. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 2 is characterized in that: In step S2, the correlation coefficient r of the biomarker ij , the conflict information A j , the change level AL of the biomarker, the change index S j and the overall information index C j The objective functions are: <h2 style=";text-align:left;direction:ltr">X<h2 style=";text-align:left;direction:ltr"> ij″ <h2 style=";text-align:left;direction:ltr"> =|AL|; C j =A j ·S j (j=1,2,3,…,n)。 5. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 2 is characterized in that: In step S2, the biomarker fine-tuned weight W j The objective function is The objective function of the standardized metal concentration is 6. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 1, characterized in that: In step S4, after the step of taking the principal component with the largest variance in the data as the dimensionality reduction result, it also includes classifying the dimensionality reduction result according to the numerical quartiles to obtain the PCA classification result; wherein the PCA classification results are mild, moderate, severe and serious in order.
7. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 1, characterized in that: Step S5 also includes performing balanced data set processing and cross-validation processing on the results of the training set.
8. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 1, characterized in that: In step S6, the machine learning algorithm is obtained as follows: The target index is used to evaluate the model performance, and the optimal solution is screened to obtain the corresponding machine learning algorithm; wherein the target index includes area under the curve, accuracy, sensitivity, specificity, precision, F1 score, macro F1 score, Kappa coefficient and Hamming loss.
9. The method for ecological risk assessment of soil heavy metal pollution based on machine learning according to claim 1, characterized in that: In step S6, the machine learning algorithm is random forest.
10. An application of the soil heavy metal pollution ecological risk assessment method based on machine learning as described in any one of claims 1 to 9 in the classification planning of soil heavy metal pollution areas.
Citation Information
Cited By
Port ecological risk assessment method and system based on biotoxicity and heavy metal fusion
CN120611981A
Safety assessment method for new reconstruction and expansion land related to underground concrete storage pool
CN121562005A
Soil heavy metal ecological risk microorganism rapid detection method and system
CN121884991A
A rapid detection method and system for soil heavy metal ecological risk microorganisms
CN121884991B
Method for repairing heavy metal contaminated soil through combination of microorganisms and plants
CN122209807A