Pollution assessment method based on XGBoost collaborative kriging

By combining XGBoost with the co-kriging method to generate comprehensive auxiliary variables, the problem of high computational complexity of the traditional co-kriging method in multivariate processing is solved, and efficient and accurate pollution assessment is achieved.

CN119671068BActive Publication Date: 2025-09-23CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510188090.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-09-23
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The traditional co-kriging method has high computational complexity when dealing with multiple auxiliary variables, and it is difficult to provide high-precision and low-cost pollution assessment solutions under the influence of highly nonlinear characteristics and multivariate interactions.

Method used

Combining the XGBoost machine learning algorithm with the cokriging interpolation method, XGBoost is used to extract the most relevant information from multiple auxiliary variables, generate comprehensive auxiliary variables, reduce the complexity of cross-covariance calculation, and use the cokriging method for pollution assessment.

Benefits of technology

It reduces computational complexity, improves data modeling efficiency and accuracy, is suitable for spatial distribution analysis in complex geological fields, and improves the accuracy and efficiency of pollution assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119671068B_ABST
    Figure CN119671068B_ABST
Patent Text Reader

Abstract

The present invention provides a pollution assessment method based on XGBoost co-kriging, taking the spatial distribution prediction of heavy metal mercury pollution in soil as an example, and fully considering the influence of multiple auxiliary variables on mercury content. The present invention first adopts the XGBoost regression model to train the nonlinear relationship between each auxiliary variable and the mercury content, extracts and integrates the information of all auxiliary variables, and generates a comprehensive one-dimensional auxiliary variable. Subsequently, the comprehensive auxiliary variable is used as an auxiliary factor for co-kriging interpolation to optimize the spatial distribution prediction of soil mercury content. Compared with the co-kriging method that directly uses multiple auxiliary variables, the present invention can more effectively reduce the dimension, improve the prediction accuracy and model stability, and at the same time improve the computational efficiency, providing a highly efficient and accurate prediction method for soil pollution assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental data processing, and in particular to a pollution assessment method based on XGBoost and collaborative kriging, which is applicable to soil pollution monitoring and assessment and other environment-related fields. Background Art

[0002] With the deepening of geological research, the understanding and prediction of natural phenomena are becoming increasingly important. The spatial distribution of soil pollution is often affected by multiple environmental factors, such as soil composition, meteorological conditions, terrain characteristics, and human activities, and there are complex interactions between these factors. Traditional cokriging interpolation methods are widely used in spatial data analysis, especially in the fields of geological resource evaluation and groundwater pollution monitoring. Due to its advantages in processing spatially correlated data, it has achieved good application results. However, when multiple auxiliary variables are involved, the computational complexity of the cokriging method increases significantly. It is necessary to calculate the cross-covariance matrix between the main variable and each auxiliary variable separately, which increases the difficulty of data modeling and calculation. At the same time, traditional methods often find it difficult to provide high-precision and low-computational-cost solutions when faced with the highly nonlinear characteristics of the data and the interaction of multiple variables.

[0003] Against this backdrop, this paper proposes a data prediction technology that combines the XGBoost machine learning algorithm with the cokriging interpolation method. XGBoost, with its powerful feature selection and regression modeling capabilities, can extract the information most relevant to a target pollution variable (such as soil mercury content) from multiple auxiliary variables and integrate it into a single comprehensive auxiliary variable, effectively reducing the complexity of calculating cross-covariances between multiple variables. This method avoids the tedious process of independently calculating multiple auxiliary variables in the cokriging method while preserving the influence of multiple sources of information on the primary variable, improving the efficiency and accuracy of data modeling.

[0004] This method, implemented in Python, forms an efficient pollution assessment software tool. It is particularly suitable for spatial distribution analysis in complex geoscience fields, including but not limited to geological exploration, environmental pollution assessment, and meteorological pollution forecasting. Compared to traditional cokriging methods, this technique significantly reduces the computational complexity of cokriging interpolation by using XGBoost to reduce the dimensionality of multivariate features, while improving the accuracy and efficiency of data prediction. It is particularly effective when working with datasets with significant nonlinear characteristics and multivariate coupling. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a pollution assessment method based on XGBoost collaborative kriging.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0007] The pollution assessment method based on XGBoost collaborative kriging includes the following steps:

[0008] S1. Collect soil samples from the target area, record the geographic coordinates of each sampling point, and test the content of pollutants and other elements in each soil sample. Clean and preprocess all data and divide the data set into a training set and a test set.

[0009] S2. Build an XGBoost regression model, use the training set to learn the nonlinear relationship between cofactors and target pollutants, use the test set to verify and optimize the model, integrate the effects of all cofactors, and generate a comprehensive variable to characterize their overall impact on the target mercury element;

[0010] S3, based on the XGBoost comprehensive prediction variables obtained through training in S2 as auxiliary variables, the co-kriging method is used to perform spatial distribution interpolation prediction of mercury pollution elements;

[0011] S4. Solve the co-kriging equation to obtain the predicted spatial distribution value of soil mercury content and generate a spatial distribution map of soil pollution based on the predicted results.

[0012] Furthermore, the XGBoost regression model in S2 is expressed as:

[0013]

[0014] in, is a comprehensive auxiliary variable for prediction, It is A base learner, is the total number of base learners.

[0015] Furthermore, the S3 specifically includes the following steps:

[0016] S31, calculating the variogram of each of the main variable and the auxiliary variable, and calculating the cross-variogram between the main variable and the auxiliary variable to characterize the spatial correlation and co-dependence of the variables;

[0017] S32, constructing a combined covariance matrix based on the main variable covariance matrix, the auxiliary variable covariance matrix, and the cross covariance matrix between the two;

[0018] S33. Combine the spatial information of the main variable and the auxiliary variable in the combined covariance matrix, construct a co-kriging equation, and obtain the weight coefficient by solving the equation.

[0019] Furthermore, the variation function of the main variable in S31 is expressed as:

[0020]

[0021] in, For location The value of the main variable, For distance, is the mathematical expectation, is the variance function of the main variable;

[0022] The variogram of the auxiliary variable is expressed as:

[0023]

[0024] Where, is the variogram of the auxiliary variable, yes variogram of the auxiliary variables at the location;

[0025] The cross-variogram function is expressed as:

[0026]

[0027] Where, is the cross-variogram function.

[0028] Furthermore, the combination matrix in S32 is expressed as:

[0029]

[0030]

[0031]

[0032] in, Is the main variable and The mutual covariance between

[0033] is an auxiliary variable and The mutual covariance between

[0034] Is the main variable and auxiliary variables The mutual covariance between .

[0035] Furthermore, the co-kriging equation in S33 is expressed as:

[0036]

[0037] in, Is the main variable and The covariance between

[0038] is an auxiliary variable and The covariance between

[0039] Is the main variable and auxiliary variables The covariance between

[0040] The main variable at position x and the reference point The mutual covariance between

[0041] The main and auxiliary variables at position x and reference point The mutual covariance between

[0042] 、 are the weight coefficients of the main variable and the auxiliary variable respectively.

[0043] Furthermore, the specific calculation of the predicted value of the spatial distribution of soil mercury content in S4 is as follows:

[0044]

[0045] in, For the location The predicted main variable value of

[0046] The main variable is at point The observed value at

[0047] For auxiliary variables at point The observed value at

[0048] 、 are the weight coefficients of the main variable and the auxiliary variable,

[0049] are the number of observation points for the main variable and the auxiliary variable, respectively.

[0050] The present invention has the following beneficial effects:

[0051] This paper combines XGBoost with cokriging methods to effectively introduce the synergistic influence of multiple auxiliary variables while maintaining the spatial autocorrelation of the main variables. Cokriging can integrate multivariate information, but as the number of variables increases, the computational complexity and data dimensionality increase, posing challenges to modeling efficiency. After the introduction of XGBoost, the complexity of cokriging calculations is reduced by integrating multiple auxiliary variables into a single comprehensive variable, while retaining the spatial autocorrelation of the main variables and the synergistic correlation between multiple variables. This method improves prediction accuracy while reducing computational costs and enhancing the stability and adaptability of the model, providing an efficient and accurate technical means for pollution assessment and geospatial data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of the pollution assessment method based on XGBoost collaborative kriging in the present invention. DETAILED DESCRIPTION

[0053] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0054] The pollution assessment method based on XGBoost co-kriging, such as Figure 1 As shown, the following steps are included:

[0055] S1. Data collection and preprocessing: Collect soil samples from the target area, record the geographic coordinates of each sampling point, and test the content of pollutants and other elements in each soil sample. Clean and preprocess all data, and divide the data set into training and test sets:

[0056] First, soil samples are collected from a specific area. The geographic coordinates of each sampling point are recorded, along with the corresponding concentrations of pollutants (e.g., mercury) and other elements (e.g., pH, organic matter, and various oxides) in the soil samples. Appropriate data cleaning and transformation strategies are then employed to prepare data suitable for model training. The dataset is divided into training and test sets, with 30% of the data used for testing and 70% for training.

[0057] S2. XGBoost Regression Model Construction and Training: Build an XGBoost regression model, using the training set to learn the nonlinear relationship between other elements (cofactors) and the target pollutant element (for example, mercury). Use the test set to validate and optimize the model. Finally, integrate the effects of all cofactors to generate a composite variable that represents their overall impact on the target mercury element:

[0058] First, create an XGBoost regression model and train the model. The model can be expressed as:

[0059]

[0060] in, is the predicted value, It is A base learner, is the total number of base learners. XGBoost minimizes an objective function, which is usually composed of a loss function and a regularization term. The loss function measures the difference between the predicted value and the true value. A commonly used loss function is the mean squared error (MSE):

[0061]

[0062] The objective function of XGBoost can be expressed as:

[0063]

[0064] in, is a regularization term used to control model complexity and is usually defined as:

[0065]

[0066] here, is the number of leaf nodes in the tree, ω j is the weight of leaf node j, and is the regularization parameter. XGBoost iteratively adds base learners via the gradient boosting algorithm. For each round k, the model is updated by minimizing the objective function:

[0067] The gradient formula is:

[0068]

[0069] The Hessian matrix formula is:

[0070]

[0071] XGBoost builds a new tree for the current residual (gradient) and selects features at each node by splitting using the gradient and the Hessian matrix. and split point , so that the objective function is minimized:

[0072]

[0073] in, and They are the total gradient and Hessian matrix of the current node, respectively. and It's the left and right

[0074] The gradient and Hessian matrix of the node. After multiple rounds of iteration, the final prediction model consists of multiple trees:

[0075]

[0076] At each prediction, the input features This includes variables that may affect mercury levels, such as soil pH. The outputs from each tree are summed to form the final prediction. Finally, the model performance is evaluated by calculating the mean squared error (MSE) using the mean squared error formula, which measures the difference between the predicted value and the true value. The mean squared error formula is:

[0077]

[0078] The smaller the MSE value, the better the model performance. The R² score is then calculated using the coefficient of determination formula to measure how well the model fits the data. The coefficient of determination formula is:

[0079]

[0080] The closer the R² value is to 1, the better the model explains the variation in the data. Output the MSE and R² scores to evaluate model performance.

[0081] S3, Co-kriging modeling: Based on the XGBoost comprehensive prediction variables obtained through S2 training as auxiliary variables, the co-kriging method is used to perform spatial distribution interpolation prediction of mercury pollution elements:

[0082] The predicted data from the XGBoost method is then combined with cokriging to create a cokriging object. The primary variable is the primary variable to be predicted (mercury content data), and the auxiliary variable is the composite auxiliary variable predicted using XGBoost. In Python, pykrige does not currently support cokriging directly, but it can be implemented manually or using other libraries (such as scikit-gstat, pykrige extensions, or custom implementations). The specific steps are as follows:

[0083] S31. Variogram calculation: Calculate the variograms of the primary and auxiliary variables, and the cross-variograms between the primary and auxiliary variables to characterize the spatial correlation and co-dependence of the variables:

[0084] Use the Variogram class of the skgstat library to calculate the variogram of each of the primary and auxiliary variables, and the cross-variogram between the primary and auxiliary variables. The variogram is used to describe the degree of variation in spatial data and is usually defined as:

[0085]

[0086] in, For location The value of the main variable, For distance, For expectations, is the variance function of the main variable.

[0087] The variogram of the auxiliary variable is expressed as:

[0088]

[0089] in, is the variogram of the auxiliary variable, yes Variogram of the auxiliary variable at location.

[0090] The cross-variogram function is expressed as:

[0091]

[0092] in, is the cross-variogram function.

[0093] S32. Covariance matrix construction: Construct a combined covariance matrix based on the main variable covariance matrix, the auxiliary variable covariance matrix, and the cross covariance matrix between the two:

[0094]

[0095]

[0096]

[0097] in, Is the main variable and The mutual covariance between is an auxiliary variable and The mutual covariance between Is the main variable and auxiliary variables The mutual covariance between .

[0098] S33. Establishment of Co-Kriging Equation: Combine the spatial information of the main variables and auxiliary variables in the combined covariance matrix to construct the Co-Kriging equation, and obtain the weight coefficient by solving the equation:

[0099]

[0100] in, Is the main variable and The covariance between is an auxiliary variable and The covariance between Is the main variable and auxiliary variables The covariance between The main variable is in position and reference points The mutual covariance between The main and auxiliary variables are in position and reference points The mutual covariance between 、 are the weight coefficients of the main variable and the auxiliary variable respectively.

[0101] S4. Solve the co-kriging equation to obtain the predicted spatial distribution value of soil mercury content and generate a soil pollution spatial distribution map of the predicted results:

[0102] In this embodiment, the predicted spatial distribution value of soil mercury content is obtained by solving the following equation:

[0103]

[0104] Where, For the location The predicted main variable value of The main variable is at point The observed value at For auxiliary variables at point The observed value at 、 are the weight coefficients of the main variable and the auxiliary variable, are the number of observation points for the main variable and the auxiliary variable, respectively.

[0105] After obtaining the predicted values ​​for the primary variable (mercury content), we used Matplotlib to generate a more accurate spatial distribution map of mercury content. This map shows areas of high and low mercury concentrations. Heat maps were used to visually demonstrate the differences in prediction performance between different methods.

[0106] The root mean square error (RMSE) and coefficient of determination (R²) were used to evaluate the model's predictive performance. The results of using cokriging alone and combining it with XGBoost-cokriging were compared. In this case study, XGBoost-cokriging outperformed cokriging alone in terms of reducing RMSE and increasing R².

[0107] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0108] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0110] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0111] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. The pollution assessment method based on XGBoost collaborative kriging is characterized by: The steps include: S1. Collect soil samples from the target area, record the geographic coordinates of each sampling point, and test the content of pollutants and other elements in each soil sample. Clean and preprocess all data and divide the data set into training and test sets. S2. Build an XGBoost regression model, use the training set to learn the nonlinear relationship between cofactors and target pollutants, use the test set to verify and optimize the model, integrate the effects of all cofactors, and generate a comprehensive variable to characterize their overall impact on the target mercury element. The specific method is as follows: S21. Create an XGBoost regression model and train the model. The model can be expressed as: in, is the predicted value, It is A base learner, is the total number of base learners; S22. Create an XGBoost objective function, iteratively add base learners through the gradient boosting algorithm, and update the base learners in each round using the objective function, where the objective function is expressed as: Where, is the i-th output result, is the predicted value of the i-th output result, represents the loss function, It is A base learner, is the total number of base learners; is a regularization term used to control the complexity of the model, expressed as: ; in, is the number of leaf nodes in the tree, ω j Is a leaf node The weight, and is the regularization parameter. XGBoost iteratively adds base learners through the gradient boosting algorithm. For each round of base learners, the model is updated by minimizing the objective function. S23, XGBoost builds a new tree for the current residual, using the gradient and Hessian matrix to split, at each node Select features and split points to minimize the objective function, which can be expressed as: in, and They are the total gradient and Hessian matrix of the current node, respectively. and are the gradient and Hessian matrices of the left and right child nodes; S24. After multiple rounds of iteration, the final prediction model of multiple tree outputs is obtained, and the output of each tree is obtained and summed up to form the final comprehensive variable prediction result, which is expressed as: Where, is the predicted value, It is A base learner, is the total number of base learners; S3, based on the XGBoost comprehensive variable obtained through training in S2 as an auxiliary variable, uses the co-kriging method to perform spatial distribution interpolation prediction of mercury pollution elements, wherein the comprehensive variable is generated by nonlinearly integrating multiple auxiliary factors through XGBoost, and is used to replace the independent calculation of the original multiple auxiliary variables in co-kriging, specifically including the following steps: S31, calculating the variogram of each of the main variable and the auxiliary variable, and calculating the cross-variogram between the main variable and the auxiliary variable to characterize the spatial correlation and co-dependence of the variables; S32, constructing a combined covariance matrix based on the main variable covariance matrix, the auxiliary variable covariance matrix, and the cross covariance matrix between the two; S33, combining the spatial information of the main variable and the auxiliary variable in the combined covariance matrix, constructing a co-kriging equation, and obtaining a weight coefficient by solving the equation; S4. Solve the co-kriging equation to obtain the predicted spatial distribution value of soil mercury content and generate a spatial distribution map of soil pollution based on the predicted results.

2. The pollution assessment method based on XGBoost collaborative kriging according to claim 1, characterized in that: The variation function of the main variable in S31 is expressed as: in, For location The value of the main variable, For distance, is the mathematical expectation, is the variance function of the main variable; The variogram of the auxiliary variable is expressed as: Where, is the variogram of the auxiliary variable, yes variogram of the auxiliary variables at the location; The cross-variogram function is expressed as: Where, is the cross-variogram function.

3. The pollution assessment method based on XGBoost collaborative kriging according to claim 1, characterized in that: The combination matrix in S32 is expressed as: in, Is the main variable and The mutual covariance between is an auxiliary variable and The mutual covariance between Is the main variable and auxiliary variables The mutual covariance between .

4. The pollution assessment method based on XGBoost collaborative kriging according to claim 1, characterized in that: The co-kriging equation in S33 is expressed as: in, Is the main variable and The covariance between is an auxiliary variable and The covariance between Is the main variable and auxiliary variables The covariance between The main variable at position x and the reference point The mutual covariance between The main and auxiliary variables at position x and reference point The mutual covariance between 、 are the weight coefficients of the main variable and the auxiliary variable respectively.

5. The pollution assessment method based on XGBoost collaborative kriging according to claim 1, characterized in that: The specific calculation of the predicted spatial distribution value of soil mercury content in S4 is as follows: in, For the location The predicted main variable value of The main variable is at point The observed value at For auxiliary variables at point The observed value at 、 are the weight coefficients of the main variable and the auxiliary variable, are the number of observation points for the main variable and the auxiliary variable, respectively.

Citation Information

Patent Citations

  • Interpolation method of seismic local velocity data based on collaborative Kriging method

    CN110941014A