Machine Learning-Based Methods, Devices, Equipment, and Media for Predicting Heavy Metal Concentration Distribution in Industrial Sites

By constructing a multi-dimensional covariate system and combining various machine learning algorithms, the accuracy and adaptability issues of heavy metal concentration distribution prediction in industrial sites were solved, achieving high-precision three-dimensional distribution prediction and supporting refined management and remediation of contaminated sites.

CN121884992BActive Publication Date: 2026-05-26CENT SOUTH UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-03-18
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for predicting heavy metal concentration distribution in industrial sites suffer from problems such as wasted data resources, low prediction accuracy, poor spatial continuity, and low visualization, especially when sampling points are sparse, making it difficult to accurately characterize the three-dimensional spatial distribution features.

Method used

By acquiring multi-source data, including soil sampling data, geophysical exploration data, and site-related basic data, a multi-dimensional covariate system is constructed. Then, random forest algorithm, extreme gradient boosting algorithm, and lightweight gradient boosting machine algorithm are used to construct and select the optimal prediction model for the distribution range of heavy metal concentration, and finally, three-dimensional prediction is performed.

Benefits of technology

It achieves high-precision prediction of heavy metal concentration distribution range, adapts to complex scenarios, provides a basis for refined management and remediation of contaminated sites, and improves the accuracy and reliability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884992B_ABST
    Figure CN121884992B_ABST
Patent Text Reader

Abstract

This application discloses a machine learning-based method, apparatus, equipment, and medium for predicting heavy metal concentration distribution in industrial sites, relating to the field of environmental remediation technology. The method involves: acquiring multi-source data from industrial sites, including soil sampling data, geophysical exploration data, and relevant site-specific basic data, to construct a multi-dimensional covariate system. Multiple machine learning algorithms, such as random forest, extreme gradient boosting, and lightweight gradient boosting, are employed to construct and select the optimal heavy metal concentration distribution range prediction model. Finally, a high-precision three-dimensional prediction of the heavy metal concentration distribution range in industrial sites is achieved, yielding the three-dimensional distribution results of heavy metal concentration. By combining multi-source data and multiple machine learning algorithms, the accuracy and reliability of heavy metal concentration distribution range prediction are improved, providing a scientific basis for the refined management and remediation of contaminated sites, and adapting to complex scenarios of compound contaminated sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental remediation technology, and in particular to a method, apparatus, equipment and medium for predicting the concentration distribution of heavy metals in industrial sites based on machine learning. Background Technology

[0002] Currently, predictions of heavy metal concentration distribution in industrial sites mostly employ traditional geostatistical interpolation methods, such as Kriging interpolation and inverse distance weighted interpolation. These methods use measured concentration data from soil sampling points, combined with spatial interpolation algorithms, to achieve planar or simple three-dimensional predictions of the heavy metal concentration distribution range. While some studies have introduced machine learning algorithms for heavy metal concentration distribution range prediction, they often use only a single machine learning model, and the input feature variables are mostly limited to basic data such as soil physicochemical properties and simple spatial distances, without integrating multi-source data such as geophysical surveys. Furthermore, existing prediction methods lack sufficient refinement in data preprocessing, failing to systematically extract features and handle redundancy from multi-source data, and failing to perform differentiated model training and hyperparameter optimization based on the characteristics of different machine learning algorithms. Some three-dimensional prediction methods also suffer from poor adaptability to grid division and soil stratification, as well as spatial heterogeneity of pollution, achieving only simple spatial interpolation extensions.

[0003] Traditional geostatistical interpolation methods heavily rely on the number and distribution of soil sampling points. When sampling points are sparse, the accuracy and spatial continuity of the prediction results are poor, making it difficult to accurately depict the three-dimensional spatial distribution characteristics of heavy metal pollution. The predictive ability of a single machine learning model is limited and easily affected by single feature variables and poor data quality, resulting in weak generalization ability of the prediction results. The integration and feature mining of multi-source data are insufficient, failing to fully utilize the indicative role of geophysical exploration data and spatial information of site pollution sources in heavy metal distribution, resulting in a waste of data resources. The grid design of three-dimensional prediction has poor adaptability to the actual site conditions, failing to accurately reflect the differences in heavy metal concentrations in different soil layers and areas with different pollution heterogeneity, and the visualization of the prediction results is low, making it difficult to intuitively support pollution remediation decisions.

[0004] Therefore, how to integrate multi-source data to construct a refined feature system and combine various machine learning algorithms to achieve high-precision three-dimensional distribution prediction of heavy metal concentration in industrial sites has become an urgent problem to be solved. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, equipment, and medium for predicting the distribution of heavy metal concentrations in industrial sites based on machine learning, aiming to solve the technical problem of how to combine multi-source data and machine learning algorithms to improve the accuracy of predicting the distribution range of heavy metal concentrations in industrial sites.

[0006] To achieve the above objectives, this application proposes a machine learning-based method for predicting heavy metal concentration distribution in industrial sites, comprising:

[0007] Acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data;

[0008] The multi-source data is preprocessed to construct a multi-dimensional covariate system;

[0009] Based on the multi-dimensional covariate system, various machine learning algorithms are used to construct various heavy metal concentration distribution range prediction models, including random forest algorithm, extreme gradient boosting algorithm and lightweight gradient boosting machine algorithm.

[0010] Multiple heavy metal concentration distribution range prediction models were trained and validated, and the optimal prediction model was selected.

[0011] The optimal prediction model is used to make a three-dimensional prediction of the distribution range of heavy metal concentration in industrial sites, and the three-dimensional distribution results of heavy metal concentration in industrial sites are obtained.

[0012] In one embodiment, the step of preprocessing the multi-source data to construct a multi-dimensional covariate system includes:

[0013] Outlier detection was performed on pollutant content data in soil sampling data to obtain standardized basic soil data. The outlier detection was divided into removing outlier data that exceeded the preset confidence level using the Grubbs criterion and supplementing missing data using the K-nearest neighbor interpolation method based on soil type grouping.

[0014] Inversion is performed based on apparent resistivity data from geophysical exploration data to generate a true resistivity profile.

[0015] High-resolution resistivity covariate data were obtained by using ordinary Kriging interpolation based on the actual resistivity profile.

[0016] Based on the pollution source coordinates in the site-related basic data and the three-dimensional sampling coordinates of the soil samples, the three-dimensional shortest straight-line distance from each sampling point to each pollution source is calculated, and a multi-dimensional pollution source distance covariate matrix is ​​generated.

[0017] The pollutant concentration data in the standardized soil baseline data are processed to obtain spatial regionalization index covariate data;

[0018] The spatial clustering characteristic covariate data are obtained by calculating and processing the three-dimensional sampling coordinates of the soil samples and the concentration data of each pollutant in the standardized soil basic data.

[0019] By performing feature correlation analysis, redundancy processing and integration are performed on the soil physicochemical property covariates, the high-resolution resistivity covariate data, the multidimensional pollution source distance covariate matrix, the spatial regionalization index covariate data, and the spatial clustering characteristic covariate data in the standardized soil basic data to construct a multidimensional covariate system. The multidimensional covariate system includes soil physicochemical properties, geophysical attributes, spatial distance, and spatial distribution characteristics.

[0020] In one embodiment, the step of processing the pollutant concentration data in the standardized soil baseline data to obtain spatial regionalization index covariate data includes:

[0021] Extract all pollutant concentration data and the three-dimensional coordinates of the corresponding sampling points from the standardized soil baseline data, and construct concentration-coordinate datasets according to pollutant type;

[0022] For each pollutant's concentration-coordinate dataset, a spatial K-fold cross-validation strategy is used to partition the data, dividing the entire dataset into K spatially mutually exclusive subsets;

[0023] For each target sampling point of a feature to be calculated, the spatial subset to which the target sampling point belongs is set as the validation set, and the remaining K-1 spatial subsets are set as the training set.

[0024] The concentration information of the target sampling point itself is removed, and a local variation function model is obtained by fitting the training set with regional variable theory. The model parameters of the local variation function model are determined, including range, sill value and nugget value.

[0025] Based on the local variation function model and the model parameters, the three-dimensional ordinary kriging method is used to interpolate the local spatial region where the target sampling point is located, generating a subset of the continuous distribution field of pollutant concentration in the local region;

[0026] The subset of the continuous distribution field of pollutant concentration in the local area is imported into a geographic information processing tool. The natural fault classification method is used to perform preliminary regional division of the subset of the continuous distribution field of pollutant concentration, resulting in multiple local division schemes with different numbers of sub-regions.

[0027] Calculate the variance optimal fit value and inflection point feature value for each group of local partitioning schemes, select the local partitioning scheme with the largest variance optimal fit value and the inflection point feature value that meets the preset conditions, and determine the corresponding optimal number of local sub-regions.

[0028] For the division results corresponding to the optimal number of local sub-regions, the arithmetic mean of pollutant concentrations at all interpolation points in each local sub-region is calculated. The arithmetic mean of pollutant concentrations is used as the local spatial regionalization index of the corresponding local sub-region, and the local spatial regionalization index is assigned to the target sampling point in the corresponding local sub-region to obtain the pollutant spatial regionalization index feature value corresponding to the sampling point.

[0029] Point-by-point feature calculations were performed on all sampling points to obtain spatial regionalization index covariate sub-data for each pollutant.

[0030] By integrating the spatial regionalization index covariate sub-data of all pollutants, spatial regionalization index covariate data is obtained.

[0031] In one embodiment, the step of calculating and processing the three-dimensional sampling coordinates of the soil sample and the concentration data of each pollutant in the standardized soil baseline data to obtain spatial clustering characteristic covariate data includes:

[0032] The three-dimensional sampling coordinates of the soil samples were extracted, and the X, Y, and Z axis data of the three-dimensional sampling coordinates were normalized by the min-max normalization method to obtain the standardized three-dimensional sampling coordinates.

[0033] Based on the standardized three-dimensional sampling coordinates, a distance bandwidth type spatial weight matrix is ​​constructed;

[0034] The spatial statistical analysis tool is used to import the concentration data of each pollutant in the standardized soil basic data, the standardized three-dimensional sampling coordinates, and the distance bandwidth spatial weight matrix. The local statistics corresponding to each pollutant concentration are calculated to obtain the Z value corresponding to each sampling point.

[0035] Set a preset positive threshold and a preset negative threshold, and determine the sampling points with Z values ​​greater than the preset positive threshold as pollution hotspots, and determine the sampling points with Z values ​​less than the preset negative threshold as pollution colds;

[0036] Based on the standardized three-dimensional sampling coordinates, the pollution hotspots, and the pollution coldspots, the first Euclidean distance from each sampling point to the nearest pollution hotspot and the second Euclidean distance from each sampling point to the nearest pollution coldspot are calculated respectively.

[0037] Integrating the first Euclidean distance and the second Euclidean distance yields spatial clustering feature covariate data.

[0038] In one embodiment, the step of training and validating multiple prediction models for the distribution range of heavy metal concentrations and selecting the optimal prediction model includes:

[0039] Key feature variables were extracted from the multidimensional covariate system, and a model training dataset was constructed by combining measured soil pollutant concentration data.

[0040] The training dataset of the model is divided into training set and validation set by adopting a spatial block cross-validation strategy. The spatial block cross-validation strategy is to divide the industrial site into three-dimensional spatial blocks, and the sampling points in each spatial block are allocated as a whole to the training set or validation set.

[0041] The initial random forest prediction model, the initial extreme gradient boosting prediction model, and the initial extreme gradient boosting prediction model are trained and optimized using the training set to obtain the optimized random forest prediction model, the optimized extreme gradient boosting prediction model, and the optimized lightweight gradient boosting machine prediction model.

[0042] The optimization random forest prediction model, the optimization extreme gradient boosting prediction model, and the optimization lightweight gradient boosting machine prediction model are calculated using the validation set to obtain the coefficient of determination and root mean square error.

[0043] The model with the best overall performance is selected as the optimal prediction model based on the coefficient of determination and root mean square error.

[0044] In one embodiment, the step of training and optimizing the initial random forest prediction model, the initial extreme gradient boosting prediction model, and the initial extreme gradient boosting prediction model using the training set to obtain the optimized random forest prediction model, the optimized extreme gradient boosting prediction model, and the optimized lightweight gradient boosting prediction model includes:

[0045] The training set is input into the initial random forest prediction model. Recursive feature elimination combined with 5-fold cross-validation is used to select the feature subset that contributes the most to the prediction of pollutant concentration. The hyperparameters of decision tree number, maximum depth and minimum number of split samples are optimized simultaneously to obtain the optimized random forest prediction model.

[0046] The training set is input into the initial extreme gradient boosting prediction model. The error change is monitored by the early stopping mechanism, and the learning rate, subsampling rate and L1 / L2 regularization coefficient hyperparameters are optimized to obtain the optimized extreme gradient boosting prediction model.

[0047] The training set is input into the initial lightweight gradient booster prediction model. Based on histogram optimization and leaf growth strategy, the hyperparameters of the upper limit of the number of leaves, the minimum data segmentation threshold, and the feature score calculation method are adjusted to obtain the optimized lightweight gradient booster prediction model.

[0048] In one embodiment, the step of using the optimal prediction model to perform three-dimensional prediction of the heavy metal concentration distribution range of an industrial site, and obtaining the three-dimensional distribution result of heavy metal concentration in the industrial site, includes:

[0049] Based on the geographical boundaries, stratum distribution depth, and three-dimensional coordinates of sampling points of the industrial site, a three-dimensional grid of the industrial site is generated. The node spacing in the X and Y directions of the three-dimensional grid is set according to the spatial heterogeneity of site pollution, and the node spacing in the Z direction matches the soil layer thickness.

[0050] The multidimensional covariate system is mapped to each node of the three-dimensional grid using three-dimensional kriging interpolation to form a grid covariate dataset.

[0051] Input the grid covariate dataset into the optimal prediction model and output the predicted pollutant concentration value for each grid node;

[0052] Based on the training and verification error distribution pattern of the optimal prediction model, the confidence interval and confidence level corresponding to the predicted pollutant concentration value of each grid node are calculated.

[0053] Set preset screening thresholds for each pollutant based on preset screening criteria;

[0054] The predicted pollutant concentration of each grid node is combined with the confidence interval for comprehensive judgment, and the set of grid nodes that exceed the preset screening threshold and whose lower edge of the confidence interval is not lower than the preset screening threshold is marked.

[0055] By combining the site stratigraphic structure data, the set of grid nodes exceeding the standard is spatially topologically reconstructed to generate a three-dimensional pollution dataset, wherein the three-dimensional pollution dataset includes the spatial coordinates of grid nodes, soil layer attribution, the level of exceeding the standard concentration, and the corresponding confidence level.

[0056] Interactive 3D volume rendering technology is used to perform correlation rendering and visualization processing on the 3D pollution dataset by combining confidence level and concentration level, so as to obtain the 3D distribution results of heavy metal concentration in industrial sites.

[0057] Furthermore, to achieve the above objectives, this application also proposes a machine learning-based heavy metal concentration distribution prediction device for industrial sites, the machine learning-based heavy metal concentration distribution prediction device for industrial sites comprising:

[0058] The acquisition module is used to acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data.

[0059] The processing module is used to preprocess the multi-source data and construct a multi-dimensional covariate system;

[0060] The model building module is used to construct various heavy metal concentration distribution range prediction models based on the multi-dimensional covariate system and employing various machine learning algorithms, including random forest algorithm, extreme gradient boosting algorithm and lightweight gradient boosting machine algorithm.

[0061] The model selection module is used to train and validate various prediction models for the concentration distribution range of heavy metals, and select the optimal prediction model.

[0062] The results module is used to perform three-dimensional prediction of the heavy metal concentration distribution range of the industrial site using the optimal prediction model, and obtain the three-dimensional distribution results of heavy metal concentration in the industrial site.

[0063] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the machine learning-based method for predicting the concentration distribution of heavy metals in industrial sites as described above.

[0064] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites as described above.

[0065] This application constructs a multi-dimensional covariate system by acquiring multi-source data from industrial sites, including soil sampling data, geophysical exploration data, and site-related basic data. Multiple machine learning algorithms, such as random forest, extreme gradient boosting, and lightweight gradient boosting, are employed to construct and select the optimal prediction model for heavy metal concentration distribution range. Ultimately, a high-precision three-dimensional prediction of the heavy metal concentration distribution range of industrial sites is achieved, yielding the three-dimensional distribution results. By combining multi-source data and multiple machine learning algorithms, the accuracy and reliability of heavy metal concentration distribution range prediction are improved, providing a scientific basis for the refined management and remediation of contaminated sites and adapting to complex scenarios of compound contaminated sites. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 This is a flowchart illustrating the first embodiment of the machine learning-based method for predicting heavy metal concentration distribution in industrial sites according to this application.

[0068] Figure 2 This is a flowchart illustrating the second embodiment of the machine learning-based method for predicting heavy metal concentration distribution in industrial sites according to this application.

[0069] Figure 3 This is a schematic diagram of the module structure of the heavy metal concentration distribution prediction device for industrial sites based on machine learning, according to the first embodiment of the machine learning-based method for predicting heavy metal concentration distribution in industrial sites of this application.

[0070] Figure 4 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the machine learning-based method for predicting the concentration distribution of heavy metals in industrial sites in this application embodiment.

[0071] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0072] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0073] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0074] This application proposes a machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites. This method combines multi-source data with advanced machine learning algorithms to achieve high-precision and high-reliability prediction of the distribution range of heavy metal concentrations in industrial sites, and effectively analyzes pollution sources. The main solution of this application is as follows: First, acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data. Second, preprocess the multi-source data to construct a multi-dimensional covariate system. Third, based on the multi-dimensional covariate system, construct multiple heavy metal concentration distribution range prediction models using various machine learning algorithms, including random forest, extreme gradient boosting, and lightweight gradient boosting algorithms. Fourth, train and validate the multiple heavy metal concentration distribution range prediction models, and select the optimal prediction model. Fifth, use the optimal prediction model to perform a three-dimensional prediction of the heavy metal concentration distribution range of the industrial site, obtaining the three-dimensional distribution result of heavy metal concentrations in the industrial site.

[0075] Based on the above, this application also provides a machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the machine learning-based method for predicting heavy metal concentration distribution in industrial sites according to this application.

[0076] In this embodiment, the machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites includes steps S10 to S50:

[0077] Step S10: Obtain multi-source data of the industrial site.

[0078] It should be noted that multi-source data includes soil sampling data, geophysical exploration data, and site-related basic data. Soil sampling data refers to the data obtained through laboratory testing after soil samples are collected at sampling points within the industrial site in accordance with soil environmental monitoring technical specifications. This data covers pH value, turbidity, heavy metal content (including cadmium, lead, copper, nickel, zinc, total chromium, hexavalent chromium, arsenic, and mercury), 22 volatile organic compounds (VOCs), 3 semi-volatile organic compounds (SOCs), and petroleum hydrocarbon content, and is core data directly reflecting the degree of soil pollution and the types of pollutants. Geophysical exploration data refers to the apparent resistivity data collected through an electrical resistivity system after setting up survey lines within the industrial site using the high-density resistivity method, a non-invasive geophysical exploration technique. This data reflects the conductivity characteristics of the underground medium, and heavy metal pollution can lead to increased pore water ion concentration in the soil, thereby altering conductivity; therefore, it has a potential correlation with soil pollution status. Site-related basic data refers to basic information closely related to the industrial site obtained through on-site investigation and data collection. This includes the site's geographical coordinates, stratum distribution information (i.e., the type of soil at different depths, such as sandy soil at 0 to -4 meters, loam at -4 to -9 meters, and clay at -9 to -12 meters), historical production activity information (i.e., the production business and operating period that the site has carried out, such as an electroplating plant from 1998 to 2018 and a waste recycling station from 2018 to 2019), and pollution source location information (i.e., the specific locations of wastewater treatment areas, hazardous waste areas, and waste storage points in each workshop).

[0079] Specifically, following the technical specifications for soil environmental monitoring, 64 soil sampling points, including 9 deepened sampling points, were first established within the industrial site. Five to ten soil samples were collected from each sampling point, and these samples were sent to a laboratory for testing to obtain soil sampling data such as pH value, turbidity, and the content of various pollutants. Next, six survey lines ranging from 130 to 190 meters in length were laid out within the industrial site. An electrical resistivity system using the Wenner-Schlumberger apparatus was employed to collect apparent resistivity data, yielding geophysical exploration data. The Wenner-Schlumberger apparatus is a commonly used electrode arrangement device in high-density resistivity methods, consisting of four electrodes arranged in a specific order. The two outer electrodes are power supply electrodes, used to supply a stable current to the ground, while the two inner electrodes are measuring electrodes, used to detect the potential difference in the underground medium caused by the current. Finally, through on-site investigation and data collection, basic site data, including geographical coordinates, stratigraphic distribution, historical production activities, and the location of pollution sources, were obtained.

[0080] Step S20: Preprocess the multi-source data to construct a multi-dimensional covariate system.

[0081] It should be noted that the multi-dimensional covariate system refers to the integration of various types of preprocessed data to form a comprehensive dataset covering multiple influencing factors. It includes four types of covariates: soil physicochemical properties, spatial distribution of resistivity, distance from pollution sources, and spatially derived variables. This system can comprehensively characterize various key factors affecting soil pollution distribution. Covariates are variables that have a potential correlation with soil pollutant concentrations and can provide auxiliary explanations for predicting the distribution range of heavy metal concentrations. These variables can be the soil's own physicochemical properties, the physical characteristics of underground media, or spatial relationships with pollution sources.

[0082] Specifically, firstly, outlier identification and removal were performed on the soil sampling data, and missing values ​​were handled using the mean imputation method to obtain standardized soil sampling data. Next, geophysical exploration data underwent two-dimensional inversion using RES2DINV software, followed by Gaussian-Newton inversion using the pyGIMLi platform to generate true resistivity profile data. Then, three-dimensional Kriging interpolation was used to construct resistivity spatial distribution data. Following this, based on relevant site-specific data, Python was used to calculate the Euclidean distances from each sampling point to each workshop, wastewater treatment area, hazardous waste area, and waste recycling station, obtaining pollution source distance data. Finally, spatial regionalization indices were calculated using three-dimensional Kriging interpolation and the natural fault classification method, combined with local Getis-Ord Gibbs analysis. The statistical measures construct spatially derived data such as distance from pollution hotspots, distance from pollution coldspots, and hotspot Z-values; finally, standardized soil sampling data, resistivity spatial distribution data, pollution source distance data, and spatially derived data are integrated to construct a multi-dimensional covariate system.

[0083] Step S30: Based on a multi-dimensional covariate system, various machine learning algorithms are used to construct multiple prediction models for the distribution range of heavy metal concentrations.

[0084] It should be noted that the machine learning algorithms include Random Forest, Extreme Gradient Boosting, and Lightweight Gradient Boosting Machine. Random Forest is a machine learning algorithm based on ensemble learning. It generates multiple training subsets through Bootstrap sampling, randomly selects feature subsets when splitting the decision tree nodes of each subset, constructs multiple independent decision trees, and finally outputs the prediction result by averaging the predictions of all decision trees. It has strong anti-overfitting and generalization abilities. Extreme Gradient Boosting is an ensemble learning algorithm based on gradient boosting decision trees. It sequentially constructs multiple weak learners (decision trees), with each tree learning the prediction residuals of the previous tree. A regularization term is introduced into the objective function to control model complexity, achieving both high prediction accuracy and efficient computational performance. Lightweight Gradient Boosting Machine is an efficient algorithm based on the gradient boosting framework. It employs a histogram-based decision tree algorithm and a depth-limited leaf growth strategy, suitable for modeling large-scale, high-dimensional datasets. The heavy metal concentration distribution range prediction models are all regression models. Their core purpose is to predict continuous values ​​of heavy metal pollutant concentrations in industrial sites, rather than classifying pollution levels.

[0085] Specifically, firstly, characteristic variables such as soil physicochemical properties, resistivity spatial distribution, pollution source distance, and spatially derived variables are extracted from the multi-dimensional covariate system. These characteristic variables are then correlated with the corresponding measured soil pollutant concentration data to form the basic dataset for model training. Utilizing the core physical law that heavy metal pollution alters soil pore water ion concentration, thereby affecting the soil's electrical conductivity (manifested as resistivity changes), the quantitative correlation trend between resistivity data and pollutant concentrations and soil physicochemical properties (such as pH and organic matter content) is clarified. Secondly, a statistical coupling relationship between the two is constructed. Through feature correlation analysis and regression analysis, the correlation strength between the resistivity covariate and each pollutant concentration and soil physicochemical property covariate is quantified, forming correlation constraint rules. Then, during the training of various initial machine learning models (random forest, extreme gradient boosting, and lightweight gradient boosting machine), the above physical mechanism relationship and statistical coupling relationship are integrated as constraints into the training process: in the model feature input stage, soil chemical characteristic variables with high correlation to the resistivity covariate and conforming to the physical mechanism are preferentially retained, while redundant features that are unrelated or violate the physical laws of pollution migration are eliminated. Next, based on the Random Forest algorithm, initial values ​​for hyperparameters such as the number of decision trees and maximum depth were set, and multiple training subsets were generated through Bootstrap sampling to construct an initial Random Forest prediction model. Then, based on the Extreme Gradient Boosting algorithm, a regularization term was introduced, and initial values ​​for hyperparameters such as learning rate and subsampling rate were set to construct an initial Extreme Gradient Boosting prediction model. Then, based on the Lightweight Gradient Boosting Machine algorithm, a histogram decision tree and leaf growth strategy were adopted, and initial values ​​for hyperparameters such as the upper limit of the number of leaves and the minimum data splitting threshold were set to construct an initial Lightweight Gradient Boosting Machine prediction model. Finally, the training dataset was adapted to be input into the three initial prediction models respectively, completing the construction of prediction models for various heavy metal concentration distribution ranges.

[0086] Step S40: Train and validate various heavy metal concentration distribution range prediction models, and select the optimal prediction model.

[0087] It should be noted that the optimal prediction model refers to the model that performs best across all evaluation metrics among the heavy metal concentration distribution range prediction models trained using three algorithms: Random Forest, Extreme Gradient Boosting, and Lightweight Gradient Boosting Machine. This model can most accurately capture the correlation between covariates and pollutant concentrations, and its prediction results best match the actual pollution situation, thus meeting the needs of predicting the heavy metal concentration distribution range in industrial sites.

[0088] Further, step S40 includes: First, extracting key feature variables from the multi-dimensional covariate system and constructing a model training dataset by combining it with measured soil pollutant concentration data; then, using a spatial block cross-validation strategy, dividing the model training dataset to obtain a training set and a validation set. It should be noted that key feature variables refer to covariates selected from the multi-dimensional covariate system that are strongly correlated with soil pollutant concentrations and have a significant impact on the prediction results. Measured soil pollutant concentration data refer to the actual pollutant concentration values ​​obtained through laboratory testing of soil samples, including the specific content of pollutants such as cadmium, lead, copper, nickel, zinc, total chromium, hexavalent chromium, arsenic, mercury, and petroleum hydrocarbons, which are the core reference standard for model training. The model training dataset refers to the dataset formed by associating and integrating the key feature variables with the corresponding measured soil pollutant concentration data. The spatial block cross-validation strategy involves dividing the industrial site into three-dimensional spatial blocks, with sampling points within each block being allocated as a whole to the training set or validation set. Specifically, by recursive feature elimination or correlation threshold screening, key feature variables that contribute most to the prediction of pollutant concentration are extracted from the multi-dimensional covariate system. These are then matched one-to-one with the synchronously acquired measured soil pollutant concentration data to construct a model training dataset, ensuring that the input features correspond spatially to the target variables and that the information is complete. Secondly, a spatial block cross-validation strategy is used to divide the dataset: the industrial site is first uniformly divided into non-overlapping three-dimensional spatial blocks along the X, Y, and Z directions according to preset side lengths. Then, all sampling points within each spatial block are randomly assigned to the training or validation set as a whole, ensuring that data from the same spatial block do not appear simultaneously in training and validation, thereby avoiding overfitting caused by spatial autocorrelation and improving the model's generalization performance.

[0089] Next, the initial random forest prediction model, the initial extreme gradient boosting prediction model, and the initial extreme gradient boosting prediction model were trained and optimized using the training set to obtain the optimized random forest prediction model, the optimized extreme gradient boosting prediction model, and the optimized lightweight gradient boosting prediction model. Furthermore, the training set was input into the initial random forest prediction model, and recursive feature elimination combined with 5-fold cross-validation was used to select the feature subset that contributed the most to pollutant concentration prediction. Simultaneously, the hyperparameters for the number of decision trees, maximum depth, and minimum number of split samples were optimized to obtain the optimized random forest prediction model. The training set was input into the initial extreme gradient boosting prediction model, and an early stopping mechanism was used to monitor error changes. The hyperparameters for the learning rate, subsampling rate, and L1 / L2 regularization coefficient were optimized to obtain the optimized extreme gradient boosting prediction model. The training set was input into the initial lightweight gradient boosting prediction model, and the hyperparameters for the upper limit of the number of leaves, the minimum data splitting threshold, and the feature score calculation method were adjusted based on histogram optimization and leaf growth strategies to obtain the optimized lightweight gradient boosting prediction model. It's important to note that recursive feature elimination is a feature selection method. It recursively removes features and builds a model, selecting the subset of features that contribute most to the prediction results based on model performance scores. Redundant or low-contribution features are gradually eliminated, improving model training efficiency and prediction accuracy. Five-fold cross-validation divides the training set into five subsets with similar data sizes. The model is trained using four subsets and validated using one subset, repeated five times, and the average performance is used as the model evaluation result. This effectively avoids the randomness of single validations and improves the reliability of hyperparameter optimization. Early stopping is a regularization strategy during model training. It monitors the validation set error in real time and stops training when the error no longer decreases after several consecutive rounds, preventing the model from overfitting the training set. Histogram optimization is a core optimization strategy in the lightweight gradient boosting machine algorithm. It discretizes continuous feature values ​​into multiple histogram intervals and calculates the optimal split point based on these intervals, significantly reducing the computational complexity of high-dimensional covariates and improving model training efficiency. The leaf growth strategy is the decision tree construction strategy in the Lightweight Gradient Boosting Machine algorithm. It restricts the decision tree to grow step by step from leaf nodes and sets a depth limit to avoid excessive complexity and balance model accuracy and computational efficiency. Hyperparameters are parameters in machine learning models that cannot be automatically learned through training and need to be preset or optimized manually. Different algorithms have different core hyperparameters, which directly affect the model's structure and performance and are the core objects of model optimization.Specifically, the training set is input into the initial random forest prediction model, and a recursive feature elimination process is initiated. Using 5-fold cross-validation, the feature subset that contributes the most to pollutant concentration prediction is selected round by round. Simultaneously, the hyperparameters for the number of decision trees, maximum depth, and minimum number of split samples are adjusted. After optimization, an optimized random forest prediction model is obtained. Next, the training set is input into the initial extreme gradient boosting prediction model, and an early stopping mechanism is enabled to monitor the validation set error changes in real time. Based on this, the learning rate, subsampling rate, and L1 / L2 regularization coefficient hyperparameters are adjusted. After optimization, an optimized extreme gradient boosting prediction model is obtained. Finally, the training set is input into the initial lightweight gradient boosting prediction model. High-dimensional covariates are processed based on histogram optimization, and the hyperparameters for the upper limit of the number of leaves, minimum data splitting threshold, and feature score calculation method are adjusted using a leaf growth strategy. After optimization, an optimized lightweight gradient boosting prediction model is obtained.

[0090] Then, the validation set is used to calculate the coefficients of determination (R²) and root mean square error (RMSE) for the optimized random forest, extreme gradient boosting, and lightweight gradient boosting prediction models. Specifically, the performance of the optimized random forest, extreme gradient boosting, and lightweight gradient boosting prediction models is evaluated using the validation set. The prediction accuracy and error level of each model are quantified by calculating the coefficients of determination (R²) and root mean square error (RMSE). The coefficients of determination reflect the goodness of fit of the model to the data, while the RMSE measures the deviation between the model's predictions and the actual values.

[0091] Finally, the model with the best overall performance was selected as the optimal prediction model based on the coefficient of determination and root mean square error. Specifically, the model with the best overall performance was selected as the optimal prediction model based on the comprehensive evaluation results of the coefficient of determination and root mean square error. The model with the highest coefficient of determination and the lowest root mean square error was chosen to ensure that it has good generalization ability and prediction accuracy on untrained data, thus providing reliable model support for the three-dimensional prediction of the distribution range of heavy metal concentrations in industrial sites.

[0092] Step S50: Use the optimal prediction model to make a three-dimensional prediction of the heavy metal concentration distribution range of the industrial site, and obtain the three-dimensional distribution result of heavy metal concentration in the industrial site.

[0093] It should be noted that 3D prediction refers to the process of predicting pollutant concentrations at different locations and in different soil layers of an industrial site from three spatial dimensions: longitude, latitude, and depth. The 3D distribution results of heavy metal concentrations at an industrial site refer to a dataset of heavy metal concentration distribution data presented in a 3D visualization format, including predicted pollutant concentration values ​​for each 3D grid node, spatial coordinates of areas exceeding standards, and pollution boundaries for different soil layers. The 3D grid refers to a regular spatial grid generated based on the geographical boundaries of the industrial site, the depth of the geological strata, and the 3D coordinates of the sampling points. This grid divides the underground space of the entire industrial site into several cubic units of equal or unequal volume, each unit corresponding to a 3D coordinate node. Grid covariate data refers to the data formed by mapping the feature variables in a multi-dimensional covariate system to each node of the 3D grid through interpolation methods, ensuring that each grid node possesses complete covariate characteristics.

[0094] Further, step S50 includes: generating a three-dimensional grid of the industrial site based on its geographical boundaries, stratigraphic depth, and the three-dimensional coordinates of sampling points; mapping the multi-dimensional covariate system to each node of the three-dimensional grid using three-dimensional kriging interpolation to form a grid covariate dataset; inputting the grid covariate dataset into the optimal prediction model to output the predicted pollutant concentration value for each grid node; calculating the confidence interval and confidence level corresponding to the predicted pollutant concentration value for each grid node based on the training and verification error distribution law of the optimal prediction model; setting preset screening thresholds for each pollutant according to preset screening criteria; and processing each grid node... The predicted pollutant concentration values ​​at each point are combined with confidence intervals for comprehensive judgment, marking the set of grid nodes that exceed the preset screening threshold and whose lower edge of the confidence interval is not lower than the preset screening threshold. The set of grid nodes that exceed the standard is then spatially topologically reconstructed by combining site stratigraphic data to generate a three-dimensional pollution dataset. The three-dimensional pollution dataset includes the spatial coordinates of the grid nodes, soil layer attribution, exceedance concentration level, and corresponding confidence level. Interactive three-dimensional volume rendering technology is used to perform correlation rendering and visualization processing on the three-dimensional pollution dataset, combining confidence level and concentration level, to obtain the three-dimensional distribution results of heavy metal concentration in the industrial site.

[0095] It should be noted that spatial topology reconstruction refers to the process of reorganizing and reconstructing the topological attributes of the set of out-of-standard grid nodes, such as their spatial location, adjacency relationships, and soil layer affiliation, based on site stratigraphic data. This process transforms discrete out-of-standard nodes into continuous three-dimensional pollution area boundaries. A three-dimensional pollution dataset is a structured data set that integrates the spatial coordinates of out-of-standard grid nodes, their associated soil layer information, and the level of out-of-standard concentration. It is a core data carrier capable of comprehensively characterizing the three-dimensional spatial distribution of pollution. Interactive 3D volume rendering technology is a technique that transforms the numerical information in a 3D dataset into visualized 3D graphics. It allows users to view the pollution distribution from different perspectives through interactive operations such as rotation, sectioning, and zooming, intuitively presenting the three-dimensional spatial characteristics of pollution.

[0096] Specifically, firstly, by combining the clearly defined geographical boundaries of the industrial site, measured stratum depth data, and the three-dimensional sampling coordinates of all soil samples, a three-dimensional grid covering the entire industrial site is constructed. The node spacing in the X and Y directions of the 3D grid is set according to the spatial heterogeneity of site pollution; smaller spacing is used in areas with high pollution spatial variability to improve prediction accuracy, while the spacing is appropriately increased in areas with low variability to reduce computational load. The node spacing in the Z direction matches the soil stratification thickness. Then, using three-dimensional kriging interpolation, the constructed multi-dimensional covariate system is successively substituted into each node of the 3D grid, realizing the spatial transformation of covariate data from discrete sampling points to continuous grid nodes. Inter-grid mapping is performed to form a complete grid covariate dataset. This is done to ensure that the optimal prediction model can obtain the complete feature input of each grid node, guaranteeing the continuity and comprehensiveness of the prediction. Next, the entire grid covariate dataset is input into the selected optimal prediction model. The model calculates and outputs the predicted pollutant concentration value for each grid node. Simultaneously, based on the error distribution patterns generated during the training and validation of the optimal prediction model, statistical analysis methods are used to calculate the confidence interval and confidence level corresponding to the predicted pollutant concentration value for each grid node. This quantifies the uncertainty of the prediction results and improves their reliability. Finally, based on site pollution risk assessment standards and relevant environmental regulations... To ensure compliance with pre-defined screening standards, preset screening thresholds were set for each pollutant. Then, the predicted pollutant concentration of each grid node was comprehensively judged based on its corresponding confidence interval. Only grid nodes whose predicted pollutant concentration exceeded the preset screening threshold, and whose lower edge of the confidence interval was not lower than the preset screening threshold, were marked, forming a set of grid nodes exceeding the standard. This was done to eliminate misjudgments caused by prediction uncertainty and ensure the authenticity and reliability of the marked exceeding areas. Subsequently, combined with detailed site stratigraphic data, the marked set of grid nodes was spatially topologically reconstructed to complete the soil layer attribution information of the grid nodes, classify the exceeding concentration levels, and integrate the spatial coordinates of the grid nodes. Information such as soil layer attribution, exceedance concentration levels, and corresponding confidence levels are used to generate a complete three-dimensional pollution dataset, enabling precise correlation between pollution data and site stratigraphic characteristics. Finally, interactive three-dimensional volume rendering technology is employed to visualize the three-dimensional pollution dataset, linking confidence levels with concentration levels in the rendering. Different exceedance concentration levels are distinguished by different colors, and within the same concentration level, the color intensity is adjusted according to the confidence level. This generates an interactive, scalable, and layered three-dimensional distribution result of heavy metal concentrations in industrial sites, allowing staff to intuitively grasp the three-dimensional spatial distribution range, concentration levels, and prediction reliability of pollution, providing accurate data support for pollution remediation decisions.

[0097] This embodiment acquires multi-source data from industrial sites, including soil sampling data, geophysical exploration data, and relevant site-specific basic data, and constructs a multi-dimensional covariate system. Multiple machine learning algorithms, such as random forest, extreme gradient boosting, and lightweight gradient boosting, are employed to construct and select the optimal prediction model for heavy metal concentration distribution range. Ultimately, a high-precision three-dimensional prediction of the heavy metal concentration distribution range of industrial sites is achieved, yielding the three-dimensional distribution results. By combining multi-source data and various machine learning algorithms, the accuracy and reliability of heavy metal concentration distribution range prediction are improved, providing a scientific basis for the refined management and remediation of contaminated sites and adapting to complex scenarios of compound contaminated sites.

[0098] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 The method for predicting the concentration distribution of heavy metals in industrial sites based on machine learning, step S20, further includes steps S201 to S207:

[0099] Step S201: Perform outlier detection on the pollutant content data in the soil sampling data to obtain standardized soil basic data.

[0100] It should be noted that outlier detection is divided into two parts: using the Grubbs criterion to remove outliers exceeding a pre-set confidence level, and using K-nearest neighbor interpolation based on soil type grouping to fill in missing data. The Grubbs criterion is an outlier detection method based on a normal distribution. It calculates the Grubbs statistic for each data point in the dataset and compares it with the critical value corresponding to the pre-set confidence level to identify and remove outliers that deviate from the overall data distribution. The confidence level refers to the probability standard for judging whether data is an outlier; that is, it is considered that normal data should fall within the statistical interval corresponding to that probability. Commonly used confidence levels include 95% and 99%, etc. The higher the pre-set confidence level, the more stringent the outlier judgment. K-nearest neighbor interpolation based on soil type grouping first groups the soil sampling data according to soil type, such as sandy soil, loam, and clay. Then, within each group, the K samples closest to the missing data sampling point are selected, and the supplementary value for the missing value is calculated by weighted average. This method takes into account the homogeneity of soil types and improves the accuracy of missing value filling. Standardized soil baseline data refers to soil sampling data that, after outlier removal and missing value supplementation, exhibits a statistically consistent data distribution and has no missing data.

[0101] Specifically, pollutant content data is extracted from soil sampling data. The Grubbs statistic for each data point is calculated based on the Grubbs criterion, and outliers exceeding a pre-set confidence level threshold are removed. Next, the datasets after outlier removal are grouped by soil type, and missing data sampling points within each group are identified. Then, within each soil type group, the K closest valid samples to the missing data sampling points are selected, and the missing values ​​are calculated and filled using a weighted average. Finally, the processed data are integrated to obtain standardized soil baseline data without outliers or missing values.

[0102] Step S202: Based on the apparent resistivity data in the geophysical exploration data, perform inversion to generate the true resistivity profile.

[0103] It should be noted that apparent resistivity data refers to the raw data reflecting the electrical conductivity of the subsurface medium, collected by an electrical resistivity system using the high-density resistivity method. This data is a preliminary characterization of the subsurface medium's resistivity, but it does not completely eliminate interference from external factors such as surface conditions and electrode arrangement, and therefore does not represent the true resistivity of the subsurface medium. Inversion refers to the process of processing the collected raw apparent resistivity data using specific mathematical algorithms and software tools to eliminate the influence of external interference factors and calculate the true resistivity of the subsurface medium at different depths and locations. The true resistivity profile refers to the two-dimensional or three-dimensional profile data obtained after inversion processing, which truly reflects the electrical conductivity of the subsurface medium. This profile data is continuous and complete, clearly showing the differences in resistivity distribution among soil layers at different depths.

[0104] Specifically, the raw apparent resistivity data was first extracted from geophysical survey data, ensuring the data format met the input requirements of the inversion software. Next, the apparent resistivity data was imported into the RES2DINV software and processed using the least squares method for two-dimensional inversion, initially eliminating some external interference. Then, the two-dimensional inversion results were imported into the pyGIMLi platform and further optimized using a Gauss-Newton inversion algorithm under smoothing constraints to accurately calculate the true resistivity of the subsurface medium. Finally, the inversion results were integrated to generate a continuous and complete true resistivity profile.

[0105] Step S203: Based on the actual resistivity profile, ordinary Kriging interpolation is used to calculate and obtain high-resolution resistivity covariate data.

[0106] It should be noted that high-resolution resistivity covariate data refers to resistivity data with a spatial resolution far higher than the original true resistivity profile after ordinary kriging interpolation. The data points are more densely distributed and have stronger spatial continuity, which can accurately characterize the subtle changes in the resistivity of underground media in three-dimensional space and be incorporated into the multi-dimensional covariate system as a core covariate.

[0107] Specifically, firstly, resistivity values ​​and corresponding spatial coordinates are extracted from the actual resistivity profile to construct a resistivity spatial sample dataset. Next, the variogram of the sample dataset is calculated to fit the optimal spatial structure model. Then, based on the spatial range of the industrial site and the preset resolution, interpolation grid nodes are set. Finally, ordinary kriging interpolation is used to perform unbiased optimal estimation of the resistivity values ​​at each interpolation grid node, generating high-resolution resistivity covariate data with dense spatial distribution and strong continuity.

[0108] Step S204: Based on the pollution source coordinates in the site-related basic data and the three-dimensional sampling coordinates of the soil samples, calculate the shortest straight-line distance in three-dimensional space from each sampling point to each pollution source, and generate a multi-dimensional pollution source distance covariate matrix.

[0109] It should be noted that pollution source coordinates refer to the three-dimensional spatial coordinates of various pollution sources extracted from the site's relevant basic data. This includes spatial location information such as longitude, latitude, and depth. Pollution sources cover the core areas within the site that generate pollutants, such as workshops, wastewater treatment areas, hazardous waste storage areas, and waste recycling points, and serve as the benchmark for distance calculation. Three-dimensional sampling coordinates refer to the three-dimensional spatial coordinates of each sampling point recorded during soil sample collection, accurately representing the horizontal position and vertical depth of the sampling point within the site. The shortest straight-line distance in three-dimensional space refers to the straight-line distance between the soil sampling point and the pollution source in three-dimensional space. The multi-dimensional pollution source distance covariate matrix is ​​a structured data matrix constructed with sampling points as rows, various pollution sources as columns, and corresponding shortest straight-line distances in three-dimensional space as matrix elements. Each value in the matrix is ​​an independent covariate, comprehensively representing the spatial distance relationship between each sampling point and different pollution sources.

[0110] Specifically, firstly, the three-dimensional spatial coordinates of all pollution sources were extracted from the site's relevant basic data, and the three-dimensional sampling coordinates of all soil samples were compiled to ensure a unified coordinate benchmark for subsequent calculations, guaranteeing consistency and accuracy. Next, for each soil sampling point, the shortest three-dimensional spatial distance between it and each pollution source within the site was calculated sequentially. This calculation quantifies the spatial association between each sampling point and the pollution sources, providing key spatial characteristic variables for subsequent analysis. Finally, using sampling points as the row dimension and various pollution sources as the column dimension, the calculated shortest three-dimensional spatial distances were integrated as matrix elements according to row and column correspondences to generate a multi-dimensional pollution source distance covariate matrix. This matrix comprehensively reflects the spatial distance relationship between each sampling point and all pollution sources, providing crucial spatial feature inputs for machine learning models and helping them more accurately capture the patterns and characteristics of pollution diffusion.

[0111] Step S205: Process the pollutant concentration data in the standardized soil basic data to obtain spatial regionalization index covariate data.

[0112] It should be noted that the spatial regionalization index covariate data is covariate data formed by combining the spatial regionalization index as the core indicator with the three-dimensional coordinates of the sampling points. The spatial regionalization index is a quantitative index characterizing the aggregation characteristics, variability, and distribution patterns of soil pollutant concentrations in three-dimensional space. It integrates the spatial autocorrelation, heterogeneity, and gradient change characteristics of pollutant concentrations, transforming discrete concentration data into a continuous index reflecting spatial distribution patterns and intuitively demonstrating the regional distribution of pollution. Spatial autocorrelation refers to the spatial correlation characteristics of soil pollutant concentrations, i.e., the similarity or difference in pollutant concentrations between adjacent sampling points.

[0113] Further, step S205 includes: extracting all pollutant concentration data and the corresponding three-dimensional coordinates of sampling points from the standardized soil baseline data, and constructing concentration-coordinate datasets according to pollutant type; for each pollutant concentration-coordinate dataset, using a spatial K-fold cross-validation strategy to partition the data, dividing the entire dataset into K spatially mutually exclusive subsets; for each target sampling point of a feature to be calculated, setting the spatial subset to which the target sampling point belongs as the validation set, and setting the remaining K-1 spatial subsets as the training set; removing the concentration information of the target sampling point itself, fitting a local variation function model based on the training set and combining it with regionalized variable theory, and determining the model parameters of the local variation function model, where the model parameters include range, sill value, and nugget value; based on the local variation function model and model parameters, using three-dimensional ordinary kriging to interpolate the local spatial region where the target sampling point is located, generating a subset of the continuous distribution field of pollutant concentration in the local region; and continuously distributing the pollutant concentration in the local region... The field subsets were imported into a geographic information processing tool, and the natural fault classification method was used to perform preliminary regional division of the continuous distribution field subsets of pollutant concentrations, resulting in multiple local division schemes with different numbers of sub-regions. The variance optimal fit value and inflection point feature value of each local division scheme were calculated, and the local division scheme with the largest variance optimal fit value and the inflection point feature value meeting the preset conditions was selected, and the corresponding optimal number of local sub-regions was determined. For the division results corresponding to the optimal number of local sub-regions, the arithmetic mean of pollutant concentrations at all interpolation points in each local sub-region was calculated. The arithmetic mean of pollutant concentrations was used as the local spatial regionalization index of the corresponding local sub-region, and the local spatial regionalization index was assigned to the target sampling points in the corresponding local sub-region to obtain the pollutant spatial regionalization index feature value corresponding to the sampling point. Point-by-point feature calculation was performed on all sampling points to obtain the spatial regionalization index covariate sub-data for each pollutant. The spatial regionalization index covariate sub-data of all pollutants were integrated to obtain the spatial regionalization index covariate data.

[0114] It should be noted that regionalized variable theory is a core theory in geostatistics. It posits that spatial variables such as soil pollutant concentrations possess both random and structural characteristics, exhibiting both local random variation and continuous structural changes within a certain spatial range. This theory forms the theoretical basis for fitting variogram models and conducting spatial interpolation. The variogram model is a mathematical model characterizing the spatial variability of regionalized variables. Its core parameters—range (maximum distance of spatial correlation), sill (total variability), and nugget (random variability)—precisely quantify the spatial variability patterns of pollutant concentrations. The natural fault classification method is a classification approach based on the natural distribution characteristics of data. It divides regions by identifying natural breaks in data values, minimizing data differences within the same sub-region and maximizing data differences between different sub-regions, thus closely reflecting the actual spatial clustering characteristics of pollutant concentrations. The variance optimal fit value is a quantitative indicator measuring the goodness of fit of the regional division scheme. A larger value indicates that the division scheme closely matches the natural distribution patterns of pollutant concentrations and accurately reflects the rationality of the regional division. The inflection point characteristic value is a key indicator for determining the rationality of the number of regional divisions. When the value meets the preset conditions, it indicates that the corresponding number of sub-regions can take into account both the precision and practicality of regional division, avoiding division that is too fine or too coarse.

[0115] Specifically, all pollutant concentration data were extracted from standardized soil baseline data, along with the corresponding 3D coordinates of soil sampling points for each set of concentration data. Based on different pollutant types, the concentration data and corresponding 3D coordinates were categorized and integrated to construct independent concentration-coordinate datasets for each pollutant. This approach ensures separate processing of different pollutants, preventing data confusion and ensuring the accuracy of subsequent feature construction. Then, for each pollutant's concentration-coordinate dataset, a spatial K-fold cross-validation strategy was employed to partition the data, dividing the entire dataset into K spatially mutually exclusive subsets. This ensures that sampling points within each subset are uniformly distributed in 3D space, avoiding partitioning bias caused by spatial autocorrelation and laying the foundation for subsequent local feature calculations.Next, for each target sampling point whose spatial regionalization index features need to be calculated, the spatial subset to which the target sampling point belongs is set as the validation set, and the remaining K-1 spatial subsets are used as the training set. Simultaneously, the concentration information of the target sampling point itself is completely removed, and only the concentration-coordinate data within the training set is used. A local variation function model is fitted using regionalization variable theory, and the three core model parameters—range, sill value, and nugget value—are precisely determined. This is done to completely avoid data leakage, ensuring that the feature construction of the target sampling point does not depend on its own concentration information, thus improving the reliability of the features. Afterwards, based on the proposed... The obtained local variation function model and its core parameters are analyzed using three-dimensional ordinary kriging. Interpolation is performed only on the local spatial region where the target sampling point is located, generating a subset of the continuous pollutant concentration distribution field for that local spatial region. This avoids interpolation calculations of the concentration at the target sampling point itself, thus preventing the interpolation results from interfering with the feature calculations of the target sampling point. Subsequently, the generated subset of the continuous pollutant concentration distribution field is imported into a geographic information processing tool. The natural fault classification method is used to perform preliminary regional division of this subset, resulting in multiple local division schemes with different numbers of sub-regions. This provides a backup for selecting the optimal regional division scheme. Next, the optimal variance fit value and inflection point characteristic value of each local partitioning scheme are calculated. Through comprehensive comparison, the local partitioning scheme with the largest optimal variance fit value and the inflection point characteristic value meeting the preset conditions is selected, and the optimal number of local sub-regions corresponding to this scheme is determined to ensure that the regional partitioning not only conforms to the local pollutant concentration distribution pattern, but also takes into account the rationality and practicality of the partitioning. Then, for the local partitioning results corresponding to the optimal number of local sub-regions, the arithmetic mean of pollutant concentrations at all interpolation points in each local sub-region is calculated one by one. This arithmetic mean is used as the local spatial regionalization index of the corresponding local sub-region, and then the local spatial regionalization index is assigned... For each target sampling point within a corresponding local sub-region, the spatial regionalization index feature value of the pollutant corresponding to that sampling point is obtained. Then, following the same steps as above, feature calculations are performed point-by-point for all sampling points to ensure that the construction of the spatial regionalization index feature of each sampling point follows the principle of no data leakage. Finally, the spatial regionalization index covariate sub-data corresponding to each pollutant is obtained. Finally, the spatial regionalization index covariate sub-data of all pollutants are integrated, and the data format and benchmark are unified to obtain complete spatial regionalization index covariate data, providing spatial regionalization feature support for the subsequent construction of a multi-dimensional covariate system.

[0116] Step S206: Calculate and process the three-dimensional sampling coordinates of the soil sample and the concentration data of each pollutant in the standardized soil baseline data to obtain spatial clustering characteristic covariate data.

[0117] It should be noted that the spatial aggregation characteristic covariate data are calculated using spatial statistical methods based on the three-dimensional sampling coordinates and pollutant concentration data of soil samples. These covariates quantify the degree, extent, and hotspot / colds distribution characteristics of pollutants in three-dimensional space. Spatial aggregation characteristics refer to the non-uniform distribution of soil pollutants in three-dimensional space, manifested as significantly higher pollutant concentrations in some areas forming pollution hotspots, and significantly lower concentrations in other areas forming pollution colds, while also exhibiting varying degrees of differences in aggregation density and extent.

[0118] Further, step S206 includes: extracting the three-dimensional sampling coordinates of the soil sample, and normalizing the X, Y, and Z axis data of the three-dimensional sampling coordinates using the min-max normalization method to obtain standardized three-dimensional sampling coordinates; constructing a distance bandwidth spatial weight matrix based on the standardized three-dimensional sampling coordinates; importing the concentration data of each pollutant, the standardized three-dimensional sampling coordinates, and the distance bandwidth spatial weight matrix from the standardized soil basic data using a spatial statistical analysis tool, calculating the local statistics corresponding to each pollutant concentration, and obtaining the Z value corresponding to each sampling point; setting a preset positive threshold and a preset negative threshold, identifying sampling points with Z values ​​greater than the preset positive threshold as pollution hotspots, and identifying sampling points with Z values ​​less than the preset negative threshold as pollution coldspots; calculating the first Euclidean distance from each sampling point to the nearest pollution hotspot and the second Euclidean distance from each sampling point to the nearest pollution coldspot based on the standardized three-dimensional sampling coordinates, pollution hotspots, and pollution coldspots; and integrating the first Euclidean distance and the second Euclidean distance to obtain spatial clustering feature covariate data.

[0119] It should be noted that the min-max normalization method is a linear data normalization method that eliminates dimensional differences by mapping data values ​​to the interval between 0 and 1. The distance-bandwidth spatial weight matrix is ​​a weight matrix constructed based on the three-dimensional spatial distance between sampling points. By setting a bandwidth threshold, it distinguishes the degree of spatial correlation between sampling points, and the weight values ​​quantify the spatial proximity relationship between them. The weight values ​​of the distance-bandwidth spatial weight matrix are assigned as follows: when the three-dimensional spatial distance between two sampling points is less than or equal to the bandwidth threshold, a first weight value is assigned; when the three-dimensional spatial distance between two sampling points is greater than the bandwidth threshold, a second weight value is assigned. In this embodiment, the first weight value is 1, and the second weight value is 0. Local statistics are quantitative indicators used in spatial statistical analysis to measure the correlation between the concentration of pollutants at a single sampling point and its surrounding sampling points. The corresponding Z-value reflects the degree of anomaly of the sampling point's concentration relative to the surrounding area; a larger Z-value indicates that the concentration at that point is significantly higher than the surrounding area, and vice versa. Three-dimensional Euclidean distance is the straight-line distance between two points in three-dimensional space, accurately representing the actual distance between a sampling point and hot / cold spots in three-dimensional space, reflecting the spatial correlation between the sampling point and pollution accumulation points.

[0120] Specifically, firstly, the three-dimensional sampling coordinates of the soil samples were extracted, and the X, Y, and Z axis data of the three-dimensional sampling coordinates were normalized using the min-max normalization method, scaling the value of each coordinate axis to between 0 and 1 to obtain standardized three-dimensional sampling coordinates. This process eliminates the influence of different coordinate axis dimensions, ensuring the consistency and accuracy of subsequent calculations. Next, based on the standardized three-dimensional sampling coordinates, a distance-bandwidth spatial weight matrix was constructed. A bandwidth threshold was set; when the three-dimensional spatial distance between two sampling points is less than or equal to the bandwidth threshold, it is assigned 1, indicating that the two points are spatially adjacent; when the distance is greater than the bandwidth threshold, it is assigned 0, indicating that the two points are not spatially adjacent. This weight matrix can effectively reflect the spatial proximity relationship between sampling points, providing a basis for local spatial statistical analysis. Then, a spatial statistical analysis tool was called to import the pollutant concentration data, standardized three-dimensional sampling coordinates, and distance-bandwidth spatial weight matrix from the standardized soil baseline data, and calculate the local statistics (such as Getis-Ord Gibbs) corresponding to each pollutant concentration. The Z-values ​​for each sampling point were obtained using statistical methods. The Z-values ​​reflect the spatial clustering characteristics of the sampling points. Preset positive and negative thresholds were set; sampling points with Z-values ​​greater than the preset positive threshold were identified as pollution hotspots, and those with Z-values ​​less than the preset negative threshold were identified as pollution colds. This threshold screening process clearly identifies pollution hotspots and colds. Based on standardized 3D sampling coordinates, pollution hotspots, and pollution colds, the first Euclidean distance from each sampling point to the boundary of the nearest hotspot region and the second Euclidean distance from each sampling point to the boundary of the nearest coldspot region were calculated. These distances reflect the spatial correlation between the sampling points and pollution hotspots and coldspots. Finally, the first and second Euclidean distances were integrated to obtain spatial clustering characteristic covariate data. This covariate data effectively reflects the spatial clustering characteristics of the sampling points and the pollution diffusion trend, providing important spatial characteristic information for subsequent heavy metal concentration distribution range prediction models.

[0121] Step S207: Through feature correlation analysis, redundancy processing and integration are performed on soil physicochemical property covariates, high-resolution resistivity covariates, multi-dimensional pollution source distance covariate matrix, spatial regionalization index covariates, and spatial aggregation characteristic covariates in the standardized soil basic data to construct a multi-dimensional covariate system.

[0122] It should be noted that characteristic correlation analysis refers to a statistical method that analyzes the degree of linear or nonlinear association between different covariates by calculating quantitative indicators such as Pearson correlation coefficient and Spearman correlation coefficient. The multidimensional covariate system includes soil physicochemical properties, geophysical attributes, spatial distance, and spatial distribution characteristics.

[0123] Specifically, a feature correlation analysis was first performed on the soil physicochemical properties covariate data, high-resolution resistivity covariate data, multi-dimensional pollution source distance covariate matrix, spatial regionalization index covariate data, and spatial clustering characteristic covariate data in the standardized soil baseline data. This process identified highly correlated feature pairs by calculating the correlation coefficient matrix between each covariate, thereby determining which features contain redundant information. Correlation analysis helps to select the most representative and independent features, avoiding the inclusion of too many highly correlated variables in the model, thus improving the model's stability and predictive ability.

[0124] Next, based on the results of the correlation analysis, redundant covariates are processed. For example, if the correlation coefficient between two covariates exceeds a preset threshold (e.g., 0.8), one of the covariates can be retained, or they can be merged into a new feature through feature fusion (e.g., averaging or principal component analysis). This process aims to reduce data dimensionality while retaining as much useful information as possible.

[0125] Then, the covariate data, after redundancy processing, are integrated to construct a multi-dimensional covariate system. This system covers soil physicochemical properties (such as organic matter content, porosity, pH value, etc.), geophysical properties (such as resistivity), spatial distances (such as distances to pollution sources, hot spots, and cold spots), and spatial distribution characteristics (such as spatial regionalization index). By integrating these different types of covariates, a comprehensive and independent set of features is formed, which can reflect the characteristics and distribution patterns of soil pollution from multiple perspectives.

[0126] Finally, the constructed multidimensional covariate system was used as input features for subsequent machine learning models, providing rich information support for predicting the distribution range of heavy metal concentrations. This multidimensional covariate system not only considers the physicochemical properties of the soil itself, but also combines spatial location and geophysical information, enabling it to more comprehensively capture the complexity and nonlinear characteristics of pollution distribution, thereby improving the accuracy and reliability of the heavy metal concentration distribution range prediction model.

[0127] This embodiment generates standardized soil baseline data by processing soil sampling data through outlier detection and K-nearest neighbor interpolation; it generates a true resistivity profile using geophysical survey data and obtains high-resolution resistivity covariate data through Kriging interpolation; it calculates the distance between sampling points and pollution sources to construct a pollution source distance covariate matrix; and it processes pollutant concentration data to obtain spatial regionalization index and spatial aggregation characteristic covariate data. Finally, it integrates multi-source data through feature correlation analysis to construct a multi-dimensional covariate system including soil physicochemical properties, geophysical attributes, spatial distance, and distribution characteristics. By comprehensively integrating multi-source data, it effectively reduces data redundancy, improves the accuracy and reliability of heavy metal concentration distribution range prediction, and provides a scientific basis for the refined management and remediation of contaminated sites.

[0128] Based on the first embodiment of this application, this application also provides a machine learning-based device for predicting the distribution of heavy metal concentrations in industrial sites. Please refer to... Figure 3 The device includes:

[0129] The acquisition module 10 is used to acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data.

[0130] Processing module 20 is used to preprocess multi-source data and construct a multi-dimensional covariate system.

[0131] The model building module 30 is used to construct various heavy metal concentration distribution range prediction models based on a multi-dimensional covariate system and employing multiple machine learning algorithms, including random forest algorithm, extreme gradient boosting algorithm, and lightweight gradient boosting machine algorithm.

[0132] The model selection module 40 is used to train and validate various heavy metal concentration distribution range prediction models and select the optimal prediction model.

[0133] The results module 50 is used to make a three-dimensional prediction of the distribution range of heavy metal concentration in an industrial site using the optimal prediction model, and to obtain the three-dimensional distribution results of heavy metal concentration in the industrial site.

[0134] The machine learning-based heavy metal concentration distribution prediction device for industrial sites provided in this application employs the machine learning-based heavy metal concentration distribution prediction method for industrial sites described in the above embodiments. It addresses the technical problem of how to combine multi-source data and machine learning algorithms to improve the accuracy of predicting the distribution range of heavy metal concentrations in industrial sites. Compared with existing technologies, the beneficial effects of the machine learning-based heavy metal concentration distribution prediction device for industrial sites provided in this application are the same as those of the machine learning-based heavy metal concentration distribution prediction method for industrial sites provided in the above embodiments. Furthermore, other technical features of the machine learning-based heavy metal concentration distribution prediction device for industrial sites are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0135] This application provides a machine learning-based heavy metal concentration distribution prediction device for industrial sites. The machine learning-based heavy metal concentration distribution prediction device for industrial sites includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the machine learning-based heavy metal concentration distribution prediction method for industrial sites in the above embodiment 1.

[0136] The following is for reference. Figure 4 The diagram illustrates a structural schematic of a machine learning-based heavy metal concentration distribution prediction device suitable for implementing embodiments of this application. The machine learning-based heavy metal concentration distribution prediction device for industrial sites in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The machine learning-based heavy metal concentration distribution prediction device for industrial sites shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0137] like Figure 4As shown, the machine learning-based industrial site heavy metal concentration distribution prediction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the machine learning-based industrial site heavy metal concentration distribution prediction device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the machine learning-based industrial site heavy metal concentration distribution prediction device to wirelessly or wiredly communicate with other devices to exchange data. Although various machine learning-based industrial site heavy metal concentration distribution prediction devices are shown in the figures, it should be understood that implementation or possession of all of them is not required. More or fewer may be implemented alternatively.

[0138] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0139] The machine learning-based heavy metal concentration distribution prediction device for industrial sites provided in this application employs the machine learning-based heavy metal concentration distribution prediction method for industrial sites described in the above embodiments. It addresses the technical problem of how to combine multi-source data and machine learning algorithms to improve the accuracy of predicting the distribution range of heavy metal concentrations in industrial sites. Compared with existing technologies, the beneficial effects of the machine learning-based heavy metal concentration distribution prediction device for industrial sites provided in this application are the same as those of the machine learning-based heavy metal concentration distribution prediction method for industrial sites provided in the above embodiments. Furthermore, other technical features of this machine learning-based heavy metal concentration distribution prediction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0140] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0141] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0142] This application provides a computer-readable medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites as described in the above embodiments.

[0143] The computer-readable medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, or any combination thereof. More specific examples of computer-readable media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable medium may be any tangible medium containing or storing a program that can be executed by instructions, used by a device, or used in conjunction with it. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0144] The aforementioned computer-readable medium may be included in a machine learning-based industrial site heavy metal concentration distribution prediction device; or it may exist independently and not assembled into a machine learning-based industrial site heavy metal concentration distribution prediction device.

[0145] The aforementioned computer-readable medium carries one or more programs that, when executed by a machine learning-based industrial site heavy metal concentration distribution prediction device, enable the machine learning-based industrial site heavy metal concentration distribution prediction device to write computer program code for performing the operations of this application in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of this application. In this regard, all blocks in the flowcharts or block diagrams may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that all blocks in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using dedicated hardware-based implementations that perform the specified functions or operations, or using a combination of dedicated hardware and computer instructions.

[0147] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0148] The readable medium provided in this application is a computer-readable medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-described machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites. This solves the technical problem of how to combine multi-source data and machine learning algorithms to improve the accuracy of predicting the distribution range of heavy metal concentrations in industrial sites. Compared with the prior art, the beneficial effects of the computer-readable medium provided in this application are the same as those of the machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites provided in the above embodiments, and will not be repeated here.

[0149] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites.

[0150] The computer program product provided in this application solves the technical problem of how to improve the accuracy of predicting the distribution range of heavy metal concentrations in industrial sites by combining multi-source data and machine learning algorithms. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the machine learning-based method for predicting the distribution of heavy metal concentrations in industrial sites provided in the above embodiments, and will not be repeated here.

[0151] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for predicting heavy metal concentration distribution in industrial sites based on machine learning, characterized in that, The method includes: Acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data; The multi-source data is preprocessed to construct a multi-dimensional covariate system; Based on the multi-dimensional covariate system, various machine learning algorithms are used to construct various heavy metal concentration distribution range prediction models, including random forest algorithm, extreme gradient boosting algorithm and lightweight gradient boosting machine algorithm. Multiple heavy metal concentration distribution range prediction models were trained and validated, and the optimal prediction model was selected. The optimal prediction model is used to make a three-dimensional prediction of the distribution range of heavy metal concentration in industrial sites, and the three-dimensional distribution results of heavy metal concentration in industrial sites are obtained. The steps of preprocessing the multi-source data and constructing a multi-dimensional covariate system include: Outlier detection was performed on pollutant content data in soil sampling data to obtain standardized basic soil data. The outlier detection was divided into removing outlier data that exceeded the preset confidence level using the Grubbs criterion and supplementing missing data using the K-nearest neighbor interpolation method based on soil type grouping. Inversion is performed based on apparent resistivity data from geophysical exploration data to generate a true resistivity profile. High-resolution resistivity covariate data were obtained by using ordinary Kriging interpolation based on the actual resistivity profile. Based on the pollution source coordinates in the site-related basic data and the three-dimensional sampling coordinates of the soil samples, the three-dimensional shortest straight-line distance from each sampling point to each pollution source is calculated, and a multi-dimensional pollution source distance covariate matrix is ​​generated. The pollutant concentration data in the standardized soil baseline data are processed to obtain spatial regionalization index covariate data; The spatial clustering characteristic covariate data are obtained by calculating and processing the three-dimensional sampling coordinates of the soil samples and the concentration data of each pollutant in the standardized soil basic data. By performing feature correlation analysis, redundancy processing and integration are performed on the soil physicochemical property covariates, the high-resolution resistivity covariate data, the multidimensional pollution source distance covariate matrix, the spatial regionalization index covariate data, and the spatial clustering feature covariate data in the standardized soil basic data to construct a multidimensional covariate system. The multidimensional covariate system includes soil physicochemical properties, geophysical attributes, spatial distance, and spatial distribution characteristics. The step of using the optimal prediction model to perform three-dimensional prediction of the heavy metal concentration distribution range of the industrial site, and obtaining the three-dimensional distribution result of heavy metal concentration in the industrial site, includes: Based on the geographical boundaries, stratum distribution depth, and three-dimensional coordinates of sampling points of the industrial site, a three-dimensional grid of the industrial site is generated. The node spacing in the X and Y directions of the three-dimensional grid is set according to the spatial heterogeneity of site pollution, and the node spacing in the Z direction matches the soil layer thickness. The multidimensional covariate system is mapped to each node of the three-dimensional grid using three-dimensional kriging interpolation to form a grid covariate dataset. Input the grid covariate dataset into the optimal prediction model and output the predicted pollutant concentration value for each grid node; Based on the training and verification error distribution pattern of the optimal prediction model, the confidence interval and confidence level corresponding to the predicted pollutant concentration value of each grid node are calculated. Set preset screening thresholds for each pollutant based on preset screening criteria; The predicted pollutant concentration of each grid node is combined with the confidence interval for comprehensive judgment, and the set of grid nodes that exceed the preset screening threshold and whose lower edge of the confidence interval is not lower than the preset screening threshold is marked. By combining the site stratigraphic structure data, the set of grid nodes exceeding the standard is spatially topologically reconstructed to generate a three-dimensional pollution dataset, wherein the three-dimensional pollution dataset includes the spatial coordinates of grid nodes, soil layer attribution, the level of exceeding the standard concentration, and the corresponding confidence level. Interactive 3D volume rendering technology is used to perform correlation rendering and visualization processing on the 3D pollution dataset by combining confidence level and concentration level, so as to obtain the 3D distribution results of heavy metal concentration in industrial sites.

2. The method as described in claim 1, characterized in that, The step of processing the pollutant concentration data in the standardized soil baseline data to obtain the spatial regionalization index covariate data includes: Extract all pollutant concentration data and the three-dimensional coordinates of the corresponding sampling points from the standardized soil baseline data, and construct concentration-coordinate datasets according to pollutant type; For each pollutant's concentration-coordinate dataset, a spatial K-fold cross-validation strategy is used to partition the data, dividing the entire dataset into K spatially mutually exclusive subsets; For each target sampling point of a feature to be calculated, the spatial subset to which the target sampling point belongs is set as the validation set, and the remaining K-1 spatial subsets are set as the training set. The concentration information of the target sampling point itself is removed, and a local variation function model is obtained by fitting the training set with regional variable theory. The model parameters of the local variation function model are determined, including range, sill value and nugget value. Based on the local variation function model and the model parameters, the three-dimensional ordinary kriging method is used to interpolate the local spatial region where the target sampling point is located, generating a subset of the continuous distribution field of pollutant concentration in the local region; The subset of the continuous distribution field of pollutant concentration in the local area is imported into a geographic information processing tool. The natural fault classification method is used to perform preliminary regional division of the subset of the continuous distribution field of pollutant concentration, resulting in multiple local division schemes with different numbers of sub-regions. Calculate the variance optimal fit value and inflection point feature value for each group of local partitioning schemes, select the local partitioning scheme with the largest variance optimal fit value and the inflection point feature value that meets the preset conditions, and determine the corresponding optimal number of local sub-regions. For the division results corresponding to the optimal number of local sub-regions, the arithmetic mean of pollutant concentrations at all interpolation points in each local sub-region is calculated. The arithmetic mean of pollutant concentrations is used as the local spatial regionalization index of the corresponding local sub-region, and the local spatial regionalization index is assigned to the target sampling point in the corresponding local sub-region to obtain the pollutant spatial regionalization index feature value corresponding to the sampling point. Point-by-point feature calculations were performed on all sampling points to obtain spatial regionalization index covariate sub-data for each pollutant. By integrating the spatial regionalization index covariate sub-data of all pollutants, spatial regionalization index covariate data is obtained.

3. The method as described in claim 1, characterized in that, The step of calculating and processing the three-dimensional sampling coordinates of the soil sample and the concentration data of each pollutant in the standardized soil baseline data to obtain spatial clustering characteristic covariate data includes: The three-dimensional sampling coordinates of the soil samples were extracted, and the X, Y, and Z axis data of the three-dimensional sampling coordinates were normalized by the min-max normalization method to obtain the standardized three-dimensional sampling coordinates. Based on the standardized three-dimensional sampling coordinates, a distance bandwidth type spatial weight matrix is ​​constructed; The spatial statistical analysis tool is used to import the concentration data of each pollutant in the standardized soil basic data, the standardized three-dimensional sampling coordinates, and the distance bandwidth spatial weight matrix. The local statistics corresponding to each pollutant concentration are calculated to obtain the Z value corresponding to each sampling point. Set a preset positive threshold and a preset negative threshold, and determine the sampling points with Z values ​​greater than the preset positive threshold as pollution hotspots, and determine the sampling points with Z values ​​less than the preset negative threshold as pollution colds; Based on the standardized three-dimensional sampling coordinates, the pollution hotspots, and the pollution coldspots, the first Euclidean distance from each sampling point to the nearest pollution hotspot and the second Euclidean distance from each sampling point to the nearest pollution coldspot are calculated respectively. Integrating the first Euclidean distance and the second Euclidean distance yields spatial clustering feature covariate data.

4. The method as described in claim 1, characterized in that, The steps of training and validating various heavy metal concentration distribution range prediction models and selecting the optimal prediction model include: Key feature variables were extracted from the multidimensional covariate system, and a model training dataset was constructed by combining measured soil pollutant concentration data. The training dataset of the model is divided into training set and validation set by adopting a spatial block cross-validation strategy. The spatial block cross-validation strategy is to divide the industrial site into three-dimensional spatial blocks, and the sampling points in each spatial block are allocated as a whole to the training set or validation set. The initial random forest prediction model, the initial extreme gradient boosting prediction model, and the initial extreme gradient boosting prediction model are trained and optimized using the training set to obtain the optimized random forest prediction model, the optimized extreme gradient boosting prediction model, and the optimized lightweight gradient boosting machine prediction model. The optimization random forest prediction model, the optimization extreme gradient boosting prediction model, and the optimization lightweight gradient boosting machine prediction model are calculated using the validation set to obtain the coefficient of determination and root mean square error. The model with the best overall performance is selected as the optimal prediction model based on the coefficient of determination and root mean square error.

5. The method as described in claim 4, characterized in that, The step of training and optimizing the initial random forest prediction model, the initial extreme gradient boosting prediction model, and the initial extreme gradient boosting prediction model using the training set to obtain the optimized random forest prediction model, the optimized extreme gradient boosting prediction model, and the optimized lightweight gradient boosting machine prediction model includes: The training set is input into the initial random forest prediction model. Recursive feature elimination combined with 5-fold cross-validation is used to select the feature subset that contributes the most to the prediction of pollutant concentration. The hyperparameters of decision tree number, maximum depth and minimum number of split samples are optimized simultaneously to obtain the optimized random forest prediction model. The training set is input into the initial extreme gradient boosting prediction model. The error change is monitored by the early stopping mechanism, and the learning rate, subsampling rate and L1 / L2 regularization coefficient hyperparameters are optimized to obtain the optimized extreme gradient boosting prediction model. The training set is input into the initial lightweight gradient booster prediction model. Based on histogram optimization and leaf growth strategy, the hyperparameters of the upper limit of the number of leaves, the minimum data segmentation threshold, and the feature score calculation method are adjusted to obtain the optimized lightweight gradient booster prediction model.

6. A machine learning-based device for predicting heavy metal concentration distribution in industrial sites, characterized in that, The device is applied to the machine learning-based method for predicting heavy metal concentration distribution in industrial sites as described in any one of claims 1-5, and the device comprises: The acquisition module is used to acquire multi-source data of the industrial site, including soil sampling data, geophysical exploration data, and site-related basic data. The processing module is used to preprocess the multi-source data and construct a multi-dimensional covariate system; The model building module is used to construct various heavy metal concentration distribution range prediction models based on the multi-dimensional covariate system and employing various machine learning algorithms, including random forest algorithm, extreme gradient boosting algorithm and lightweight gradient boosting machine algorithm. The model selection module is used to train and validate various prediction models for the concentration distribution range of heavy metals, and select the optimal prediction model. The results module is used to perform three-dimensional prediction of the heavy metal concentration distribution range of the industrial site using the optimal prediction model, and obtain the three-dimensional distribution results of heavy metal concentration in the industrial site.

7. A machine learning-based device for predicting heavy metal concentration distribution in industrial sites, characterized in that, The device includes: a memory, a processor, and a machine learning-based industrial site heavy metal concentration distribution prediction program stored in the memory and running on the processor, the machine learning-based industrial site heavy metal concentration distribution prediction program being configured to implement the steps of the machine learning-based industrial site heavy metal concentration distribution prediction method as described in any one of claims 1-5.

8. A storage medium, characterized in that, The storage medium stores a machine learning-based industrial site heavy metal concentration distribution prediction program, which, when executed by a processor, implements the steps of the machine learning-based industrial site heavy metal concentration distribution prediction method as described in any one of claims 1-5.