Rapid Investigation Method for Soil VOCs Pollution Based on MIP and Machine Learning
By combining membrane interface detection system and machine learning technology, a prediction model is constructed to conduct quantitative prediction of soil VOCs pollutant concentration, which solves the problem of MIP technology being unable to accurately quantify inversion and signal interference, and realizes the accurate characterization of the three-dimensional spatial distribution of volatile organic matter concentration.
Patent Information
- Application Number
- CN202410599068.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-05-14
AI Technical Summary
The existing MIP technology cannot accurately quantify the concentration of inversion volatile organic compounds, and the probe will cause interference and distortion of subsequent measurement signals after passing through the high-concentration contaminated area, affecting the accurate judgment of the pollution depth position.
A rapid soil VOCs pollution survey method based on membrane interface detection system and machine learning is adopted. By obtaining the actual pollutant concentration value and parameter data of soil samples, the machine learning model is trained and a prediction model is constructed to conduct efficient quantitative prediction of volatile organic matter concentration.
The accurate depiction of the three-dimensional spatial distribution of volatile organic compounds is achieved, which alleviates the problem of signal interference after the probe passes through the high concentration zone, and improves the accuracy and efficiency of pollutant concentration prediction.
Smart Images

Figure CN118711694B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of volatile organic compound concentration prediction, and particularly to a rapid investigation method for soil VOCs pollution based on MIP and machine learning. Background Art
[0002] Volatile Organic Compounds (VOCs) are common and highly harmful pollutants in chemical sites. They can come into contact with the human body through ways such as soil gas intrusion and drinking groundwater, posing a great threat to human health. Affected by conditions such as the distribution of pollution sources, soil properties, and migration pathways, the distribution of VOCs in soil and groundwater has significant spatial heterogeneity characteristics, which restricts its accurate assessment and precise control.
[0003] The Membrane Interface Probe (MIP) is a screening tool with semi - quantitative functions. The probe can be driven to a specified depth by pushing or knocking. Combined with different probes, it can be used to determine the lithology in the unsaturated zone and saturated zone and estimate the relative concentration of volatile organic compounds (VOCs). VOC molecules existing in the gaseous, dissolved, solid, or free product phases diffuse through a heated semi - permeable membrane and are transported to a specific detector on the surface. According to the intensity of the voltage response signal value of the pollutant, the concentration of the pollutant is semi - quantitatively judged as high or low. By recording the positions of the probe at different depths underground, the VOC concentration at different depth positions can be detected. The MIP detects the direct signal related to VOCs, is relatively less interfered by other underground factors, and different detectors at the back end have different sensitivities to different types of VOCs. Therefore, the types of VOCs can be distinguished to a certain extent.
[0004] However, the MIP method can only perform semi - quantitative analysis at present. It can roughly judge the degree of underground VOCs pollution through the signal value range and cannot perform precise concentration inversion. In addition, during the MIP investigation process, when the probe passes through a high - concentration pollution area, due to factors such as the adsorption and retention of VOCs in the probe and carrier gas pipeline, and the disturbance of the probe to the equilibrium state of VOCs in the soil, after the MIP probe passes through the high - concentration area, it will interfere with the pollution concentration signal of the subsequent measurement points, resulting in the distortion of the subsequent data and affecting the accurate judgment of the actual pollution depth position. Summary of the Invention
[0005] In view of this, the present invention provides a rapid investigation method for soil VOCs pollution based on MIP and machine learning to solve the problem that rapid investigation technologies such as MIP for VOCs cannot perform precise quantitative inversion, so as to achieve the precise characterization of the three - dimensional spatial distribution of the concentration of pollutants such as volatile organic compounds.
[0006] In a first aspect, the present invention provides a rapid investigation method for soil VOCs pollution based on MIP and machine learning, the method comprising:
[0007] Obtaining the actual pollutant concentration values of soil samples in several boreholes in the investigation area;
[0008] Based on the membrane interface detection system, sampling is carried out at different depths for each borehole to obtain a number of sampling points and the parameter data corresponding to each sampling point;
[0009] Combining the actual pollutant concentration values of the soil samples and the parameter data to train a machine learning model to obtain a prediction model;
[0010] Based on the membrane interface detection system, detection data is obtained, and the detection data is input into the prediction model to obtain the predicted volatile organic compound concentration value.
[0011] The rapid investigation method for soil VOCs pollution based on MIP and machine learning provided by the embodiments of the present invention conducts on-site rapid high-density investigation of soil volatile organic compounds by adopting a membrane interface detection system, constructs an inversion model, that is, a prediction model, between the rapid investigation signal, that is, the parameter data, and the actual pollutant concentration, and efficiently quantitatively predicts the pollutant concentration, which helps to quickly master the basic situation of the site pollution distribution. Combining with the machine learning algorithm model, based on the simple and effective non-linear fitting means of the machine learning algorithm model, that is, the machine learning model, it can effectively capture the quantitative relationship between the rapid investigation signal and the pollutant concentration under the influence of various factors, quantitatively predict the volatile organic compound concentration through the investigation signal of the membrane interface detection system, and then combine three-dimensional visualization modeling to efficiently and highly resolve the spatial distribution of pollutants.
[0012] In an optional implementation manner, the membrane interface detection system includes a photoionization detector and a flame ionization detector. Based on the membrane interface detection system, sampling is carried out at different depths for each borehole to obtain a number of sampling points and the parameter data corresponding to each sampling point, including:
[0013] Based on the membrane interface detection system, sampling is carried out at different depths for each borehole to obtain a number of sampling points and the sampling data corresponding to each sampling point, the sampling data including photoionization detector signal values, flame ionization detector signal values, soil conductivity, and the depth of the sampling point;
[0014] Based on the sampling data of the sampling points, the characteristic variables corresponding to the sampling points are obtained, and the sampling data and the characteristic variables constitute the parameter data.
[0015] The volatile organic compound concentration prediction method provided by the embodiments of the present invention combines the advantages of in-situ, rapid, low-cost, and less interference factors of the membrane interface detection system in the investigation of volatile organic compound concentrations, sets characteristic variables to break through the limitation that the investigation results of the original membrane interface detection system cannot accurately and quantitatively invert the pollutant concentration, combines soil conductivity, photoionization detector signal values, flame ionization detector signal values, collection depth, etc. through machine learning, establishes a quantitative prediction model from rapid investigation data to pollutant concentration, and combines three-dimensional spatial modeling to accurately depict the three-dimensional spatial distribution of pollutant concentration.
[0016] In an alternative embodiment, obtaining the characteristic variables corresponding to the sampling points based on the sampling data of the sampling points includes:
[0017] Taking the sampling points corresponding to several maximum values in the photoionization detector signal values as the photoionization detection peak points;
[0018] Taking the sampling points corresponding to several maximum values in the flame ionization detector signal values as the flame ionization detection peak points;
[0019] Determining the characteristic variables corresponding to the sampling points based on the photoionization detection peak points and the flame ionization detection peak points, where the characteristic variables include: the signal value of the previous PID peak point of the sampling point, the signal value of the previous FID peak point of the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
[0020] The volatile organic compound concentration prediction method provided by the embodiments of the present invention solves the problem of the lag effect caused by the residual volatile organic compounds in the pipeline after the MIP probe passes through the high-concentration pollution area by setting characteristic variables.
[0021] In an alternative embodiment, training the machine learning model with the actual pollutant concentration values and parameter data of soil samples to obtain the prediction model includes:
[0022] Dividing the actual pollutant concentration values and parameter data of the soil samples into a training set and a test set according to a preset ratio;
[0023] Training the machine learning model with the training set to obtain the prediction model;
[0024] Evaluating the effect of the prediction model based on the test set.
[0025] The volatile organic compound concentration prediction method provided by the embodiments of the present invention trains the prediction model through the machine learning model combined with the machine learning algorithm and the characteristic variables. The main influencing factors of the lag effect are considered in the machine learning model, which can alleviate the influence of the lag effect to a certain extent and further improve the accuracy of the prediction results.
[0026] In an alternative embodiment, obtaining detection data based on the membrane interface detection system and inputting the detection data into the prediction model to obtain the predicted volatile organic compound concentration value includes:
[0027] Obtaining a number of prediction points based on the prediction area, obtaining the detection data of the prediction points based on the membrane interface detection system, and obtaining the prediction characteristic variables corresponding to the prediction points based on the detection data;
[0028] Inputting the detection data and prediction characteristic variables of the prediction points into the prediction model to predict the concentration of volatile organic compounds, and obtaining the predicted volatile organic compound concentration value of the prediction points.
[0029] The method for predicting the concentration of volatile organic compounds provided by the embodiments of the present invention obtains detection data through the membrane interface detection system, inputs the detection data into the prediction model to predict the concentration of volatile organic compounds, and can obtain the concentration value and spatial coordinates of volatile organic compounds without actual drilling and sampling detection, saving the time, manpower and material resources of exploration, and improving the efficiency of the investigation of the concentration of volatile organic compounds.
[0030] In an alternative embodiment, obtaining detection data based on the membrane interface detection system and inputting the detection data into the prediction model to obtain the predicted volatile organic compound concentration value further includes:
[0031] Visualizing the three-dimensional spatial distribution of the volatile organic compound concentration based on the volatile organic compound concentration value and detection data of the prediction points;
[0032] Inputting the volatile organic compound concentration value of the prediction points, the longitude, latitude and depth of the prediction points detected by the membrane interface detection system into three-dimensional visualization software, and performing three-dimensional spatial interpolation in combination with the formation lithology to obtain the visualization of the three-dimensional spatial distribution of the volatile organic compound concentration.
[0033] The method for predicting the concentration of volatile organic compounds provided by the embodiments of the present invention can achieve low-cost, high-efficiency and high-precision characterization of the spatial characteristics of pollution through encrypted investigation and prediction using the membrane interface detection system, analyze the spatial distribution characteristics of VOCS pollutants, and can provide auxiliary support for pollution source analysis, migration process judgment, etc.
[0034] In a second aspect, the present invention provides a device for predicting the concentration of volatile organic compounds, the device includes:
[0035] A membrane interface detection module for obtaining the actual pollutant concentration values of soil samples in a number of boreholes in the investigation area;
[0036] A model construction module for sampling each borehole at different depths based on the membrane interface detection system to obtain a number of sampling points and the sampling data corresponding to each sampling point;
[0037] A model training module, configured to train a machine learning model by combining the actual pollutant concentration values of soil samples and sampling data to obtain a prediction model;
[0038] A concentration value prediction module, configured to obtain detection data based on a membrane interface detection system, input the detection data into the prediction model, and obtain a predicted volatile organic compound concentration value.
[0039] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the volatile organic compound concentration prediction method according to the first aspect or any corresponding embodiment thereof.
[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the volatile organic compound concentration prediction method according to the first aspect or any corresponding embodiment thereof.
[0041] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the volatile organic compound concentration prediction method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0042] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is a schematic flowchart of a volatile organic compound concentration prediction method according to an embodiment of the present invention;
[0044] Figure 2 is a schematic flowchart of another volatile organic compound concentration prediction method according to an embodiment of the present invention;
[0045] Figure 3 is a schematic flowchart of yet another volatile organic compound concentration prediction method according to an embodiment of the present invention;
[0046] Figure 4 is a schematic flowchart of yet another volatile organic compound concentration prediction method according to an embodiment of the present invention;
[0047] Figure 5 is a structural block diagram of a volatile organic compound concentration prediction device according to an embodiment of the present invention;
[0048] Figure 6 It is a schematic diagram of the hardware structure of the computer device according to an embodiment of the present invention;
[0049] Figure 7 It is the process technical route of the volatile organic compound concentration prediction method according to an embodiment of the present invention;
[0050] Figure 8 It is the fitting coefficient and scatter plot of the random forest algorithm model (RF) according to an embodiment of the present invention on the test set;
[0051] Figure 9 It is the fitting coefficient and scatter plot of the gradient boosting tree algorithm model (LightGBM) according to an embodiment of the present invention on the test set;
[0052] Figure 10 It is the fitting coefficient and scatter plot of the generalized regression neural network algorithm model (GRNN) according to an embodiment of the present invention on the test set;
[0053] Figure 11 It is the fitting coefficient and scatter plot of the linear regression algorithm model (Linear) according to an embodiment of the present invention on the test set;
[0054] Figure 12 It is the line chart comparing the predicted values and actual values of four machine learning models according to an embodiment of the present invention;
[0055] Figure 13 It is the three-dimensional spatial distribution diagram of the volatile organic compound concentration values predicted by the gradient boosting tree algorithm model (LightGBM) according to an embodiment of the present invention;
[0056] Figure 14 It is the three-dimensional spatial distribution diagram of the volatile organic compound concentration values predicted by the generalized regression neural network algorithm model (GRNN) according to an embodiment of the present invention;
[0057] Figure 15 It is the three-dimensional spatial distribution diagram of the volatile organic compound concentration values predicted by the linear regression algorithm model (Linear) according to an embodiment of the present invention;
[0058] Figure 16 It is the three-dimensional spatial distribution diagram of the volatile organic compound concentration values predicted by the random forest algorithm model (RF) according to an embodiment of the present invention. Detailed implementation manners
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] In the prediction of soil volatile organic compound (VOC) concentration, the ground penetrating radar method emits electromagnetic pulses underground through equipment on the ground surface. It is mainly an indirect pollution detection method based on different physical property interfaces and burial depths of media. Therefore, it can only be identified when there are obvious physical property differences between polluted soil and non-polluted soil. As a result, its application is highly limited and it is extremely vulnerable to interference from different physical properties such as underground pipelines, storage tanks, and ditches, leading to distorted detection results. The high-density resistivity method determines the resistivity distribution of underground media by sending current underground and measuring the voltage difference between multiple electrodes, thereby evaluating soil pollution. However, it is extremely vulnerable to interference from soil humidity, temperature changes, electrolyte concentration, and structures such as underground pipeline storage tanks. The detection and inversion effect is poor for complex areas of stratum structure and underground construction. The optical image analysis technique (OIP) can achieve three-dimensional visualization of pollutants. However, as the depth increases, light attenuation is severe, resulting in weakened visualization ability for deep soil structure and pollutants. For turbid, highly viscous, or dark-colored soil, its imaging quality and resolution may be affected. OIP has good recognition of organic pollutants with obvious fluorescence characteristics such as polycyclic aromatic hydrocarbons (PAHs), but not all pollutants can produce fluorescence reactions. Therefore, there may be blind spots in the detection of certain specific pollutants. The principle and limitations of the laser-induced fluorescence detector (LIP) are similar to those of OIP. It is only sensitive to NAPL containing fluorescent components or pollutants with added fluorescent labels, and cannot detect pollutants without fluorescent active substances. Moreover, the particle size, humidity, temperature, and background fluorescence of the soil may all affect the intensity of the fluorescence signal, thereby affecting the accuracy of the detection results. Compared with the above four methods, the MIP method can directly transmit VOCs in the soil to the backend detector for detection through a pipeline transmission system. The correlation between the signal and the pollutant concentration is more direct, less affected by external factors, and less prone to misjudgment, with obvious advantages. However, this method can currently only perform semi-quantitative analysis, making a rough judgment on the degree of underground VOC pollution through the signal value range, and cannot perform precise concentration inversion. In addition, during the MIP investigation process, after the probe passes through the high-concentration pollution area, due to various factors such as the adsorption and retention of VOCs in the probe and the carrier gas pipeline, and the disturbance of the probe to the equilibrium state of VOCs in the soil, it will cause interference to the pollution concentration signal of the subsequent measurement points after the MIP probe passes through the high-concentration area, resulting in distortion of the subsequent data and affecting the accurate judgment of the actual pollution depth position. The embodiments of the present invention provide a rapid investigation method for soil VOC pollution based on MIP and machine learning, which constructs a prediction model specifically to achieve precise quantitative inversion, thereby achieving the effect of accurately depicting the three-dimensional spatial distribution of pollutant concentration and solving the problem of hysteresis effect caused by VOC residue in the pipeline.
[0061] According to an embodiment of the present invention, an embodiment of a rapid soil VOCs pollution investigation method based on MIP and machine learning is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0062] In this embodiment, a rapid soil VOCs pollution investigation method based on MIP and machine learning is provided, which can be used for mobile terminals such as mobile phones and tablets. Figure 1 It is a flowchart of a volatile organic compound concentration prediction method according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps:
[0063] Step S101, obtain the actual pollutant concentration values of soil samples in several boreholes in the investigation area.
[0064] In this embodiment, a typical chemical industrial site is selected as the investigation area. According to the analysis of pollution history and hydrogeological data, combined with the topographic map of the investigation area and the on-site reconnaissance situation, survey lines are arranged perpendicular to the underground water flow direction. The spacing between survey line measurement points is controlled within the range of 10 to 20 meters. For key pollution areas, encrypted detection can be carried out at intervals of 1 to 5 meters. Dispersedly select some measurement points to carry out soil borehole sampling surveys, record geological data according to the drilling results, with a vertical interval of 0.5 to 3 meters, to obtain the actual pollutant concentration values of soil samples in several boreholes. The actual pollutant is taken as total petroleum hydrocarbon (TPH) as an example to obtain the actual pollutant concentration value of petroleum hydrocarbon, and the three-dimensional spatial information of the borehole including longitude, latitude and depth is obtained through real-time kinematic (RTK) technology.
[0065] Step S102, based on the membrane interface detection system, sample each borehole at different depths to obtain several sampling points and the parameter data corresponding to each sampling point.
[0066] Use a vehicle-mounted MIP system to sample each borehole in the investigation area at different depths. The vehicle-mounted MIP detector advances at a speed of 30 cm / min, and the MIP probe has a residence time of about 1 minute at each depth, so that the volatile substances in the formation diffuse through the membrane and flow upward into the MIP detection system. The detector collects and records data once every 1.5 cm vertical interval in each borehole, that is, about 66 signal values are collected per meter on average in the same borehole, so as to obtain the parameter data of sampling points at different depths.
[0067] Step S103: Train the machine learning model with the actual pollutant concentration value and parameter data of the soil sample to obtain a prediction model.
[0068] Based on the machine learning model, i.e., the Light Gradient Boosting Machine (LightGBM), Random Forests (RF), Generalized Regression Neural Network (GRNN), and Linear Regression algorithms, train them respectively to obtain four trained machine learning models. A single trained machine learning model can directly be the prediction model, or two or more of the four trained machine learning models can be integrated as the prediction model. To obtain the optimal prediction model fitting effect, first use the model hyperparameter tuning method based on Bayesian optimization, i.e., Hyperopt, to train the four algorithm models respectively, so as to automatically tune the model parameters, find the optimal parameter combination in the given parameter space, and make the Root Mean Square Error (RMSE) of training reach the minimum value; add the actual pollutant concentration value and parameter data of the soil sample into the prediction model for training, compare the output of the machine learning model with the actual pollutant concentration value, and repeatedly train the parameters of the machine learning model according to the comparison result to obtain the mapping relationship between the parameter data and the output of the machine learning model. The machine learning model training is completed until the output of the machine learning model is the same as the actual pollutant concentration value, and the machine learning model at this time is used as the prediction model.
[0069] Step S104: Obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted volatile organic compound concentration value.
[0070] After obtaining the prediction model that can reflect the actual pollutant concentration value of the soil sample, use this prediction model to predict the volatile organic compound concentration value of the detection area where actual drilling sampling has not been carried out. Use the vehicle-mounted MIP system to obtain the detection data of the detection area, and input the detection data into the prediction model to obtain the volatile organic compound concentration value corresponding to the detection data.
[0071] The volatile organic compound concentration prediction method provided in this embodiment can conduct on-site rapid high-density surveys of soil volatile organic compounds by using a membrane interface detection system, and construct an inversion model between the rapid survey signals and pollutant concentrations to obtain a prediction model, so as to efficiently and quantitatively predict pollutant concentrations based on detection data, which helps to quickly grasp the basic situation of site pollution distribution, and combine three-dimensional visualization modeling to efficiently and highly resolve the spatial distribution of pollutants, such as Figure 7 is the process technical route of the volatile organic compound concentration prediction method of the embodiment of the present invention.
[0072] In this embodiment, a rapid soil VOCs pollution survey method based on MIP and machine learning is provided, which can be used for the above-mentioned mobile terminals, such as mobile phones, tablets, etc., Figure 2 is the flowchart of the volatile organic compound concentration prediction method according to the embodiment of the present invention, such as Figure 2 As shown, this process includes the following steps:
[0073] Step S201, obtain the actual pollutant concentration values of soil samples in several boreholes in the survey area. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.
[0074] Step S202, based on the membrane interface detection system, sample each borehole at different depths to obtain several sampling points and the parameter data corresponding to each sampling point.
[0075] Specifically, the above step S202 includes:
[0076] Step S2021, based on the membrane interface detection system, sample each borehole at different depths to obtain several sampling points and the sampling data corresponding to each sampling point. The sampling data includes photoionization detector signal value, flame ionization detector signal value, soil electrical conductivity, and the depth of the sampling point.
[0077] The sampling data is a part of the parameter data, providing a data basis for obtaining the prediction model.
[0078] The photoionization detector is Photo Ionization Detector (PID), and the photoionization detector signal value is the PID signal value; the flame ionization detector is Flame Ionization Detector (FID), and the flame ionization detector signal value is the FID signal value, and the soil electrical conductivity is Electrical Conductivity (EC).
[0079] Step S2022, obtain the characteristic variables corresponding to the sampling points based on the sampling data of the sampling points. The sampling data and the characteristic variables constitute the parameter data.
[0080] Both the sampled data and the characteristic variables are input variables for training the prediction model. The sampled data is obtained by the in-vehicle MIP system, and the characteristic variables are obtained by further processing the sampled data.
[0081] In some alternative embodiments, the above step S2022 includes:
[0082] Step a1: Take the sampling points corresponding to several maximum values in the photoionization detector signal value as the photoionization detection peak points.
[0083] The photoionization detection peak points are the PID peak points. Several maximum values, i.e., peaks, are equivalent to the maximum points in the function curve. The peak is compared with its previous signal value and the next signal value. When the signal value is higher than both the previous signal value and the next signal value, then it is a peak. Here, the previous and the next represent the previous and the next in the vertical direction of the sampling points in the borehole. The concentration of volatile organic compounds is not evenly distributed but is sometimes low, sometimes high, and sometimes low. Therefore, there may be multiple peaks in the same borehole. When the PID signal value is higher than both the previous PID signal value and the next PID signal value, this PID signal value is the PID peak point, and there are multiple PID peak points in the same borehole.
[0084] According to the MIP signal response value, after importing it into origin, the maximum values, i.e., peaks, within the regional range and their depths can be determined through the peak search function, and the distances between the maximum values can be calculated.
[0085] Step a2: Take the sampling points corresponding to several maximum values in the flame ionization detector signal value as the flame ionization detection peak points.
[0086] The flame ionization detection peak points are the FID peak points. When the FID signal value is higher than both the previous FID signal value and the next FID signal value, this FID signal value is the FID peak point, and there are multiple FID peak points in the same borehole.
[0087] When the concentration of volatile organic compounds in the borehole is 0 or too high, the concentration of volatile organic compounds tends to be average, that is, the PID and FID signal values of the sampling points tend to be stable, and there are no PID peak points or FID peak points in some boreholes.
[0088] Step a3: Determine the characteristic variables corresponding to the sampling points based on the photoionization detection peak points and the flame ionization detection peak points. The characteristic variables include: the signal value of the previous PID peak point of the sampling point, the signal value of the previous FID peak point of the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
[0089] Each sampling point has its own characteristic variables. There are four characteristic variables, namely: the signal value of the previous PID peak point of the sampling point, the signal value of the previous FID peak point of the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
[0090] In the vertical direction of each sampling point, the direction towards the ground is the front, and the direction towards the bottom of the ground is the back. When the distribution of volatile organic compounds in the borehole is uneven and the PID and FID signal values have large fluctuations, there are multiple PID peak points and FID peak points in front of the sampling point at this time. Generally speaking, there is no sampling point and no peak point in front of the first sampling point. Therefore, a sampling point has no characteristic variables. There may be multiple PID peak points and FID peak points behind other sampling points. The previous PID peak point or FID peak point refers to the peak point closest to the front of the sampling point, and the next PID peak point and FID peak point refer to the peak point closest to the back of the sampling point; when the distribution of volatile organic compounds in the borehole is even and the PID and FID signal values have no fluctuations, there is no PID peak point or FID peak point in this borehole. Therefore, some sampling points have no characteristic variables.
[0091] So far, the parameter data of each sampling point are eight input variables, namely four sampling data and four characteristic variables: PID signal value, FID signal value, soil conductivity, depth of the sampling point, signal value of the previous PID peak point of the sampling point, signal value of the previous FID peak point of the sampling point, distance between the sampling point and the previous PID peak point, and distance between the sampling point and the previous FID peak point.
[0092] Step S203: Train the machine learning model with the actual pollutant concentration value and parameter data of the soil sample to obtain a prediction model. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.
[0093] Step S204: Obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted volatile organic compound concentration value. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.
[0094] The method for predicting the concentration of volatile organic compounds provided in this embodiment uses MIP surveys to obtain direct detection signals of VOCs rather than indirect signals, and is not affected by underground pipelines, storage tanks, structures, landfill impurities, etc. The survey results are stable and reliable, and it is not restricted when used in chemical pollution sites where there are often complex underground water pipeline and storage tank structures.
[0095] By constructing a machine learning model, the problem of the lag effect of MIP caused by the residual VOCs in the pipeline is solved. There are two factors affecting the residual problem of VOCs in the pipeline: the concentration of VOCs in the high-concentration area passed by the probe and the time after passing through the high-concentration area. The present invention constructs four new input variables, namely feature variables, that is, the signal value of the previous PID peak point before the sampling point, the signal value of the previous FID peak point before the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point. Thus, the main influencing factors of the delay effect are considered in the machine learning model, the influence of the delay effect is alleviated to a certain extent, and the accuracy of the prediction result is further improved.
[0096] In this embodiment, a rapid soil VOCs pollution survey method based on MIP and machine learning is provided, which can be used for the above-mentioned mobile terminals, such as mobile phones, tablets, etc. Figure 3 It is a flowchart of a volatile organic compound concentration prediction method according to an embodiment of the present invention, as Figure 3 shown, and the process includes the following steps:
[0097] Step S301, obtain the actual pollutant concentration values of soil samples in several boreholes in the survey area. For details, please refer to Figure 1 step S101 of the embodiment shown, which will not be elaborated here.
[0098] Step S302, based on the membrane interface detection system, sample each borehole at different depths to obtain a number of sampling points and the parameter data corresponding to each sampling point. For details, please refer to Figure 1 step S102 of the embodiment shown, which will not be elaborated here.
[0099] Step S303, train the machine learning model by combining the actual pollutant concentration values of the soil samples and the parameter data to obtain a prediction model.
[0100] Specifically, the above step S303 includes:
[0101] Step S3031, split the actual pollutant concentration values of the soil samples and the parameter data into a training set and a test set according to a preset ratio.
[0102] Before splitting according to the preset ratio, first perform mean processing on the parameter data of the sampling points.
[0103] Taking 30 cm as a unit, calculate the mean value of each of the eight input variables of the sampling points included in every 30 cm in the vertical direction of the borehole as a data point, which is equivalent to normalizing the sampling points within 30 cm into one data point, and using this data point to represent the parameter data situation of the sampling points within 30 cm.
[0104] Split the data set composed of parameter data and the actual pollutant concentration values of soil samples. 75% of the data set is used as the training set for training, and 25% of the data set is used as the test set to evaluate the generalization effect of the model.
[0105] According to needs, the ratio of the training set to the test set can also be 80% and 20% or 60% and 40%, etc. Generally speaking, the number of the training set generally accounts for more than 60%.
[0106] Step S3032: Train the machine learning model with the training set to obtain the prediction model.
[0107] Input the training set into the machine learning model respectively, that is, train based on the Gradient Boosting Tree algorithm (Light GradientBoosting Machine, LightGBM), Random Forest algorithm, Generalized Regression Neural Network algorithm, and Linear Regression algorithm respectively to obtain the output results of the four algorithm models. According to actual needs, select two or more algorithm models with better performance for integration, and the integration result is the prediction model.
[0108] For example, using python programming, input the prepared data into any of the above four machine learning models, call the program package for model training. The input data can be selected in various file formats such as csv and excel. When the output result of the prediction model is the same as the actual pollutant concentration value of the soil sample, it means that the mapping relationship between the input data and the output result is obtained, indicating successful training.
[0109] Use the following linear equation to represent the output result of the linear regression algorithm model:
[0110]
[0111] Where represents the predicted target value, that is, the output result, w represents the weight parameter (Weight), b is the bias (Bias) or intercept (Intercept), and x is the input data. After solving for the unknown parameters w and b, it means that the trained linear regression algorithm model is obtained.
[0112] Four trained machine learning models are obtained respectively. Two or more trained machine learning models with better performance are selected for integration, and the integration result is the prediction model. Each trained machine learning model can be a prediction model. In order to achieve better prediction results, two or more trained machine learning models can also be selected for integration to obtain a new prediction model. In Python, scikit-learn is used to implement the stacking integration of the trained machine learning models. The stacking regressor function is used. Two or more trained machine learning models are selected as the base learners, and an integration algorithm is used as the meta-learner. The prediction results of each trained machine learning model are used as new input variables and input into the meta-learner for integrated learning. The meta-learner outputs the final prediction result of the entire integrated learning, and this prediction result is the prediction result of the prediction model. Among them, the integration algorithm of the meta-learner can adopt a linear regression algorithm or other simple algorithms.
[0113] Such as Figure 12 is the line chart comparing the predicted values and actual values of the four machine learning models in the embodiments of the present invention; combining the prediction results with the delay effect mitigation effect, where the dotted line of the line represented by Stacking is the predicted value obtained after the integration of the Linear algorithm model and the LightBGM algorithm model, and the solid line of the line represented by Measured is the measured value, that is, the true concentration value.
[0114] Step S3033, evaluate the effect of the prediction model based on the test set.
[0115] The test set is used to test the prediction performance of the prediction model. The prediction model composed of a single machine learning model can be evaluated, or the prediction model integrated by two or more machine learning models can be evaluated.
[0116] The coefficient of determination R 2 is used as an evaluation index to measure the goodness of fit of the prediction model. Such as Figure 8 is the fitting coefficient and scatter plot of the random forest algorithm model (RF) in the test set of the embodiments of the present invention; Figure 9 is the fitting coefficient and scatter plot of the gradient boosting tree algorithm model (LightGBM) in the test set of the embodiments of the present invention; Figure 10 is the fitting coefficient and scatter plot of the generalized regression neural network algorithm model (GRNN) in the test set of the embodiments of the present invention; Figure 11 is the fitting coefficient and scatter plot of the linear regression algorithm model (Linear) in the test set of the embodiments of the present invention; it can be seen that the LightGBM algorithm model and the GRNN algorithm model obtain a higher degree of fit, and the root mean square error (RMSE) is used as the error statistical index.
[0117] The coefficient of determination R is expressed by the following calculation formula 2 :
[0118]
[0119] The root mean square error (RMSE) is expressed by the following calculation formula:
[0120]
[0121] where Z(x i ) is the output result, i.e., the predicted volatile organic compound concentration value, and Z * (x i ) is the actual concentration value of the pollutant, is the mean value of the actual concentration values of the pollutants, and n is the number of samples.
[0122] Step S304: Obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted volatile organic compound concentration value. For details, please refer to step S104 of the embodiment shown in Figure 1 which will not be elaborated here.
[0123] The method for predicting the concentration of volatile organic compounds provided in this embodiment obtains a prediction model with better prediction effect through training and effect evaluation of a mature machine learning model. The prediction model takes into account influencing factors such as the detection signal delay effect of volatile organic compounds during the detection process, can alleviate the influence of the delay effect to a certain extent, and further improves the accuracy of the prediction result; calculating the mean value of the sampling points within the unit range to obtain data points can avoid accidental errors caused by abnormal value fluctuations and blank values in the sampling data.
[0124] In this embodiment, a rapid soil VOCs pollution survey method based on MIP and machine learning is provided, which can be used for the above-mentioned mobile terminals, such as mobile phones, tablets, etc., Figure 4 is a flowchart of the method for predicting the concentration of volatile organic compounds according to the embodiment of the present invention. As shown in Figure 4 the process includes the following steps:
[0125] Step S401: Obtain the actual pollutant concentration values of soil samples in several boreholes in the survey area. For details, please refer to step S101 of the embodiment shown in Figure 1 which will not be elaborated here.
[0126] Step S402: Based on the membrane interface detection system, perform sampling at different depths for each borehole to obtain several sampling points and the sampling data corresponding to each sampling point. For details, please refer to step S102 of the embodiment shown in Figure 1 which will not be elaborated here.
[0127] Step S403: Train the machine learning model with the actual pollutant concentration values and sampling data of the soil sample to obtain a prediction model. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.
[0128] Step S404: Obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted volatile organic compound concentration value.
[0129] Specifically, the above Step S404 includes:
[0130] Step S4041: Obtain several prediction points based on the prediction area, obtain the detection data of the prediction points based on the membrane interface detection system, and obtain the prediction feature variables corresponding to the prediction points based on the detection data.
[0131] Obtain the detection data of the prediction points based on the membrane interface detection system: The prediction area refers to the surveyed area where actual drilling sampling has not been carried out. Use the vehicle-mounted MIP system to sample at different depths of the prediction points in the prediction area. Insert the MIP probe into the soil of the prediction area to obtain the detection data of the prediction points. The working method of the vehicle-mounted MIP system is the same as that for obtaining parameter data. The detection data includes PID signal value, FID signal value, soil conductivity, and the depth of the prediction point.
[0132] Obtain the prediction feature variables corresponding to the prediction points based on the detection data: The prediction feature variables obtain the prediction feature variables corresponding to the detection points based on the detection data of the prediction points. The prediction feature variables include: the signal value of the previous PID peak point of the prediction point, the signal value of the previous FID peak point of the prediction point, the distance between the prediction point and the previous PID peak point, and the distance between the prediction point and the previous FID peak point. Similar to the sampling points, the first prediction point has no prediction feature variables, and some prediction points have no feature variables.
[0133] Step S4042: Input the detection data and prediction feature variables of the prediction points into the prediction model to predict the concentration of volatile organic compounds, and obtain the predicted volatile organic compound concentration value of the prediction points.
[0134] So far, eight input variables of the prediction points are obtained: PID signal value, FID signal value, soil conductivity, depth of the prediction point, the signal value of the previous PID peak point of the prediction point, the signal value of the previous FID peak point of the prediction point, the distance between the prediction point and the previous PID peak point, and the distance between the prediction point and the previous FID peak point.
[0135] Input the eight input variables of the prediction points into the prediction model to predict the concentration of volatile organic compounds. The output result of the prediction model is the predicted volatile organic compound concentration value of the prediction points.
[0136] Step S4043: Visualize the three-dimensional spatial distribution of volatile organic compound (VOC) concentrations based on the VOC concentration values at the prediction points and the detection data.
[0137] In some alternative embodiments, the above step S4043 includes:
[0138] Step b1: Input the VOC concentration values at the prediction points, the longitude, latitude, and depth of the prediction points detected by the membrane interface detection system into the Earth Volumetric Studio (EVS) three-dimensional visualization software. Combine with the formation lithology for three-dimensional spatial interpolation to obtain the visualization of the three-dimensional spatial distribution of VOC concentrations.
[0139] The prediction points are discrete, and the obtained VOC concentration values at the prediction points are also discrete. However, the three-dimensional spatial distribution visualization is continuous. Using a program software or programming as a tool, the method of three-dimensional spatial interpolation is adopted to achieve the change from discrete to continuous for the prediction points.
[0140] The VOC concentration prediction method provided in this embodiment saves the time for actual exploration of soil VOC concentrations by predicting the detection data, accurately depicts the three-dimensional spatial distribution of pollutant concentrations through three-dimensional spatial interpolation, and can provide auxiliary support for pollution source analysis, migration process judgment, etc.
[0141] As one or more specific application embodiments of the present invention, the rapid soil VOCs pollution investigation method based on MIP and machine learning includes the following steps:
[0142] 1. Obtain the actual pollutant concentration values of soil samples in several boreholes in the investigation area.
[0143] Select a typical chemical industrial site as the investigation area. According to the analysis of pollution history and hydrogeological data, combined with the topographic map of the investigation area and the on-site inspection situation, and according to the direction of pollutant migration with groundwater in the case investigation area, a total of 6 survey lines and 28 MIP monitoring points are set. The direction of the survey lines is perpendicular to the groundwater flow direction, and the spacing between the survey line measurement points is controlled within the range of 10 to 20 meters; Soil sampling is carried out beside 9 of the MIP test points, geological data and spatial information are recorded, and the detected concentration of total petroleum hydrocarbons (TPH) in the samples is obtained, that is, the true concentration value of pollutants in the soil samples.
[0144] 2. Based on the membrane interface detection system, sample at different depths for each borehole to obtain a number of sampling points and the corresponding parameter data for each sampling point.
[0145] Collect sampling data through rapid MIP surveys, including soil electrical conductivity (EC), MIP signals (i.e., PID and FID signals), and the depth of sampling points. Process the above four types of sampling data to obtain four characteristic variables, which include: the signal value of the previous PID peak point before the sampling point, the signal value of the previous FID peak point before the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point. Each sampling point includes four sampling data and four characteristic variables, and the four sampling data and four characteristic variables constitute parameter data.
[0146] 3. Train a machine learning model using the actual pollutant concentration values of soil samples and parameter data to obtain a prediction model.
[0147] The parameter data of all sampling points and the true pollutant concentration values of soil samples form a data set. This data set is split, with 75% of the data set as the training set and 25% as the test set.
[0148] Put the training set into the machine learning model for training to obtain the mapping relationship between the model output results and the training set, and obtain the trained machine learning model. Use the test set to test the trained machine learning model. According to the test results, the prediction effect of the trained machine learning model can be known. At this time, a single trained machine learning model can be used as the prediction model, or an ensemble of two or more trained machine learning models with better prediction effects can be used as the prediction model to achieve a more accurate prediction effect.
[0149] 4. Obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted concentration value of volatile organic compounds.
[0150] After obtaining the prediction model, the detection data detected by MIP can be directly input into the prediction model to obtain the predicted concentration value of volatile organic compounds, without the need to drill soil samples to obtain pollutant concentration values.
[0151] Use a vehicle-mounted MIP system to obtain detection data for the detection area, and obtain detection data for corresponding prediction points by inserting the MIP probe into the soil.
[0152] 5. Visualize the three-dimensional spatial distribution of volatile organic compound concentrations based on the predicted concentration values of volatile organic compounds at the prediction points and the detection data.
[0153] Combined with the formation lithology, input the predicted concentration value of volatile organic compounds at the prediction point and the three-dimensional coordinates of the prediction point into three-dimensional visualization software to depict the three-dimensional spatial distribution of volatile organic compound concentrations in different formation lithologies. The three-dimensional coordinates include longitude, latitude, and the depth of the prediction point; such as Figure 13It is a three-dimensional spatial distribution map of the volatile organic compound concentration values predicted by the gradient boosting tree algorithm model (LightGBM) according to the embodiments of the present invention; Figure 14 It is a three-dimensional spatial distribution map of the volatile organic compound concentration values predicted by the generalized regression neural network algorithm model (GRNN) according to the embodiments of the present invention; Figure 15 It is a three-dimensional spatial distribution map of the volatile organic compound concentration values predicted by the linear regression algorithm model (Linear) according to the embodiments of the present invention; Figure 16 It is a three-dimensional spatial distribution map of the volatile organic compound concentration values predicted by the random forest algorithm model (RF) according to the embodiments of the present invention.
[0154] In this embodiment, a volatile organic compound concentration prediction device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0155] This embodiment provides a volatile organic compound concentration prediction device, as Figure 5 shown, including:
[0156] A membrane interface detection module 501, configured to obtain the actual pollutant concentration values of soil samples in several boreholes in the investigation area;
[0157] A model construction module 502, configured to perform sampling at different depths for each borehole based on the membrane interface detection system, and obtain several sampling points and the parameter data corresponding to each sampling point;
[0158] A model training module 503, configured to train a machine learning model by combining the actual pollutant concentration values of soil samples and the parameter data to obtain a prediction model;
[0159] A concentration value prediction module 504, configured to obtain detection data based on the membrane interface detection system, and input the detection data into the prediction model to obtain the predicted volatile organic compound concentration values.
[0160] In some alternative implementation manners, the model construction module 502 includes:
[0161] A sampling data unit, configured to perform sampling at different depths for each borehole based on the membrane interface detection system, and obtain several sampling points and the sampling data corresponding to each sampling point. The sampling data includes the photoionization detector signal value, the flame ionization detector signal value, the soil conductivity, and the depth of the sampling point.
[0162] A feature variable unit for obtaining a feature variable corresponding to a sampling point based on sampling data of the sampling point, where the sampling data and the feature variable constitute parameter data.
[0163] In some alternative embodiments, the feature variable unit includes:
[0164] A peak acquisition subunit for taking the sampling points corresponding to several maximum values in the photoionization detector signal value as photoionization detection peak points; and taking the sampling points corresponding to several maximum values in the flame ionization detector signal value as flame ionization detection peak points.
[0165] A feature variable acquisition subunit for determining a feature variable corresponding to a sampling point based on the photoionization detection peak point and the flame ionization detection peak point, where the feature variable includes: the signal value of the previous PID peak point of the sampling point, the signal value of the previous FID peak point of the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
[0166] In some alternative embodiments, the model training module 503 includes:
[0167] A data splitting unit for splitting the actual pollutant concentration value and parameter data of the soil sample into a training set and a test set according to a preset ratio.
[0168] A model construction unit for training a machine learning model using the training set to obtain a prediction model.
[0169] An evaluation model unit for evaluating the effect of the prediction model using the test set.
[0170] In some alternative embodiments, the concentration value prediction module 504 includes:
[0171] A detection unit for obtaining several prediction points based on a prediction area, obtaining detection data of the prediction points based on a membrane interface detection system, and obtaining a prediction feature variable corresponding to the prediction points based on the detection data.
[0172] A calculation unit for inputting the detection data and the prediction feature variable of the prediction points into the prediction model to predict the concentration of volatile organic compounds, and obtaining the predicted concentration value of volatile organic compounds of the prediction points.
[0173] A visualization unit for visualizing the three-dimensional spatial distribution of the concentration of volatile organic compounds based on the concentration value of volatile organic compounds of the prediction points and the detection data.
[0174] In some alternative embodiments, the visualization unit includes: inputting the volatile organic compound concentration value of the prediction point, the longitude, latitude, and depth of the prediction point detected by the membrane interface detection system into three-dimensional visualization software, and performing three-dimensional spatial interpolation in combination with the formation lithology to obtain the three-dimensional spatial distribution visualization of the volatile organic compound concentration.
[0175] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0176] The volatile organic compound concentration prediction device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0177] An embodiment of the present invention further provides a computer device having the above-mentioned Figure 5 volatile organic compound concentration prediction device shown.
[0178] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 6 , the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In
[0179] FIG. 1, one processor 10 is taken as an example.
[0180] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0181] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0182] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.
[0183] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure X Taking connection through a bus as an example.
[0184] The input device 30 may receive input digital or character information, and generate key signal inputs related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0185] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0186] A part of the present invention can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0187] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A rapid investigation method for soil VOCs pollution based on MIP and machine learning, characterized in that: The method comprises: Obtain actual pollutant concentration values of soil samples from several boreholes in the survey area; Based on the membrane interface detection system, each borehole is sampled at different depths to obtain several sampling points and parameter data corresponding to each sampling point; The machine learning model is trained by combining the actual pollutant concentration values and parameter data of the soil samples to obtain a prediction model; Acquiring detection data based on the membrane interface detection system, inputting the detection data into the prediction model, and obtaining a predicted volatile organic compound concentration value; The membrane interface detection system includes a photoionization detector and a flame ionization detector. Based on the membrane interface detection system, each borehole is sampled at different depths to obtain a number of sampling points and parameter data corresponding to each sampling point, including: Based on the membrane interface detection system, each borehole is sampled at different depths to obtain a number of sampling points and sampling data corresponding to each sampling point, wherein the sampling data includes a photoionization detector signal value, a flame ionization detector signal value, soil conductivity, and the depth of the sampling point; Acquire characteristic variables corresponding to the sampling points based on sampling data of the sampling points, wherein the sampling data and the characteristic variables constitute parameter data; The characteristic variables corresponding to the sampling points are obtained based on the sampling data of the sampling points, including: The sampling points corresponding to the several maximum values in the signal value of the photoionization detector are taken as the peak points of the photoionization detection; The sampling points corresponding to the several maximum values in the flame ionization detector signal value are taken as the flame ionization detection peak points; The characteristic variables corresponding to the sampling point are determined based on the photoionization detection peak point and the flame ionization detection peak point, and the characteristic variables include: the signal value of a PID peak point before the sampling point, the signal value of a FID peak point before the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
2. The method according to claim 1, characterized in that The machine learning model is trained by combining the actual pollutant concentration values and parameter data of the soil samples to obtain a prediction model, including: The actual pollutant concentration values and parameter data of the soil samples are divided into a training set and a test set according to a preset ratio; Use the training set to train the machine learning model to obtain a prediction model; The prediction model is evaluated based on the test set.
3. The method according to claim 1, characterized in that Acquiring detection data based on the membrane interface detection system, inputting the detection data into the prediction model, and obtaining a predicted volatile organic compound concentration value, including: Acquire a number of prediction points based on the prediction area, acquire detection data of the prediction points based on the membrane interface detection system, and acquire prediction characteristic variables corresponding to the prediction points based on the detection data; The detection data and predicted characteristic variables of the prediction point are input into the prediction model to predict the concentration of volatile organic compounds, and the predicted volatile organic compound concentration value of the prediction point is obtained.
4. The method according to claim 3, characterized in that Based on the membrane interface detection system, the detection data is obtained and input into the prediction model to obtain the predicted volatile organic compound concentration value, which also includes: Visualize the three-dimensional spatial distribution of VOC concentration based on the VOC concentration values at the prediction points and the detection data; The volatile organic compound concentration value of the predicted point and the longitude, latitude and depth of the predicted point detected by the membrane interface detection system are input into the three-dimensional visualization software, and three-dimensional spatial interpolation is performed in combination with the stratum lithology to obtain the three-dimensional spatial distribution visualization of the volatile organic compound concentration.
5. A volatile organic compound concentration prediction device, characterized in that: The device comprises: The membrane interface detection module is used to obtain the actual pollutant concentration values of soil samples in several boreholes in the survey area; The model building module samples each borehole at different depths based on the membrane interface detection system, and obtains a number of sampling points and the sampling data corresponding to each sampling point; A model training module is used to train the machine learning model by combining the actual pollutant concentration values of the soil samples and the sampling data to obtain a prediction model; A concentration value prediction module, used to obtain detection data based on the membrane interface detection system, input the detection data into the prediction model, and obtain a predicted volatile organic compound concentration value; The membrane interface detection system includes a photoionization detector and a flame ionization detector. Based on the membrane interface detection system, each borehole is sampled at different depths to obtain a number of sampling points and parameter data corresponding to each sampling point, including: Based on the membrane interface detection system, each borehole is sampled at different depths to obtain a number of sampling points and sampling data corresponding to each sampling point, wherein the sampling data includes a photoionization detector signal value, a flame ionization detector signal value, soil conductivity, and the depth of the sampling point; Acquire characteristic variables corresponding to the sampling points based on sampling data of the sampling points, wherein the sampling data and the characteristic variables constitute parameter data; The characteristic variables corresponding to the sampling points are obtained based on the sampling data of the sampling points, including: The sampling points corresponding to the several maximum values in the signal value of the photoionization detector are taken as the peak points of the photoionization detection; The sampling points corresponding to the several maximum values in the flame ionization detector signal value are taken as the flame ionization detection peak points; The characteristic variables corresponding to the sampling point are determined based on the photoionization detection peak point and the flame ionization detection peak point, and the characteristic variables include: the signal value of a PID peak point before the sampling point, the signal value of a FID peak point before the sampling point, the distance between the sampling point and the previous PID peak point, and the distance between the sampling point and the previous FID peak point.
6. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the soil VOCs pollution rapid investigation method based on MIP and machine learning as described in any one of claims 1 to 4 by executing the computer instructions.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the soil VOCs pollution rapid investigation method based on MIP and machine learning as described in any one of claims 1 to 4.
8. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the soil VOCs pollution rapid investigation method based on MIP and machine learning as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Resistivity parameter and machine learning soil organic pollution concentration prediction method
CN116738227A