Method for detecting concentration of VOCs concentration correlation prediction model in water-soil-gas based on random forest algorithm

By adopting the particle swarm optimization random forest algorithm in mining engineering, the VOCs gas concentration intelligent prediction method is solved to solve the problems of high cost and low accuracy of VOCs detection in mining engineering, and the efficient and accurate monitoring and early warning of VOCs in water, soil and air environments are realized, thus reducing the risk of occupational diseases.

CN120636618APending Publication Date: 2025-09-12ANHUI UNIV OF SCI & TECH +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510735742.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing VOCs detection methods in mining engineering have the problems of high cost, low precision, difficulty in achieving real-time monitoring and inaccurate predictions in multi-media environments. In particular, VOCs pollution monitoring in water, soil and atmosphere is difficult to achieve efficient and accurate early warning and control.

Method used

A portable, low-power online monitoring system was designed using an intelligent prediction method for VOCs gas concentration based on a particle swarm optimization random forest algorithm, combined with PID technology. By constructing a random forest prediction model and introducing a particle swarm optimization algorithm to automatically optimize the key parameters of the model, accurate monitoring and prediction of VOCs concentrations in water, soil, and air environments can be achieved.

Benefits of technology

It significantly improves the accuracy and efficiency of VOCs concentration prediction, reduces monitoring costs, is suitable for pollution warning and responsibility identification in complex environmental systems, supports real-time monitoring and emergency response, and reduces the risk of occupational diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005433226970000081
    Figure BDA0005433226970000081
  • Figure BDA0005433226970000091
    Figure BDA0005433226970000091
  • Figure BDA0005433226970000101
    Figure BDA0005433226970000101
Patent Text Reader

Abstract

The invention discloses a random forest algorithm-based water-soil-gas VOCs concentration correlation prediction model concentration detection method, which is characterized in that a random forest prediction model is constructed, and a particle swarm optimization algorithm is innovatively introduced to automatically adjust and optimize key parameters of the model, so that the prediction precision and the model performance are remarkably improved; according to the method, the optimal parameter combination of 28 particle numbers, 0.9 inertia weight, 11 iterations and the like is set, so that the problems that the RF model is easy to over-fit and the training speed is low are effectively solved, a high-precision and high-efficiency intelligent prediction solution is provided for VOCs gas concentration monitoring, and the method is suitable for popularization and application. The method can be widely applied to the fields of environmental monitoring, industrial process control and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of VOCs detection, and in particular to a concentration detection method of a VOCs concentration correlation prediction model in water, soil and air based on a random forest algorithm. Background Art

[0002] Volatile organic compounds (VOCs) are a class of organic chemicals that are easily volatile at room temperature and pressure. They are widely present in various fields such as industrial production, transportation, and daily life. Due to their toxicity, carcinogenicity, and environmental persistence, VOCs pollution has become a global environmental problem. With the acceleration of my country's industrialization process, VOCs emissions have continued to increase, causing serious complex pollution to water bodies, soil, and the atmospheric environment. Traditional VOCs detection methods such as gas chromatography-mass spectrometry (GC-MS) have high accuracy, but have limitations such as high cost, long analysis cycle, and difficulty in achieving large-scale real-time monitoring. Therefore, the development of an efficient and accurate VOCs concentration prediction model is of great practical significance.

[0003] VOCs in mining engineering exist throughout the entire life cycle of mining, ore processing and tailings disposal. VOCs emissions are multi-source and complex. They are produced through diesel equipment exhaust, blasting reactions, chemical volatilization and natural release of ores during mining, transportation, mineral processing and smelting. For example, diesel-powered equipment (such as mining machines and transport trucks) releases benzene series (BTEX), polycyclic aromatic hydrocarbons (PAHs) and aldehydes (such as formaldehyde and acrolein) during high-temperature combustion; organic solvents (such as kerosene and pine oil) volatilize in mineral processing processes such as flotation and leaching, and the secondary release of VOCs adsorbed by dust during ore crushing; the oxidative decomposition of sulfide minerals and the anaerobic degradation of organic matter in tailings ponds produce characteristic VOCs such as mercaptans and sulfides; chlorinated hydrocarbons and ester compounds volatilized from equipment lubricants, hydraulic oil leakage and anti-corrosion coatings. The types and concentrations of VOCs are significantly affected by the type of ore, mining method (open-pit / underground), and environmental conditions (temperature, humidity), and are present in multiple media including water, soil, and air. As precursors of ozone (O3) and secondary organic aerosols (SOA), VOCs are important contributors to haze and photochemical smog. Long-term exposure to VOCs can lead to respiratory diseases, neurological damage, and cancer risks (for example, benzene is classified as a Class I carcinogen by the International Agency for Research on Cancer). VOCs can also enter soil and water bodies through dry and wet deposition, inhibiting microbial activity and accumulating along the food chain, leading to the degradation of ecological functions in mining areas. Therefore, strengthening the monitoring of VOCs in mines, clarifying their sources, and taking timely measures to control them are crucial. This has important theoretical value and practical significance for achieving green mine construction and sustainable development goals.

[0004] The existing VOCs concentration detection methods have the following main disadvantages:

[0005] Fourier Transform Infrared Spectrometer: FTIR technology relies on a Fourier Transform Infrared Spectrometer (FTIR) and can be used in both open and closed modes. The core principle of this technology for detecting VOCs (volatile organic compounds) lies in the principle of light absorption: by allowing a light source to penetrate a container containing a specific concentration of VOCs, the VOC molecules selectively absorb light of specific wavelengths, resulting in a decrease in light intensity. This decrease in light intensity directly reflects the concentration and content of the VOCs. However, FTIR technology has some drawbacks, such as high equipment cost, high maintenance costs, weak anti-interference capabilities, and a large footprint.

[0006] Metal Oxide Semiconductor Sensors: Metal oxide semiconductor sensors are composed of oxides such as CuO, WO3, ZnO, NiO, and SnO2. Their operating principle is based on the interaction between VOCs and the charge carriers of the metal oxide semiconductor on the sensor surface, which in turn causes a change in sensor resistance. These sensors offer advantages such as low cost, small size, long service life, and ease of integration. However, they also have disadvantages such as low selectivity, susceptibility to aging, and sensitivity to temperature and humidity.

[0007] Hydrogen flame ionization sensor: The hydrogen flame ionization sensor uses the flame generated by hydrogen combustion as an ionization source to ionize VOCs gas, causing the substance to be tested to undergo chemical ionization, forming an ion flow, and by detecting the size of the ion flow, quantitative detection of VOCs is achieved. Therefore, an external hydrogen source is required during use. The hydrogen flame ionization detector (FID) for monitoring VOCs has the advantages of high sensitivity, fast response, wide linear range and good stability, and is suitable for continuous monitoring. The disadvantages are poor selectivity, inability to distinguish between different VOCs, and dependence on hydrogen supply, which poses a safety hazard. At the same time, it cannot detect inorganic gases and has high operating costs. Overall, FID is suitable for high-sensitivity monitoring, but it needs to be used in conjunction with other technologies to improve selectivity.

[0008] Electrochemical sensor: Electrochemical sensing technology detects the target gas through the electrochemical reaction cell inside the sensor; in the reaction cell, the gas to be measured undergoes an oxidation-reduction reaction, and then quantitative detection is achieved by measuring the current at the oxidation electrode and the reduction electrode. Electrochemical sensors for monitoring VOCs have the advantages of high sensitivity, fast response speed, low power consumption, and small size, making them suitable for portable devices. However, they have poor selectivity, are susceptible to cross-interference, and lack long-term stability, and their performance may degrade due to electrode aging or contamination. In addition, changes in ambient temperature and humidity can also affect detection accuracy. Overall, electrochemical sensors are suitable for real-time monitoring, but require regular calibration and maintenance.

[0009] Gas chromatography: Gas chromatography is a chromatographic column separation technology that uses gas as the mobile phase and requires the aid of a carrier gas for detection and analysis. The principle is to use a carrier gas to deliver the gas to be tested into the chromatographic column, and the gas to be tested reacts with the stationary phase. Due to differences in the time it takes for different gases to be adsorbed by the adsorbent or dissolved in the stationary liquid, their residence time in the chromatographic column is different, and thus the order in which the components flow out is also different. By analyzing the chromatogram, the specific composition of the gas to be tested can be obtained; however, due to the large variety and different properties of the substances to be tested in VOCs, gas chromatography technology is difficult to achieve synchronous detection; in addition, gas chromatography is a point measurement technology with limited sampling volume, and it is impossible to continuously monitor the concentration of the gas to be tested in real time. Moreover, gas chromatography instruments are expensive and difficult to design as portable devices, making them more suitable for use in laboratories. Summary of the Invention

[0010] In response to the problems existing in the above-mentioned prior art, the present invention provides a method for intelligent prediction of VOCs gas concentration based on particle swarm optimization random forest (PSO-RF). Taking the VOCs that mainly exist in the field of mining engineering as the research object, combined with PID technology, a portable, low-power VOCs online monitoring system is developed to achieve effective monitoring of VOCs in water, soil and air environments, and reveal the dynamic change law of the VOCs release process; combined with machine learning algorithms, VOCs concentration prediction is achieved. Through real-time monitoring and prediction, the VOCs exposure risk of miners is reduced, the incidence of occupational diseases is reduced, and the concept of "humanistic mining" is promoted. By constructing a random forest prediction model and innovatively introducing a particle swarm optimization algorithm to automatically tune the key parameters of the model, the prediction accuracy and model performance are significantly improved; a series of experiments are designed to verify the effectiveness of this method. Finally, through the analysis and discussion of the experimental results, the concentration of VOCs gas was successfully detected in a water-soil-air environment. The experimental results showed that the RF model optimized by PSO (PSO-RF) performed best in all evaluation indicators. This method effectively solved the problems of easy overfitting and slow training speed of the RF model by setting the optimal parameter combination of 28 particles, 0.9 inertia weight and 11 iterations, providing a high-precision and high-efficiency intelligent prediction solution for VOCs gas concentration monitoring, which can be widely used in environmental monitoring, industrial process control and other fields.

[0011] The photoionization detector first guides the gas to be measured into an ionization chamber. A vacuum ultraviolet light source then releases photons of a specific energy to excite VOC gas molecules, causing them to ionize and form ions and electrons with separate positive and negative charges. Under the influence of the electric field set within the ionization chamber, these ions and electrons migrate toward the opposite poles of the electric field, inducing a weak current signal. This current signal is accurately measured and calibrated by constructing a VOC gas concentration prediction model based on a random forest algorithm to predict the changing trend of VOC gas concentration. A particle swarm optimization (PSO) algorithm is then introduced to optimize the random forest model, constructing a PSO-RF model to infer VOC gas concentration.

[0012] By innovatively integrating machine learning with environmental mechanism models, accurate monitoring and intelligent early warning of VOCs pollution in complex environmental systems have been achieved. This technical solution is mainly suitable for multi-media environmental monitoring of typical polluted sites such as industrial areas and chemical parks, covering typical VOCs pollutants such as benzene series and halogenated hydrocarbons in the three major media of water (surface water, groundwater), soil (topsoil, deep soil) and atmosphere (near ground, factory boundary air). The innovation of this patent is to construct a physically constrained random forest algorithm architecture. Through medium coupling feature engineering and multi-source data fusion technology, it solves the technical bottlenecks of traditional methods in cross-media prediction, nonlinear modeling and data heterogeneity. It can be widely used in environmental protection supervision scenarios such as pollution early warning, responsibility identification, and remediation assessment. Compared with traditional detection methods, it achieves improved prediction accuracy and reduces monitoring costs, providing intelligent technical support for ecological environmental governance.

[0013] To achieve the above objectives, the present invention adopts a technical solution: a concentration detection method for VOCs concentration in water, soil and air based on a random forest algorithm and a correlation prediction model, comprising the following contents:

[0014] (1) Input VOCs gas concentration data, preprocess the data and divide the data set into training set and test set;

[0015] (2) Selecting eigenvalues ​​for the random forest model;

[0016] (3) Setting the initial velocity and coordinates of the particles, and using the particle swarm algorithm to perform cyclic optimization to adjust the moving speed and position of each particle in the particle swarm;

[0017] (4) Set the initial parameters of the random forest, build a decision tree based on it, and then train the random forest; evaluate the effectiveness and accuracy of the model through evaluation indicators and calculate its error rate;

[0018] (5) When the error rate of the model cannot meet the requirements, return to step (3) and continue to optimize the model through the particle swarm optimization algorithm to find the optimal parameters; once satisfactory model parameters are found or other termination conditions are met, the optimization process ends.

[0019] The beneficial effects of the present invention are:

[0020] 1. Significantly Improved Forecast Accuracy

[0021] The random forest model optimized by particle swarm optimization (PSO-RF) achieved a breakthrough prediction accuracy, with a determination coefficient R 2 As high as 0.99973, compared with the traditional RF model (R 2 =0.98824) increased by 1.15 percentage points; key error indicators were comprehensively optimized: the mean square error (MSE) dropped from 2150.1018 to 96.1486, a decrease of 95.53%; the mean absolute error (MAE) dropped from 38.1817 to 7.6165, a decrease of 80.05%. This level of accuracy significantly outperforms traditional methods such as SVR and BP neural networks.

[0022] 2. Comprehensive Optimization of Model Performance

[0023] The innovative use of the particle swarm algorithm (PSO) to intelligently optimize the key parameters of the RF model effectively addresses the overfitting and slow training speed issues of traditional RF models by achieving the optimal parameter combination of 28 particles, 0.9 inertia weight, and 11 iterations. Experiments show that the optimized model improves training efficiency by over 40% while maintaining improved generalization capabilities.

[0024] 3. Multi-dimensional technical advantages

[0025] High degree of intelligence: Automatically completes parameter optimization, reducing the workload of manual parameter adjustment; Strong adaptability: Can handle complex nonlinear relationships and is suitable for VOCs monitoring in different scenarios; Good stability: Strong robustness to data noise and outliers; Strong interpretability: Retains the feature importance analysis function of random forest;

[0026] 4. Outstanding application value

[0027] The present invention demonstrates excellent adaptability to uranium tailings of varying properties. Regardless of the uranium content or the presence of multiple heavy metals, efficient solidification of the uranium tailings can be achieved by adjusting parameters such as the CaCl2 dosage and curing time. This wide range of applicability provides an effective solution for the treatment of various uranium tailings, broadening its scope of application, increasing its practical value in the field of uranium tailings treatment, and reducing the difficulty and cost of treatment due to differences in tailings properties.

[0028] V. Economic and Environmental Benefits

[0029] Provide environmental regulatory authorities with high-precision pollution early warning tools; help companies achieve precise control of VOCs in production processes; significantly reduce monitoring costs, saving more than 50% compared to traditional methods; support real-time monitoring needs, and respond quickly enough for emergencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 The effect diagram of the VOCs concentration quantitative analysis model established by traditional support vector machine regression;

[0031] Figure 2 This is the effect diagram of the support vector machine regression VOCs concentration quantitative analysis model based on signal time domain and frequency domain feature extraction;

[0032] Figure 3 This is the effect diagram of the support vector machine regression VOCs concentration quantitative analysis model based on principal component analysis of signal time domain and frequency domain features;

[0033] Figure 4 This is a graph showing the trend of prediction error changing with evolutionary generations;

[0034] Figure 5 This is the effect diagram of the VOCs concentration quantitative analysis model based on signal time domain and frequency domain feature extraction, principal component analysis, and genetic algorithm optimized support vector machine regression;

[0035] Figure 6 The accuracy graph of the PCA-GA-SVM model under different training samples;

[0036] Figure 7 This is a flowchart for VOCs gas concentration prediction based on random forest algorithm;

[0037] Figure 8 This is the initial adjustment diagram for the number of random forest decision trees;

[0038] Figure 9 This is the quadratic adjustment graph for the number of random forest decision trees;

[0039] Figure 10 The effect diagram of the VOCs concentration quantitative analysis model established by random forest regression;

[0040] Figure 11 The effect diagram of the VOCs concentration quantitative analysis model established by BP neural network;

[0041] Figure 12 The effect diagram of the VOCs concentration quantitative analysis model established for the long short-term memory network;

[0042] Figure 13 The effect diagram of the VOCs concentration quantitative analysis model established by radial basis function neural network;

[0043] Figure 14 The effect diagram of the VOCs concentration quantitative analysis model established by the extreme learning machine;

[0044] Figure 15 Comparison chart of evaluation indicators of different models;

[0045] Figure 16 This is the flow chart of particle swarm algorithm;

[0046] Figure 17 This is the flow chart of the VOCs gas concentration prediction model based on PSO-RF;

[0047] Figure 18 This is a graph showing the changes in model evaluation indicators for different numbers of particles;

[0048] Figure 19 It is a graph of the iterative convergence process of the particle swarm algorithm;

[0049] Figure 20 This is a graph showing the changes in model evaluation indicators for different numbers of particles. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention pertains. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0052] Implementation of Support Vector Machine Regression Algorithm

[0053] Because support vector machine regression can be used to construct regression models for small sample data, a regression model was constructed using support vector machines for PID response to VOC concentration data. When using support vector machine regression for quantitative analysis of VOC concentrations, 84 different sets of VOC concentration data were first randomly arranged. 67 sets of data were used as training sets, and 17 sets of data were used as test sets. Each set of data contained 6001 values, of which the first 6000 values ​​were the PID response values ​​to VOC concentrations, and the 6001th value was the actual VOC concentration.

[0054] First, the training and test sets must be normalized to eliminate the negative effects of singular sample data. For SVMs, if data normalization is not used to make features dimensionless, the feature distribution will be elliptical, affecting model prediction accuracy during training and even causing the model training to enter an infinite loop and fail to converge.

[0055] Then set the support vector machine regression parameters and select the Gaussian radial basis kernel function as the kernel function of SVM regression. Try the effects of different kernel function parameters δ and penalty factor C on the SVM regression accuracy. Since there is no prior knowledge, first set the penalty factor C to 0.1 and the kernel function parameter δ to 1. At this time, the R 2 When R is 0.21 and C is 0.2 2 When R is 0.13 and C is 0.3 2 is 0.015, the model accuracy is not high, and C is set to 1 again. At this time, R 2 When R is 0.30 and C is 2 2 When R is 0.30 and C is 3 2 When R is 0.30 and C is 4 2 When C is 5, R 2 It is 0.29. Since it is difficult to find a mathematical relationship between the value of C and the model regression accuracy, the penalty factor C of SVM regression is temporarily set to 4. Keeping C unchanged, the kernel function parameter δ is obtained to be 0.8 using the same method.

[0056] Finally, the traditional support vector machine regression method was used to construct a model for PID signals under 84 groups of VOCs with different concentrations when the penalty factor C was 4 and the kernel function parameter δ was 0.8; the test set results of the model are shown in Figure 1 As shown in the figure. The model's mean square error is 69458.6, and its R² is 0.38. The model's predicted VOC concentration differs significantly from the actual VOC concentration. Analysis of the reasons for this large error reveals that the first reason is the uncertainty of the penalty factor C and the kernel function parameter δ, making it difficult to manually determine the optimal values ​​for these two parameters. The second reason is that the raw data used to train the model is redundant, meaning useful information cannot be well identified, while useless information can affect the accuracy of the model. Therefore, this study subsequently optimized the model from these two aspects.

[0057] PID signal feature extraction

[0058] The sampling interval of PID monitoring for VOCs is 0.02s / time, and each of the 84 groups of VOCs data with different concentrations consists of 6000 sampling points. Figure 1As can be seen, each set of data consisting of 6000 sampling points contains a lot of redundant information, which causes the SVM regression process to take a long time and the SVM regression model has low accuracy. Therefore, it is very necessary to extract the characteristic parameters of the data generated by PID.

[0059] The collected PID signal data contains a wealth of useful information reflecting VOC concentrations. These features are often expressed as time-domain waveform signals. However, in practice, using time-domain waveforms cannot accurately represent the VOC concentrations reflected by PID. If noise interferes with the PID signal acquisition process, the accuracy of determining VOC concentrations using time-domain characteristic parameters will be compromised. Therefore, in order to more comprehensively extract the characteristic parameters of the PID response to VOCs, it is necessary to extract characteristic parameters from multiple angles of the PID signal. PID signal feature parameters are extracted from both the time and frequency domains, and VOC concentrations are determined based on the time and frequency domain features of the PID signal. Experimental results show that the combined time-domain and frequency-domain feature extraction method can effectively improve the accuracy of support vector machine regression quantitative analysis of VOC concentrations.

[0060] (1) PID signal time domain feature extraction

[0061] The signal time-domain characteristic parameter analysis method determines VOC concentrations by analyzing the waveform information of the PID signal. As one of the oldest feature analysis methods, time-domain characteristic parameter analysis is highly mature, widely applicable, and straightforward, making it widely used for signal feature extraction. To extract the characteristic quantities of the PID response signal at different VOC concentrations, the amplitude analysis method within the time-domain analysis method is used to obtain the time-domain characteristic quantities of the PID signal.

[0062] The time domain characteristics of the signal are represented by 12 characteristic parameters, including the mean value, standard deviation, skewness, kurtosis, maximum value, minimum value, peak-to-peak value, root mean square, amplitude factor, waveform factor, impact factor, and margin factor of the signal generated by PID at different VOCs concentrations.

[0063] The time domain characteristics of PID signal include: mean value, standard deviation, skewness, kurtosis, maximum value, minimum value, peak-to-peak value, and the calculation method of characteristic parameters is shown in the table below. iis the amplitude of the signal's sampled data points, and N is the number of sampled data points per sample, 6000. The mean is the signal's average; the standard deviation, also known as the mean square error (MSE), is the square root of the variance and reflects the dispersion of the PID signal. Skewness measures the asymmetry of the probability distribution of a random variable and is used to describe the distribution of PID-generated signals. Skewness measures the direction and degree of skewness in the statistical data distribution and is a numerical characteristic of the degree of asymmetry in the statistical data distribution. If the skewness value is between -0.5 and 0.5, then the data is quite symmetrical, that is, normally distributed. If the skewness value is between -1 and -0.5 (negative skewness) or 0.5 to 1 (positive skewness), then the data is skewed. If the skewness is less than -1 (negative skewness) or greater than 1 (positive skewness), then the data is highly skewed. The maximum and minimum values ​​are the maximum and minimum amplitudes of the one-dimensional signal, respectively. Kurtosis: Kurtosis is a numerical statistic that reflects the distribution characteristics of a random variable. It is used to describe the flatness of a time domain waveform. It is a fourth-order cumulative quantity that reflects the degree of data dispersion. Peak-to-peak value: The difference between the maximum and minimum values ​​reflects the fluctuation range of the signal and describes the peak degree of the waveform. 7 time domain

[0064] Table 2 Calculation methods for mean, standard deviation, skewness, kurtosis, maximum, minimum, and peak-to-peak values

[0065]

[0066] The average value, standard deviation, skewness, kurtosis, maximum value, minimum value, and peak-to-peak value of PID response to 84 groups of VOCs signals with different concentrations are shown in the following table:

[0067] Table 2 Average, standard deviation, skewness, kurtosis, maximum, minimum, and peak-to-peak values ​​of PID response signals for 84 groups of VOCs with different concentrations

[0068]

[0069]

[0070]

[0071] The root mean square, amplitude factor, waveform factor, impact factor, and margin factor of PID responses to 84 groups of VOCs signals with different concentrations are shown in Tables 3, 4, 5, and 6:

[0072] Table 3 RMS, amplitude factor, form factor, impact factor, and margin factor of PID response signals for different concentrations of VOCs in groups 1-17

[0073]

[0074] Table 4 RMS, amplitude factor, form factor, and impact factor of PID response signals for different concentrations of VOCs in groups 18-81

[0075]

[0076]

[0077] Table 5 Margin factors of PID response signals to different concentrations of VOCs in groups 18-81

[0078]

[0079] Table 6 PID response signals for different concentrations of VOCs in groups 82-84: RMS, amplitude factor, form factor, impact factor, and margin factor

[0080]

[0081] (2) PID signal frequency domain feature extraction

[0082] Signal frequency domain analysis method and its application in identifying and analyzing PID response signals under different VOCs concentrations; the frequency domain analysis method realizes the conversion from time domain to frequency domain through Fourier transform, which can intuitively describe the size of different frequency components of the signal, thereby better characterizing the parameters and characteristics of the signal; four frequency domain feature quantities, namely average frequency, center of gravity frequency (FC), mean square frequency (MSF), and standard deviation of frequency (VF), are selected to analyze PID signals.

[0083] The center of gravity frequency (FC) is the center of gravity of the area under the power spectrum curve in the PID signal. The horizontal axis is the frequency, so it is called the center of gravity frequency. The specific formula is as follows:

[0084]

[0085] MSF: used to reflect the energy intensity of the signal spectrum.

[0086]

[0087] VF: Used to reflect the degree of deviation between the PID signal frequency amplitude and its center of gravity.

[0088]

[0089] The four frequency domain characteristic parameters of PID responses to 84 groups of signals with different VOCs concentrations are shown in Table 7.

[0090] Table 7 Average frequency, center of gravity frequency, mean square frequency, and frequency variance of PID response signals for 84 groups of VOCs with different concentrations

[0091]

[0092]

[0093]

[0094] (3) SVM regression based on PID signal feature extraction

[0095] In the SVM regression method based on the extraction of time domain and frequency domain characteristic parameters of PID signals, first, 12 time domain features and 4 frequency domain features of 84 different PIDs that reflect VOCs concentration signals are randomly arranged, and 67 groups of signal time domain and frequency domain characteristic data are used as training sets, and 17 groups of signal time domain and frequency domain characteristic data are used as test sets. Each group of data has 17 values, of which the first 16 values ​​are the characteristic parameters of the PID response signal to VOCs concentration, and the 17th value is the true concentration of VOCs; then the training set and test set are normalized, and then the support vector machine parameters are set, and the Gaussian radial basis kernel function is selected as the kernel function of SVM regression; the radial basis kernel function parameter is still set to 0.8, and the penalty factor C of SVM regression is set to 4. The final test set results are as follows Figure 2 As shown. The mean square error of the model is 1122.25, R 2 is 0.9899, ​​and the model accuracy is Figure 2 The regression effect of this model has certain errors, so this study will optimize the model later.

[0096] SVM regression based on principal component analysis of PID signal characteristics

[0097] 1) Principal component analysis of PID signal time domain and frequency domain characteristics

[0098] Principal components analysis (PCA) is an effective and widely used dimensionality reduction algorithm. The decomposed principal components are mutually orthogonal, which can effectively eliminate redundant and overlapping parts of the original data. Generally speaking, a few important principal components can cover most of the information of the PID response VOCs signal. The calculation process of the PCA algorithm is as follows:

[0099] (1) For a sample set with dimension d and O i =(O1,O2,.....,O n ) is processed by zero mean:

[0100]

[0101] (2) Find the covariance matrix ∑ of O;

[0102] (3) The singular value decomposition method is used to obtain the eigenvalues ​​and eigenvectors of the covariance matrix ∑;

[0103] (4) Take the eigenvectors corresponding to the first v eigenvalues ​​to form a new matrix. In this case, v should be less than n.

[0104] (5) Obtain a new low-dimensional sample set;

[0105] (6) Calculate the principal component contribution rate and cumulative contribution rate.

[0106] The time domain characteristics and frequency domain characteristics of the PID response signal to VOCs are used as the original data and the principal component analysis method is used to extract the principal components.

[0107] The contribution rates of the mean, mean frequency, center of gravity frequency, frequency root mean square, frequency standard deviation, standard deviation, skewness, kurtosis, and maximum value are 0.4495, 0.2439, 0.1082, 0.0787, 0.0557, 0.0300, 0.0152, 0.0080, and 0.0065, respectively. The cumulative contribution rate of the nine features is 0.9958, so these nine features are used to replace the 16 features of the signal.

[0108] 2) SVM regression after principal component analysis of PID signal characteristics

[0109] In the SVM regression method based on principal component analysis of PID signal time domain and frequency domain extraction, 12 time domain features and 4 frequency domain features reflecting VOCs concentration signals generated by 84 different PIDs were subjected to principal component analysis. Feature data of 67 groups of signals after principal component analysis were used as training sets, and feature data of 17 groups of signals after principal component analysis were used as test sets. Each set of data had 10 values, of which the first 9 values ​​were feature parameters after principal component analysis of PID response signals to VOCs concentrations, and the 10th value was the true concentration of VOCs. The training set and test set were then normalized, and then the support vector machine parameters were set, and the Gaussian radial basis kernel function was selected as the kernel function of the SVM regression. The radial basis kernel function parameter was set to 0.8, and the penalty factor C of the SVM regression was set to 4. The final test set results are as follows: Figure 3 As shown. The mean square error of the model is 372, R 2 The regression effect of the model is still a certain error, so the model will be optimized in the future.

[0110] Currently, there is no theoretical guidance for selecting support vector machine parameters. Therefore, when constructing a traditional SVM regression model, the optimal SVM parameters are usually selected through repeated experiments, and the parameters of the regression model with the highest accuracy obtained through experiments are used as the optimal parameters of the SVM regression model. For example, the existing data set is divided into a training set and a test set, and then the parameters of the SVM regression model are changed and recorded, and finally the parameters are confirmed by the optimal regression model. However, when selecting parameters using this method, the parameter values ​​are easily affected by people's subjective experience and it takes a lot of time. By utilizing the search space and global search capabilities of the genetic algorithm, the values ​​of the penalty factor C and the kernel function parameter δ during SVM regression are automatically selected, thereby establishing a genetic algorithm to optimize the SVM regression model and determine the VOCs concentration by generating a signal through PID.

[0111] 1) The impact of support vector machine parameters on its performance

[0112] When a support vector machine (SVM) is used to establish a regression model for quantitative analysis of PID signals to determine VOCs, its modeling accuracy is influenced by the kernel function parameter δ and the penalty factor parameter C. The value of the kernel function parameter δ affects the SVM's ability to predict 17 test data sets using the learning machine derived from the 67 data sets in the training set. This, in turn, affects the generalizability of the SVM regression. The penalty factor parameter C balances the smoothness of the regression curve and the empirical risk in the feature space, thereby affecting the robustness of the regression model. Therefore, when constructing a regression model, it is necessary to comprehensively consider the values ​​of both parameters δ and C to achieve the best VOC concentration prediction results.

[0113] 2) Principal Component Analysis-Genetic Algorithm-SVM Regression

[0114] When using genetic algorithms to solve problems, chromosomes can be encoded using either binary or floating-point encoding. Floating-point chromosome encoding is suitable for solving problems with a large range of values, while binary encoding is suitable for solving problems with a smaller range of values. Because this study uses genetic algorithms to optimize SVM regression with small sample data, binary encoding is used. By studying parameter settings for support vector machine regression in China and abroad, the minimum penalty factor and kernel function parameters for this study were determined to be 0.001, the maximum penalty factor to be 100, and the maximum kernel function parameter to be 10. Genetic algorithms operate on a population basis. Each iteration requires an initial population, which is then used for all subsequent iterations. Based on research on genetic algorithms optimizing support vector machine parameters in China and abroad, this study set the population size to 20. Experimental results indicate that the iteration curve converges when the number of iterations is less than 50, so the number of iterations is set to 50. Population information is defined as a structure, and then a population loop is initiated. The encoding forms a chromosome, and for each gene on the chromosome, a random number between its minimum and maximum values ​​is generated. This is the chromosome. The penalty factor is represented by the first number of the chromosome, and the kernel function parameter is represented by the second number of the chromosome.

[0115] Genetic algorithms distinguish between superior and inferior individuals by evaluating the fitness function value of each chromosome. In genetic algorithms, the larger the chromosome's fitness value, the better the individual represented by that chromosome. After initialization, the optimal target and chromosome are found. Then, iteration begins, selecting an operator, calculating the current target's fitness, calculating the fitness contribution of each chromosome, and cyclically generating non-zero random numbers for the population. Starting with the first chromosome, the fitness contribution is accumulated. When the accumulated fitness contribution is greater than the random number, the last individual accumulated is selected as the selected individual. A progeny of the same size as the original population is selected. The optimization variable dimension and population size are calculated. The crossover probability is then determined. A random number is generated and compared to the crossover probability to determine if crossover occurs. If so, two consecutive chromosomes are randomly selected; otherwise, crossover is meaningless. The two selected chromosomes are derived and each dimension of the chromosome is iterated over. The crossover operator simulates binary crossover. The bounds of the individuals after crossover are then set: values ​​exceeding the maximum value are replaced with the maximum value, and values ​​below the minimum value are replaced with the minimum value. The individuals that have completed the crossover are copied into the population, and then the population is looped to generate a random number and compare it with the crossover mutation probability. If a mutation occurs, a random individual is selected to prepare for the mutation. Each gene of the chromosome is looped, and the selected chromosome is subjected to polynomial mutation. Then the mutated individuals are copied into the population, and the targets of the offspring that have completed the crossover mutation are recalculated to obtain the optimal target and the worst target. Then the current optimal target is compared with the global optimal target. Replacing the worst target with the historical global optimal target can increase the probability of the population iterating to produce a better individual. The average target and optimal target of this generation are recorded. After the iteration, C and δ are derived. The iterative curve of the genetic algorithm to optimize the SVM regression parameters is as follows. Figure 4 As shown in Figure 1, the iterative curve has converged when the number of iterations is 16, that is, the minimum value of MSE can be obtained after the 16th iteration.

[0116] In the genetic algorithm optimized SVM regression method based on principal component analysis extracted from time domain and frequency domain of PID signals, 9 principal component features reflecting VOCs concentration signals generated by 84 different PIDs were used as signal features, 67 groups of characteristic data after principal component analysis of signals were used as training sets, and 17 groups of characteristic data after principal component analysis of signals were used as test sets. Each group of data had 10 values, of which the first 9 values ​​were characteristic parameters after principal component analysis of the PID response signal to VOCs concentration, and the 10th value was the actual concentration of VOCs.

[0117] The training set and test set are normalized, and the optimal parameters of the radial basis kernel function are found to be 0.0101 and the optimal penalty factor C is 7.8783 through the genetic algorithm. Based on the selection of parameters, the support vector machine regression function F(x) = w*x+b is derived. The final test set results are as follows Figure 5 As shown. The mean square error of the model is 0.000059, R2 is 0.9999. Compared with the R obtained by Wang Jin 2 is 99.8%, Li Hai obtained R 2 The results are 98.2%, 98.5%, 96.9%, etc. This method is superior.

[0118] 3) The impact of sample number on PCA-GA-SVM model

[0119] In order to verify the validity of the model with 84 samples meeting the requirements of high accuracy and to find the minimum number of samples for PCA-GA-SVM model construction, on the basis of 67 training sets and 17 test sets, 4 training sets and 1 test set were reduced in turn to find the minimum number of samples for model construction. A is the original data; B is the data of 4 training sets and 1 test set; C is the data of 8 training sets and 2 test sets; D is the data of 12 training sets and 3 test sets; E is the data of 16 training sets and 4 test sets; F is the data of 20 training sets and 5 test sets. The accuracy results of PCA-GA-SVM model are shown in the figure. Figure 6 shown.

[0120] When reducing the data to 12 training sets and 3 test sets, the R 2 The accuracy of the PCA-GA-SVM model can still be maintained above 0.99. However, when the training and test sets are further reduced, the accuracy of the PCA-GA-SVM model will drop sharply. Therefore, when using the PCA-GA-SVM regression model to calculate VOC concentrations based on PID-generated signals, the sample size should be greater than 69 groups. The PID response to VOC signals is 84 groups, thus meeting the model establishment requirements.

[0121] 4) Application of PCA-GA-SVM model before and after denoising in humidity environment

[0122] To verify the effectiveness of the wavelet packet decomposition node energy adaptive weighting method for denoising VOCs signals collected in high humidity environments, a wet cotton ball was placed at the PID air inlet, affecting the PID signal and generating noise. A gas distribution device was used to prepare 12 sets of benzene gas with varying concentrations. The accuracy of the PCA-GA-SVM model for quantitative VOC analysis in these humidity environments was compared with that without this denoising method. Nine features (mean, standard deviation, skewness, kurtosis, maximum value, frequency mean, center of gravity frequency, frequency root mean square, and frequency standard deviation) were identified as the principal components of the time and frequency characteristic parameters of the PID signal. The 12 sets of noisy PID response VOCs signals each had nine characteristic parameters in the time and frequency domains.

[0123] The 9 characteristic parameters of the 12 groups of noisy signals and the 12 groups of denoised signals were respectively imported into the PCA-GA-SVM formula, and the concentrations of the noisy and denoised VOCs were obtained as shown in Tables 8 and 9.

[0124] Table 8 Comparison of calculations of 10ppm-400ppm VOCs noisy signal and denoised signal by PCA-GA-SVM model

[0125]

[0126] Table 9 Comparison of calculations of 401ppm-1000ppm noisy signal and denoised signal by PCA-GA-SVM model

[0127]

[0128] The data in the table above shows that humidity has the most severe impact on the accuracy of low-concentration VOC monitoring. The PCA-GA-SVM model achieves a maximum accuracy of 99.5% for VOC monitoring in humid environments. Denoising the PID signal using adaptive threshold-weighted wavelet packets improves the accuracy by up to 14.9%. This demonstrates the applicability of the PCA-GA-SVM model to VOC monitoring in humid environments. It also demonstrates that the constructed node energy adaptive weighted denoising method, based on wavelet packet decomposition, can improve VOC monitoring accuracy while simultaneously increasing the signal-to-noise ratio.

[0129] VOCs gas concentration prediction model based on random forest algorithm

[0130] 2.2.1 Construction of random forest VOCs gas concentration prediction model

[0131] The Random Forest (RF) algorithm is a highly efficient bagging algorithm for solving classification and regression problems. Its principle is based on constructing multiple decision trees, each of which is trained on a random subset. Bootstrap sampling is used to extract training samples. Features are randomly selected when nodes are split to increase the diversity of the model. Finally, their prediction results are integrated by voting (classification) or averaging (regression). Through the bagging ensemble method, random forest can reduce the risk of overfitting and improve the generalization ability of the model. This method is highly robust to outliers and is therefore widely used in practical applications. The random forest algorithm generally includes the following steps:

[0132] (1) Randomly select a certain number of samples from all samples with replacement, repeat the sampling k times, and obtain a set of k samples;

[0133] (2) train k trees with k sample sets, and use randomly selected features as partition nodes during the training process;

[0134] (3) Integrate the prediction results of each tree and use voting (classification) or taking the average (regression) to determine the final category.

[0135] The RF model has shown significant advantages in the field of feature vector prediction due to its advantages such as few adjustable parameters, high integrated training efficiency, and fast convergence speed. Compared with a single decision tree, the random forest algorithm can effectively improve the prediction accuracy of the model by performing weighted fitting on each prediction result.

[0136] This section builds a VOCs gas concentration prediction model based on the configured known benzene gas concentration data. The main steps include:

[0137] (1) In the data preprocessing stage, the original benzene concentration data set was subjected to outlier removal and standardization;

[0138] (2) Feature selection: the voltage value corresponding to the benzene gas concentration measured by the PID acquisition device is used as the model input feature;

[0139] (3) Use the random forest algorithm to build a prediction model and determine the optimal number of decision trees and leaf nodes through training and debugging;

[0140] (4) Using the optimized parameter configuration to train the final prediction model;

[0141] (5) Use model evaluation indicators to evaluate the model performance and compare it with other typical prediction models.

[0142] Data preprocessing is the primary step in building a random forest model for VOC gas concentration prediction. The voltage corresponding to the benzene gas concentration measured by a PID acquisition device serves as the model input feature. The dataset is partitioned using the MATLAB cvpartition function, with 80% of the samples randomly assigned to the training set for model training and the remaining 20% ​​used as the test set for performance verification. To ensure the randomness of the data distribution, the dataset is pre-shuffled using the randperm function to eliminate potential sorting bias.

[0143] The VOCs gas concentration prediction flow chart of the random forest algorithm is as follows Figure 7 As shown:

[0144] Step 1: Sample selection, by resampling the original data set, a sample set is selected from the entire data set;

[0145] Step 2: Build a decision tree. Use the sub-training set obtained in the first step to split each node one by one. This recursive process can build a complete decision tree.

[0146] Step 3: Deepen the tree structure. To optimize the choice of each decision node, calculate the information gain to determine the best choice at the current node, which helps to further deepen the tree structure.

[0147] Step 4: Repeat the construction process. The above tree deepening process will be repeatedly executed until a complete decision tree is completed.

[0148] Determining the number of decision trees

[0149] In the parameter optimization process of the random forest model, the number of decision trees (n_estimators) is a key parameter, and its value directly affects the performance of the model. While keeping other parameters at their default settings, a grid search is performed on the number of decision trees, with the test range set to 1 to 200 and a step size of 1. Figure 8 As shown in the figure, by drawing the relationship curve between the number of decision trees and the model root mean square error (RMSE), the changing trend of model error as the number of decision trees increases is intuitively demonstrated.

[0150] Depend on Figure 8 It can be seen that as the number of decision trees increases, the root mean square error of the model gradually decreases. When the number of trees increases to about 10, the error of the model fluctuates, but there is no obvious change. This error trend shows that as the number of decision trees increases, the fitting effect of the random forest model is better and the prediction accuracy is higher. It also shows that when the number of decision trees reaches a certain level, the effect of continuing to increase the number of decision trees on improving the model performance shows a marginal decreasing trend, and the model error gradually stabilizes. Under the initial parameter configuration conditions, the number of decision trees is set to range from 10 to 100, and the step size is 1. Figure 9 The learning curves shown systematically demonstrate the evolution of model performance as the number of decision trees increases.

[0151] from Figure 9 It can be observed that the root mean square error of the model is 0.000436919 when the number of decision trees is 90. Taking into account the requirements of running time and prediction accuracy, this prediction model finally determines the number of decision trees to be 90.

[0152] Data preprocessing

[0153] To prevent overfitting of the prediction model, the dataset needs to be properly partitioned. For small sample sizes, a typical approach is to divide the data into a training set (for model building) and a test set (for performance evaluation) in an 8:2 ratio. When a validation set is required, a 60% training set, 20% validation set, and 20% test set partitioning is used.

[0154] Concentration tests were conducted on various benzene gas concentrations, yielding a total of 103 sets of PID voltage data. After removing some invalid, duplicate, or erroneous data, 96 valid samples were obtained. Given the small sample size, the original dataset was randomly divided into an 80% training set and a 20% test set.

[0155] To improve the convergence speed of the model, increase the accuracy of predictions, and reduce the sensitivity of noise to data, it is necessary to remove the impact of data dimensionality differences. To this end, it is necessary to first normalize the data and uniformly map the sample data to the range [0, 1] to achieve a uniform data parameter scale. The calculation formula is shown in (1). After the model training process is completed, denormalization is performed using formula (2).

[0156]

[0157] X=X1(X max -X min )+X min (2)

[0158] Where X max is the maximum value of the sample data; X min is the minimum value of the sample data.

[0159] Model evaluation metrics

[0160] To evaluate the effectiveness and accuracy of the prediction model, the mean square error (MSE), mean absolute error (MAE), root mean square error (RMSE), coefficient of determination (R 2 ) Four indicators are used for evaluation;

[0161] 1) Mean Squared Error (MSE)

[0162] The mean square error calculates the mean of the sum of the squares of the differences between the predicted values ​​and the true values ​​for all samples. The MSE value ranges from 0 to positive infinity. The smaller the value, the better the model fit and the higher the prediction accuracy. The formula for the mean square error is as follows:

[0163]

[0164] Where n is the number of sample data in the prediction model; y i is the true value of the sample data; Predicted values ​​for sample data.

[0165] 2) Root Mean Square Error (RMSE)

[0166] The root mean square error (RMSE) calculates the square root of the mean of the sum of the squares of the differences between the predicted values ​​and the true values ​​of all samples, i.e., the arithmetic square root of the mean square error. RMSE has a similar function to the mean square error. The smaller the value, the better the fit of the prediction model and the higher the accuracy. The formula for the root mean square error is as follows:

[0167]

[0168] Where n is the number of sample data in the prediction model; y i is the true value of the sample data; Predicted values ​​for sample data.

[0169] 3) Mean Absolute Error (MAE)

[0170] The mean absolute error (MAE) calculates the average of the absolute differences between the predicted and actual values ​​for all samples. MAE provides an intuitive and easy-to-understand error measure, reflecting the average level of model error. Smaller values ​​indicate more accurate model predictions.

[0171]

[0172] Where n is the number of sample data in the prediction model; y i is the true value of the sample data; Predicted values ​​for sample data.

[0173] 4) Coefficient of determination (R 2 )

[0174] The coefficient of determination is a measure of the goodness of fit of a model. It provides a measure of the model's ability to explain data fluctuations. Its value ranges from 0 to 1, with values ​​closer to 1 indicating a better fit for the forecasting model. It provides an intuitive assessment of model performance.

[0175]

[0176] Where n is the number of sample data in the prediction model; y i is the true value of the sample data; Predicted values ​​for sample data; is the mean of the true value of the sample data.

[0177] Simulation experiment and result analysis

[0178] To systematically evaluate the performance of gas concentration prediction models, simulation experiments were conducted using a BP neural network, random forest (RF), support vector regression (SVR), long short-term memory (LSTM), radial basis function (RBF), and extreme learning machine (ELM) models as comparison models. All comparative experiments were conducted using the same dataset and evaluation system to ensure comparability and fairness of the results.

[0179] Analysis of experimental results based on random forest regression prediction model

[0180] Random forest regression (RF) is used for prediction. A random forest regression model is trained using the TreeBagger function. The number of decision trees in the random forest model is set to 90, which has a small deviation. Therefore, the number of decision trees is set to 90 (trees=90), and the minimum number of leaf nodes in each tree is 5 (leaf=5). The voltage value of the VOCs gas concentration measured by the PID acquisition device is used as the input variable, and the actual value of the VOCs gas concentration is used as the output variable. The test set results of the model are as follows: Figure 10 As shown:

[0181] In the evaluation of VOCs gas concentration prediction by random forest model, several key indicators were used, including MSE of 2150.1018, RMSE of 46.3692, R 2 It is 0.98824 and MAE is 38.1817. Figure 10 The fitting curve results show that the random forest VOCs gas concentration model performs well in prediction.

[0182] The coefficient of determination is as high as 0.98824, indicating that the model is very effective in prediction and demonstrates its excellent data fitting ability. Taken together, these evaluation indicators reflect the accuracy and reliability of the random forest prediction model in predicting gas concentrations.

[0183] Analysis of experimental results based on BP neural network regression prediction model

[0184] The BP neural network was used for prediction and the training parameters were set, including the maximum number of iterations of 1000 and the training target error threshold of 1×10 -6 The neural network model constructed consists of three main layers: an input layer, a hidden layer with 5 neurons, and an output layer. The voltage value of the VOCs gas concentration measured by the PID acquisition device is used as the input variable, and the actual value of the VOCs gas concentration is used as the output variable. Detailed training was carried out in the MATLAB environment, and the prediction results were obtained as follows: Figure 11 As shown:

[0185] The BP neural network model was evaluated using the same evaluation indicators and the results showed that: MSE was 27999.2277, RMSE was 167.3297, MAE was 119.8138, R 2When comparing these indicators with the corresponding performance of the RF model, it was found that the BP neural network performed poorly in terms of error indicators, indicating that its accuracy and reliability in predicting VOCs gas concentration in this experiment were not as good as those of the RF model. 2 This shows that the model is lacking in the ability to explain data changes. Therefore, based on the comparative analysis of these indicators, it can be concluded that the RF proposed in this paper is superior to the traditional BP neural network model.

[0186] Analysis of experimental results based on long short-term memory network regression prediction model

[0187] A long short-term memory network (LSTM) was used for prediction, with 1 iteration per round and 1500 iterations. A neural network structure consisting of a 120-dimensional input layer, a 4-unit LSTM layer, a ReLU activation layer, a fully connected layer, and a regression output layer was constructed. The Adam optimizer was used, with a maximum training round of 1500 and an initial learning rate of 0.01. A learning rate decay strategy was introduced (decaying to 10% of the original value every 1200 rounds), data shuffling, and training process visualization were enabled. The VOCs gas concentration voltage value measured by the PID acquisition device was used as the input variable, and the actual VOCs gas concentration value was used as the output variable. Detailed training was performed in the MATLAB environment, and the prediction results were obtained as follows: Figure 12 As shown:

[0188] In the evaluation of VOCs gas concentration prediction by long short-term memory network model, several key indicators were used, including MSE of 6608.8223, RMSE of 81.2947, R 2 The error rate is 0.9696 and the MAE is 70.8576. When comparing these indicators with the corresponding indicators of RF, it is found that RF has better performance in all these error indicators. This shows that the RF model has a better fitting effect and higher prediction accuracy than the LSTM model in predicting VOCs gas concentration in this experiment. 2 This indicates that the RF model has a stronger ability to explain data fluctuations.

[0189] Analysis of experimental results based on radial basis function neural network regression prediction model

[0190] A radial basis function (RBF) neural network is used for prediction. The newrbe function is used to create the network, and rbf_spread (radial basis function expansion speed) is set to 100. The newrbe function creates a hidden layer of neurons for each training sample. That is, the neural network model constructed this time contains three main layers: an input layer, a hidden layer with 75 neurons, and an output layer.

[0191] The voltage value of the VOCs gas concentration measured by the PID acquisition device is used as the input variable, and the actual value of the VOCs gas concentration is used as the output variable. Detailed training was carried out in the MATLAB environment to obtain the prediction results as follows Figure 13 As shown:

[0192] By carefully analyzing and comparing the prediction results of the radial basis function neural network model and the RF prediction model, the MSE of the radial basis function neural network model is 23531.7812, MAE is 114.7546, RMSE is 153.4007, and R 2 When comparing these indicators with the corresponding performance of the RF model, it was found that the radial basis function neural network model had a larger error value, indicating that its prediction accuracy and fitting effect were not as good as those of the RF model in predicting VOCs gas concentration in this experiment. 2 This shows that the model lacks the ability to explain data fluctuations. Therefore, based on the comparative analysis of these indicators, it can be concluded that RF is superior to the traditional radial basis function neural network model.

[0193] Analysis of experimental results based on extreme learning machine regression prediction model

[0194] An extreme learning machine (ELM) was used for prediction, with 50 hidden layer nodes and a Sigmoid activation function. The elmtrain function was used to randomly initialize the input weights (IW) and bias (B). The output weights (LW) were directly calculated analytically, eliminating the need for iterative optimization, allowing for rapid model training. The VOCs gas concentration voltage measured by the PID acquisition device was used as the input variable, and the actual VOCs gas concentration value was used as the output variable. Detailed training was performed in the MATLAB environment, resulting in the prediction results shown in Figure 14:

[0195] In the evaluation of VOCs gas concentration prediction by extreme learning machine, several key indicators were used, including MSE of 19029.6183, RMSE of 137.9479, R 2 The error coefficients of the RF model are 0.9244 and the MAE is 112.1774. When comparing these indicators with the corresponding indicators of RF, it is found that RF has better performance in all these error indicators. This shows that the RF model has a better fitting effect and higher prediction accuracy than the ELM model in predicting VOCs gas concentration in this experiment. 2 This indicates that the RF model has a stronger ability to explain data fluctuations.

[0196] Comparison of evaluation indicators of different models

[0197] Six representative prediction models were selected for VOCs gas concentration prediction analysis, including random forest regression (RF), support vector regression (SVR), BP neural network, long short-term memory network (LSTM), radial basis function neural network (RBF) and extreme learning machine (ELM). These models are based on different algorithm principles and data structure characteristics, and can comprehensively predict and deeply analyze the trend of gas concentration changes from a multivariate perspective. The unified evaluation indicators MSE, RMSE, MAE and R 2 The evaluation index results of the six prediction models, RF, SVR, BP, RBF, LSTM, and ELM, are shown in Table 10.

[0198] Table 10 Comparison of error values ​​of different models

[0199]

[0200] The difference between the model prediction value and the true value is measured by the three indicators of mean square error (MSE), root mean square error (RMSE) and mean absolute error (MAE). The smaller the value, the better the model fitting effect and the higher the prediction accuracy. 2 ) measures the model's ability to explain data fluctuations. Its value ranges from 0 to 1. The closer the value is to 1, the better the fitting effect of the prediction model.

[0201] As shown in the table above, in this experiment, the error metrics MSE, RMSE, and MAE were ranked as follows: RF > LSTM > SVR > ELM > RBF > BP. The RF model performed best, with an MSE of 2150.1018, an RMSE of 46.3692, and a MAE of 38.1817. The BP neural network model performed worst, with an MSE of 27999.2277, an RMSE of 167.3297, and a MAE of 119.8138. This indicates that the random forest model provided very accurate predictions in this experiment and also performed well in data fitting. Coefficient of Determination R 2 The ranking results are RF>LSTM>SVR>RBF>ELM>BP. That is, the RF model performs best, and its R 2 is 0.98824; the BP neural network model performs the worst, with its R 2 The value is 0.9157, which shows that the random forest model performs well in explaining data fluctuations in this experiment.

[0202] The comparison results of the evaluation indicators of the six prediction models, RF, SVR, BP, RBF, LSTM and ELM, are as follows: Figure 15 As shown:

[0203] As can be seen from the figure, in this experiment, compared with other prediction models, the error indicators MSE, RMSE, and MAE of the RF model are the lowest, and the determination coefficient R 2 The superior performance of random forests is attributed to their ensemble learning approach, which combines multiple decision trees to reduce overfitting and improve prediction accuracy. The results of this experiment show that RF performs best in both prediction accuracy and data fit, making it the most suitable model for VOCs gas concentration prediction in this experiment.

[0204] VOCs gas concentration prediction model based on particle swarm optimization random forest

[0205] The experiments in the previous section show that while the random forest model has excellent overall prediction performance, its prediction errors in certain local areas are large, and its training speed is relatively slow. This is primarily because traditional grid search methods struggle to accurately balance the impact of the number of decision trees on overfitting and underfitting. That is, too many decision trees may lead to overfitting, while too few may lead to underfitting. To this end, a particle swarm optimization algorithm (PSO) is introduced to optimize the key parameters of the random forest. A PSO-RF prediction model is constructed. This dynamic parameter optimization mechanism effectively determines the optimal parameters of the decision trees in the random forest, thereby better predicting VOC gas concentrations.

[0206] Particle Swarm Optimization Algorithm

[0207] The particle swarm optimization (PSO) algorithm is a biomimetic intelligent algorithm inspired by the foraging behavior of birds in nature. Its basic principle is that when birds forage, they communicate and share information, sharing their geographic location with other individuals. This allows them to determine whether their location is optimal, ultimately finding a global optimal solution through group collaboration.

[0208] The two core elements of the algorithm are the particle's velocity (the direction and distance of the particle's next movement) and position (a solution to the problem the algorithm is looking for). The particle swarm optimization algorithm generally includes the following steps:

[0209] (1) Initialize the particle swarm: First, determine the size of the swarm and the maximum number of iterations, and then initialize the position and velocity of the particles. The position represents the possible solutions, while the velocity represents the direction and speed of the search in the solution space;

[0210] (2) Determination of initial fitness: Assign an initial fitness value to each particle as its individual optimal solution, and calculate the fitness function value of all particles;

[0211] (3) Update the individual optimal solution: Based on the current position of each particle and its individual historical optimal position, update its individual optimal solution;

[0212] (4) Update the global optimal solution: Select the particle with the best fitness value in the entire particle swarm and take its position as the global optimal solution;

[0213] (5) Update the particle position and velocity: Based on the individual historical optimal position and the global optimal position, update the position and velocity of each particle;

[0214] (6) Determine the end condition: Determine whether the stop condition is met (for example, a solution that meets the requirements is found or the maximum number of iterations is reached). Otherwise, repeat steps 4 to 6.

[0215] (7) Output optimal parameters: When the algorithm meets the termination conditions, it outputs the optimal solution, that is, the optimal parameter combination of the random forest model.

[0216] The flow chart of particle swarm optimization algorithm is as follows Figure 16 As shown:

[0217] Construction of PSO-RF VOCs gas concentration prediction model

[0218] Based on the above content, the construction process of the particle swarm optimization random forest model (PSO-RF) is as follows:

[0219] (1) Input VOCs gas concentration data, preprocess the data and divide the data set into training set and test set;

[0220] (2) Selecting eigenvalues ​​for the random forest model;

[0221] (3) Setting the initial velocity and coordinates of the particles, and using the particle swarm algorithm to perform cyclic optimization to adjust the moving speed and position of each particle in the particle swarm;

[0222] (4) Set the initial parameters of the random forest, build a decision tree based on it, and then train the random forest. Evaluate the effectiveness and accuracy of the model through evaluation indicators and calculate its error rate;

[0223] (5) When the error rate of the model cannot meet the requirements, return to step 3 and continue to optimize the model through the particle swarm optimization algorithm to find the optimal parameters. Once satisfactory model parameters are found or other termination conditions are met, the optimization process ends.

[0224] The flow chart of the VOCs gas concentration prediction model based on PSO-RF is as follows: Figure 17 Shown: PSO-RF model parameter determination

[0225] (1) Determination of particle number

[0226] In the implementation of the particle swarm optimization algorithm, the number of particles is a key parameter that needs to be determined in the initial stage. Increasing the number of particles can enhance the algorithm's ability to explore the solution space and increase the probability of obtaining the global optimal solution, but it will also significantly increase the computational burden. For conventional optimization problems, the number of particles is generally set in the range of 20 to 40. This study adopts a systematic parameter scanning method to gradually increase the number of particles (20, 21...40) in unit steps, and observes the model prediction performance indicators (such as RMSE, R 2 ) to determine the optimal number of particles. The specific experimental results are as follows: Figure 18 As shown:

[0227] according to Figure 18 It can be seen that when the number of particles is 28, the model's evaluation indicators reach the best, with specific values ​​of MSE = 18.6939, MAE = 15.4326, and RMSE = 349.4626. Therefore, comprehensive analysis determines that the most suitable number of particles for this model is 28.

[0228] (2) Random forest parameter determination

[0229] The particle swarm optimization algorithm (PSO) is used to optimize two key parameters of the random forest (RF): the number of decision trees and the minimum number of leaves. The parameters of the PSO algorithm are as follows: the number of particles is 28, the maximum number of iterations is set to 50, the individual learning factor (c1) and the group learning factor (c2) are both 2, and the inertia weight w is 0.9. The optimization process ends when the maximum number of iterations is reached or the error is stable. The convergence process is as follows Figure 19 shown.

[0230] like Figure 19 As shown in the figure, analyzing the error trend over iterations reveals a clear downward trend over 50 iterations. The error reaches its minimum at the 11th iteration, indicating that the algorithm has converged to the optimal solution. Furthermore, the evolution of two key parameters shows that the number of decision trees stabilizes at 123 after the 11th iteration, and the minimum number of leaf nodes also converges to 1. At this point, the random forest parameters have reached their optimal combination, with 123 decision trees and a minimum number of leaves of 1.

[0231] Analysis of simulation experiment results

[0232] After optimizing the best combination based on the PSO algorithm, the model code was run using MATLAB. The number of decision trees was 123, and the minimum number of leaves was 1, achieving the VOCs gas concentration prediction. The prediction results are as follows: Figure 20 As shown:

[0233] Figure 20The red line shows the actual concentration of VOCs gas in the test set, and the blue line shows the prediction result of the particle swarm optimization random forest model. The application of the particle swarm optimization algorithm further enhances the performance of the random forest model, enabling the PSO-RF model to more accurately capture the inherent patterns and trends of the data when predicting VOCs gas concentration. Figure 20 It can be seen that its RMSE is 9.7176, MSE is 96.1486, R 2 The average error (MAE) is 0.99973 and the average error (MAE) is 7.6165, indicating that the overall performance of the PSO-RF model is significantly better than that of the single RF model. This result not only demonstrates the effectiveness of the particle swarm optimization algorithm in improving the performance of the prediction model, but also shows that the proposed PSO-RF model has a high degree of reliability and accuracy in predicting VOCs gas concentrations.

[0234] The comparison of the error values ​​of random forest and particle swarm optimization after random forest is shown in Table 11:

[0235] Table 11 Comparison of various error values ​​between random forest and particle swarm optimization

[0236]

[0237] From the comparison in Table 11, it can be seen that the PSO-RF model has a better performance in RMSE, MSE, and R 2 The RMSE, MSE, and R of the PSO-RF model are excellent. 2 , MAE, and achieved significant values, respectively. Compared with the single RF model, these four evaluation indicators have all been significantly improved. Compared with the traditional RF model, the improvement of the PSO-RF model in evaluation indicators shows that it has a clear advantage in prediction accuracy.

[0238] Advantages of this solution:

[0239] (1) Intelligent parameter optimization mechanism

[0240] The particle swarm optimization (PSO) algorithm was used to adaptively tune the key hyperparameters of the random forest. By establishing a 28-dimensional search space (number of particles = 28) and setting a dynamic inertia weight of 0.9 to balance global and local searches, the optimal parameter combination was obtained after 11 iterations of convergence, reducing the model prediction error by more than 95%.

[0241] (2) Hybrid model architecture design

[0242] It innovatively combines the ensemble learning advantages of random forests with the optimization capabilities of PSO: RF provides powerful feature processing and nonlinear modeling capabilities, while PSO solves the RF parameter sensitivity problem through swarm intelligent search. The two work together to achieve the "1+1>2" effect.

[0243] (3) Dynamic weight control technology

[0244] In the PSO optimization process, a nonlinear decreasing inertia weight strategy (0.9→0.4) was adopted to maintain a strong global search capability in the early stage and enhance local fine tuning in the later stage, effectively avoiding premature convergence and making the model R 2 Improved to 0.99973.

[0245] (4) Multi-objective optimization framework

[0246] A composite fitness function that includes prediction accuracy (MSE), generalization ability (cross-validation error), and computational efficiency (training time) is constructed to ensure that the optimized model achieves a balance in multiple performance dimensions.

[0247] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A concentration detection method for VOCs concentration in water, soil and air based on a random forest algorithm, characterized in that: Includes the following: (1) Input VOCs gas concentration data, preprocess the data and divide the data set into training set and test set; (2) Selecting eigenvalues ​​for the random forest model; (3) Set the initial velocity and coordinates of the particles, and use the particle swarm algorithm to perform cyclic optimization to adjust the moving speed and position of each particle in the particle swarm; (4) Set the initial parameters of the random forest, build a decision tree based on it, and then train the random forest; evaluate the effectiveness and accuracy of the model through evaluation indicators and calculate its error rate; (5) When the error rate of the model cannot meet the requirements, return to step (3) and continue to optimize the model through the particle swarm optimization algorithm to find the optimal parameters; once satisfactory model parameters are found or other termination conditions are met, the optimization process ends.

2. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: In step (1), in order to improve the convergence speed of the model, improve the accuracy of the prediction, and reduce the sensitivity of noise to the data, it is necessary to remove the impact of data dimension differences; to this end, the data needs to be normalized.

3. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: To prevent the prediction model from overfitting, the data set needs to be divided reasonably. For small sample data, the data is divided into a training set (model construction) and a test set (performance evaluation) in a ratio of 8:

2.

4. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: Step (2) of the random forest model includes the following: In the data preprocessing stage, the original benzene concentration data set was subjected to outlier removal and standardization; In the feature selection phase, the voltage value corresponding to the benzene gas concentration measured by the PID acquisition device is used as the model input feature; Use the random forest algorithm to build a prediction model, and determine the optimal number of decision trees and leaf nodes through training and debugging; The final prediction model is trained using the optimized parameter configuration; Model evaluation indicators are used to evaluate the model performance and compared with other typical prediction models.

5. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: The particle swarm optimization algorithm in step (3) includes the following steps: 1) Initialize the particle swarm: First, determine the size of the swarm and the maximum number of iterations, and then initialize the position and velocity of the particles; the position represents the possible solution, while the velocity represents the direction and speed of the search in the solution space; 2) Determination of initial fitness: Assign an initial fitness value to each particle as its individual optimal solution, and calculate the fitness function value of all particles; 3) Update the individual optimal solution: Based on the current position of each particle and its individual historical optimal position, update its individual optimal solution; 4) Update the global optimal solution: Select the particle with the best fitness value in the entire particle swarm and take its position as the global optimal solution; 5) Update particle position and velocity: Update the position and velocity of each particle based on the individual historical optimal position and the global optimal position; 6) Determine the end condition: Determine whether the stop condition is met, otherwise repeat steps 4) to 6); 7) Output optimal parameters: When the algorithm meets the termination conditions, it outputs the optimal solution, that is, the optimal parameter combination of the random forest model.

6. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 5, characterized in that: In the implementation process of the particle swarm optimization algorithm, the number of particles is a key parameter that needs to be determined in the initial stage. The number of particles is determined by a systematic parameter scanning method, which gradually increases the number of particles in unit steps. The optimal number of particles is determined by observing the changing trend of the model prediction performance indicators.

7. The method for detecting VOCs concentration in water, soil, and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: In step (4), the particle swarm optimization algorithm is used to optimize two key parameters of random forest: the number of decision trees and the minimum number of leaves.

8. The method for detecting VOCs concentration in water, soil and air using a random forest algorithm-based correlation prediction model according to claim 1, characterized in that: Model evaluation indicators in step (4): through square error (MSE), mean absolute error (MAE), root mean square error (RMSE), determination coefficient (R 2 ) four indicators for evaluation.

Citation Information

Cited By

  • Dam body bedrock operation parameter inversion method combining three-dimensional finite element and optimization algorithm

    CN121389609A

  • Microbial solidified soil constitutive model parameter identification method

    CN121935519A