Machine learning-based model training system and method for predicting volatile organic compounds, and online early warning device
Through a machine learning-based model training system, using standard equipment, PID sensors and meteorological data, the mass concentration of volatile organic matter in industrial parks is predicted, which solves the problem of the inability to detect VOCs quickly and accurately in the existing technology, and realizes an efficient and economical online warning function.
Patent Information
- Application Number
- CN202510585524.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The prior art is difficult to quickly and accurately detect the concentration of volatile organic compounds (VOCs) in industrial parks, especially in mass concentration (μg/m3), resulting in the inability to realize the real alarm function.
Using a machine learning-based model training system, data preprocessing and model training are carried out by obtaining standard equipment detection data, PID sensor detection data and meteorological data of target monitoring points, and final prediction model is constructed to predict the quality concentration of standard equipment detection data.
The quality concentration of the detection data of standard equipment through PID sensor detection data and meteorological data prediction is realized, and it has the advantages of shorter monitoring frequency, small size, cheaper price, and easy to distribute points, and can be used for real-time online warning.
Smart Images

Figure CN120105252A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model training system, method and online early warning device for volatile organic compound prediction based on machine learning. Background Art
[0002] In my country, in the online monitoring of volatile organic compounds (VOCs) in industrial park grid monitoring, the mass concentration (μg / m 3 ) is used as the existing alarm limit unit to reflect and alarm the detection concentration of each organic matter.
[0003] At present, standard equipment such as gas chromatography-mass spectrometry (GC-MS) or gas chromatography-flame ionization detector (GC-FID) is usually used to detect the concentration of VOCs. They have the advantages of high sensitivity, wide linear range, good stability, etc., especially the ability to analyze single VOC species and to measure the concentration in terms of mass concentration (μg / m 3 ) However, such equipment is relatively expensive, takes a long time to analyze, is not representative enough, has high requirements for the working environment, and has high maintenance costs.
[0004] Photo ionization detector (PID) is one of the detection methods of VOCs. It uses ultraviolet light (UV) as a light source to break substances into positive and negative ions (ionization) that can be detected by the detector. The detector measures the charge of the ionized gas and converts it into a current signal. The current is amplified and displays the corresponding concentration value. After being detected, the ions recombine into the original gas and vapor. It is a widely used detector. PID has the advantages of shorter monitoring frequency (minutes or even seconds), small size, low price, and easy to deploy in large quantities. However, it cannot monitor a single VOC substance, but responds to one or more types of pollutants (such as aromatic hydrocarbons and alkanes). The total volume fraction of VOCs detected is obtained, and since the composition ratio of VOCs is unknown, the volume fraction cannot be converted into mass concentration (μg / m 3 ), can only be displayed in volume percentage and cannot be compared with the existing alarm limit unit of μg / m 3 No corresponding comparison is performed, so that the real alarm function cannot be realized.
[0005] Therefore, there is an urgent need for a rapid detection method that can measure the mass concentration (μg / m 3 ) reflects the online VOCs monitoring plan for industrial parks for each substance being tested. Summary of the invention
[0006] In order to solve the problems in the prior art, the present application provides a model training system, a model training method and an online early warning device for volatile organic compound prediction based on machine learning. The technical solution of the present application is as follows:
[0007] 1. A model training system for volatile organic compound prediction based on machine learning, comprising:
[0008] A data acquisition module, wherein the data acquisition module acquires detection data of target monitoring points as training sets and test sets; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data;
[0009] A data preprocessing module, which preprocesses the detection data to obtain a processed training set and a processed test set;
[0010] A data training module, wherein the data training module trains the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data;
[0011] A model evaluation module compares the predicted value of the standard equipment detection data obtained by using the final prediction model to predict the test set and the true value of the standard equipment detection data to obtain an evaluation result.
[0012] 2. The model training system of item 1, wherein:
[0013] The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
[0014] 3. The model training system of item 1, wherein:
[0015] The data preprocessing module comprises:
[0016] a data separation component, wherein the data separation component divides the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and divides the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively;
[0017] A feature standardization component, wherein the feature standardization component standardizes the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data;
[0018] A dimension-upgrading component is used to upscale the training feature standardized data and the test feature standardized data to obtain training feature upscaled dimension data and test feature upscaled dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upscaled dimension data, and a processed test set consisting of test label data and test feature upscaled dimension data, respectively.
[0019] 4. The model training system of item 3, wherein:
[0020] The training feature data and the test feature data are respectively standardized as follows:
[0021] The training feature data and the test feature data are made to have zero mean and unit variance.
[0022] 5. The model training system of item 4, wherein:
[0023] The training feature data and the test feature data are made to have zero mean and unit variance:
[0024] For each feature j, first calculate the mean µ of all samples j and standard deviation σ j :
[0025]
[0026]
[0027] in:
[0028] m is the number of samples;
[0029] x ij is the original value of the i-th sample on the j-th feature;
[0030] µ j is the mean of the jth feature;
[0031] σ j is the standard deviation of the jth feature.
[0032] Use the mean and standard deviation of the features to transform the original data to standardize it:
[0033]
[0034] Where: z ij is the normalized value of the i-th sample on the j-th feature.
[0035] 6. The model training system of item 3, wherein:
[0036] The training feature standardized data and the test feature standardized data are respectively upgraded to:
[0037] The training feature standardized data and the test feature standardized data are mapped to a high-dimensional space using a polynomial kernel function.
[0038] 7. The model training system of item 6, wherein:
[0039] The polynomial kernel function is:
[0040]
[0041] in,
[0042] x i is the feature vector representing the i-th sample;
[0043] x j is the feature vector representing the jth sample;
[0044] <x i ,x j > represents a vector x i and x j The inner product of
[0045] d is the degree of the kernel function.
[0046] 8. The model training system of item 1, wherein:
[0047] The data training module includes:
[0048] A base model component, wherein the base model component includes more than two base training models, each base training model is trained on the processed training set to obtain more than two base prediction models;
[0049] An integrated model component, wherein the integrated model component integrates more than two base prediction models into a final prediction model, and the final prediction model can obtain a final prediction result based on the prediction results of more than two base prediction models.
[0050] 9. The model training system of item 8, wherein:
[0051] The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model.
[0052] 10. The model training system of item 8, wherein:
[0053] The final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models:
[0054] Two or more base prediction models predict independently to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is taken as the final prediction result.
[0055] 11. The model training system of item 1, wherein:
[0056] The model evaluation module includes:
[0057] A prediction component, which can use the final prediction model to obtain a predicted value of the standard equipment detection data from the processed test set;
[0058] An evaluation component compares the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result.
[0059] 12. A model training system as described in item 11, wherein:
[0060] The evaluation component compares the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result:
[0061] The evaluation component calculates the mean square error, determination coefficient, root mean square error and / or mean absolute error between the predicted value of the standard equipment detection data and the true value of the standard equipment detection data as an evaluation result.
[0062] 13. A model training method for volatile organic compound prediction based on machine learning, comprising:
[0063] A data acquisition step, acquiring detection data of target monitoring points and dividing them into a training set and a test set; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data;
[0064] A data preprocessing step, preprocessing the detection data to obtain a processed training set and a processed test set;
[0065] A data training step, training the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data;
[0066] The model evaluation step compares the predicted value of the standard equipment test data obtained by using the final prediction model to predict the test set with the actual value of the standard equipment test data to obtain an evaluation result.
[0067] 14. The model training method of item 13, wherein:
[0068] The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
[0069] 15. The machine learning-based model training method according to item 13, wherein:
[0070] The data preprocessing step includes:
[0071] a data separation sub-step, dividing the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and dividing the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively;
[0072] A feature standardization sub-step of standardizing the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data respectively;
[0073] In the dimension upgrading sub-step, the dimension of the training feature standardized data and the test feature standardized data are upgraded to obtain training feature upgraded dimension data and test feature upgraded dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upgraded dimension data, and a processed test set consisting of test label data and test feature upgraded dimension data, respectively.
[0074] 16. The model training method of item 15, wherein:
[0075] The training feature data and the test feature data are respectively standardized as follows:
[0076] The training feature data and the test feature data are made to have zero mean and unit variance.
[0077] 17. The model training method of item 16, wherein:
[0078] The training feature data and the test feature data are made to have zero mean and unit variance:
[0079] For each feature j, first calculate the mean µ of all samples j and standard deviation σ j :
[0080]
[0081]
[0082] in:
[0083] m is the number of samples;
[0084] x ij is the original value of the i-th sample on the j-th feature;
[0085] µ j is the mean of the jth feature;
[0086] σ j is the standard deviation of the jth feature;
[0087] Use the mean and standard deviation of the features to transform the original data to standardize it:
[0088]
[0089] Where: z ij is the normalized value of the i-th sample on the j-th feature.
[0090] 18. The model training method of item 15, wherein:
[0091] The training feature standardized data and the test feature standardized data are respectively upgraded to:
[0092] The training feature standardized data and the test feature standardized data are mapped to a high-dimensional space using a polynomial kernel function.
[0093] 19. The model training method of item 18, wherein:
[0094] The polynomial kernel function is:
[0095]
[0096] in,
[0097] x i is the feature vector representing the i-th sample;
[0098] x j is the feature vector representing the jth sample;
[0099] <x i ,x j > represents a vector x i and x j The inner product of
[0100] d is the degree of the kernel function.
[0101] 20. The model training method of item 13, wherein:
[0102] The data training step comprises:
[0103] A base model sub-step, training the processed training set by two or more base training models respectively to obtain two or more base prediction models;
[0104] In the integrated model sub-step, the final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models.
[0105] 21. The model training method of item 20, wherein:
[0106] The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model.
[0107] 22. The model training method of item 20, wherein:
[0108] The final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models:
[0109] Two or more base prediction models predict independently to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is taken as the final prediction result.
[0110] 23. The model training method of item 13, wherein:
[0111] The model evaluation step includes:
[0112] In the prediction sub-step, the final prediction model is used to obtain the predicted value of the standard equipment detection data from the processed test set;
[0113] An evaluation sub-step is to compare the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result.
[0114] 24. The model training method of item 23, wherein:
[0115] Compare the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain the evaluation result:
[0116] The mean square error, determination coefficient, mean absolute error and / or mean absolute percentage error between the predicted value of the standard equipment test data and the true value of the standard equipment test data are calculated as the evaluation result.
[0117] 25. A volatile organic compound online early warning device based on machine learning, comprising:
[0118] More than one detection module, different detection modules all include PID sensors, and each detection module detects the concentration of volatile organic compounds at a corresponding target monitoring point to obtain PID sensor detection data;
[0119] A data acquisition module, wherein the data acquisition module acquires the detection data of the PID sensor and the meteorological data of the target monitoring point as prediction data;
[0120] A prediction module, wherein the prediction module is deployed with a final prediction model, and the prediction module obtains a prediction result of a corresponding target monitoring point from the prediction data based on the final prediction model;
[0121] Wherein, the final prediction model is obtained by the model training system described in any one of items 1 to 12 or the model training method described in any one of items 13 to 24.
[0122] 26. The online early warning device for volatile organic compounds based on machine learning as described in item 25, wherein:
[0123] The detection module also includes a meteorological detection device, which detects the meteorological data of the target monitoring point.
[0124] 27. The online warning device for volatile organic compounds based on machine learning as described in item 25 or 26, wherein:
[0125] The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
[0126] Through the above-mentioned machine learning-based model training system, method and online early warning device for volatile organic compound prediction of the present application, a qualified final prediction model can be obtained, and the qualified final prediction model can be used to predict the predicted value (μg / m 3 PID detection has the advantages of shorter monitoring frequency (minutes or even seconds), small size, low price, and easy deployment of large numbers of points. In addition, the predicted value of the standard equipment detection data is the mass concentration of each substance (μg / m 3 ), which can be used to determine the specific detection value of each substance for alarm.
[0127] The above description is only an overview of the technical solution of the present application. In order to make the technical means of the present application clearer and to enable those skilled in the art to implement it according to the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0128] Figure 1 : A schematic diagram of a model training system for volatile organic compound prediction based on machine learning in an embodiment of the present application;
[0129] Figure 2: A schematic diagram of a model training method for volatile organic compound prediction based on machine learning in an embodiment of the present application;
[0130] Figure 3 : Schematic diagram of an online early warning device for volatile organic compounds based on machine learning in an embodiment of the present application. DETAILED DESCRIPTION
[0131] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0132] Application Overview
[0133] As mentioned above, for online VOCs monitoring in industrial parks, standard equipment such as gas chromatography-mass spectrometry (GC-MS) or gas chromatography-flame ionization detector (GC-FID) is generally used for detection. Although it can detect and analyze individual VOC species and the mass concentration in the existing alarm limit unit (μg / m 3 ) to reflect, but such equipment is relatively expensive, has high maintenance costs, and the analysis time is long, resulting in insufficient time representativeness. In addition, PID is also a commonly used VOCs detection solution, but it can only detect the total volume fraction of VOCs, not a single VOC, and cannot detect mass concentration (μg / m 3 ) reflects the detection results of a single VOC, and is used for VOCs over-limit alarm. Since the composition ratio of VOCs is unknown in the PID monitoring solution, technicians generally cannot accurately convert the PID monitoring results (volume fraction) into mass concentration (μg / m 3 ), while the bias of detecting the mass concentration of a single VOC in VOCs and its over-limit alarm is used.
[0134] For online VOCs monitoring in industrial parks, the following technical solution not only has the advantages of shorter monitoring frequency (minutes or even seconds) of PID, small size, low price, and easy deployment of large numbers of points, but also has the prediction results that can be actually used to detect the mass concentration (μg / m 3 ) and the level of over-limit alarm.
[0135] Exemplary Systems
[0136] Figure 1 A machine learning-based model training system for volatile organic compound prediction is illustrated.
[0137] like Figure 1As shown, this embodiment provides a model training system for volatile organic compound prediction based on machine learning, which includes:
[0138] A data acquisition module, wherein the data acquisition module acquires detection data of target monitoring points as training sets and test sets; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data;
[0139] A data preprocessing module, which preprocesses the detection data to obtain a processed training set and a processed test set;
[0140] A data training module, wherein the data training module trains the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data;
[0141] A model evaluation module compares the predicted value of the standard equipment detection data obtained by using the final prediction model to predict the test set and the true value of the standard equipment detection data to obtain an evaluation result.
[0142] Regarding the collection time and quantity of the detection data, it can be the recent historical data of the target monitoring point, such as historical data within a month, a quarter, or a year. In addition, those skilled in the art know that it is necessary to collect an appropriate amount of sample data and an appropriate sample collection time distribution to facilitate subsequent machine learning training and testing.
[0143] In this application, “standard equipment” refers to equipment that can detect the mass concentration (μg / m 3 ) equipment, specifically, gas chromatography-mass spectrometry (GC-MS) and gas chromatography-flame ionization detector (GC-FID) which are commonly used.
[0144] Regarding meteorological data, one or more of wind direction, wind speed, temperature, humidity, and air pressure can be selected. In this embodiment, the meteorological data selected by the inventor includes the wind direction, wind speed, temperature, humidity, and air pressure of the target monitoring point at the same time, which can achieve very good prediction results for a single VOC mass concentration (for a single VOC mass concentration (especially benzene compounds) can have a good prediction effect). Benzene compounds are organic compounds containing a benzene ring structure, also known as aromatic compounds, such as toluene, xylene, etc.
[0145] Through the above-mentioned machine learning-based model training system for volatile organic compound prediction of the present application (hereinafter referred to as "the system of the present application"), the detection data (including training set and test set) of the target monitoring point can be acquired through the data acquisition module for subsequent processing; the detection data (including training set and test set) can be preprocessed (such as standardization, etc.) through the data preprocessing module for subsequent machine learning; through the data training module, the PID sensor detection data and meteorological data in the processed training set are regressed with the standard equipment detection data to construct a final prediction model that predicts the standard equipment detection data from the PID sensor detection data and meteorological data; the model evaluation module can evaluate the final prediction model. Thereby, the qualified final prediction model can be used to predict the predicted value (μg / m 3 PID detection has the advantages of shorter monitoring frequency (minutes or even seconds), small size, low price, and easy deployment of large numbers of points. In addition, the predicted value of the standard equipment detection data is the mass concentration of each substance (μg / m 3 ), which can be used to determine the specific detection value of each substance for alarm, thereby solving the problems of the above-mentioned prior art.
[0146] In addition, by setting the above-mentioned PID sensors at multiple points in the industrial park, the system of this application can separately predict the prediction results of VOCs at multiple points, so as to facilitate timely and comprehensive understanding of the pollution status and trends in the entire park.
[0147] In addition, it should be noted that the correction coefficient is an important parameter of the PID sensor, and the correction coefficient is a measure of the sensitivity of the PID to a specific gas. The system of the present application has a better prediction effect for compounds with lower correction coefficients. Since the correction coefficients of the current main VOCs (such as benzene series, etc.) for PID sensors are relatively low, the system of the present application can be used in most industrial parks. Of course, for special scenarios, technicians in this field can decide whether to adopt the system of the present application based on the industrial distribution of the industrial park. For example, benzene series compounds are the main VOCs in the industrial park where the applicant is conducting experiments, and the system of the present application has a good prediction effect on benzene series compounds.
[0148] Regarding the data preprocessing module, Figure 1 As shown, it includes:
[0149] a data separation component, wherein the data separation component divides the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and divides the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively;
[0150] A feature standardization component, wherein the feature standardization component standardizes the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data;
[0151] A dimension-upgrading component is used to upscale the training feature standardized data and the test feature standardized data to obtain training feature upscaled dimension data and test feature upscaled dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upscaled dimension data, and a processed test set consisting of test label data and test feature upscaled dimension data, respectively.
[0152] Before subsequent training or testing of the detection data used as training sets and test sets, they need to be processed first to facilitate better training and testing results. Specifically, in this solution, in order to first divide the training set and the test set, the training set is divided into training label data (corresponding to standard equipment detection data) and training feature data (corresponding to PID sensor detection data and meteorological data) through the data separation component, and the test set is divided into test label data (corresponding to standard equipment detection data) and test feature data (corresponding to PID sensor detection data and meteorological data), so as to facilitate the construction of a model for predicting standard equipment detection data from PID sensor detection data and meteorological data; in addition, in order to achieve better training results, the training feature data and the test feature data are further standardized through the feature standardization component to accelerate the convergence speed of the machine learning algorithm and improve the performance of the model; and the standardized data (training feature standardized data and test feature standardized data) need to be further dimensioned, and the sample points are projected to a higher dimension to separate the mixed sample points, thereby preventing underfitting and improving the accuracy of the model.
[0153] In this embodiment, respectively standardizing the training feature data and the test feature data specifically means: making the training feature data and the test feature data have zero mean and unit variance, specifically:
[0154] For each feature j, first calculate the mean µ of all samples j and standard deviation σ j :
[0155]
[0156]
[0157] in:
[0158] m is the number of samples;
[0159] x ij is the original value of the i-th sample on the j-th feature;
[0160] µ j is the mean of the jth feature;
[0161] σ j is the standard deviation of the jth feature.
[0162] Use the mean and standard deviation of the features to transform the original data to standardize it:
[0163]
[0164] Where: z ij is the normalized value of the i-th sample on the j-th feature.
[0165] In addition, the dimension of the training feature standardized data and the test feature standardized data are respectively upgraded as follows: the training feature standardized data and the test feature standardized data are mapped to a high-dimensional space using a polynomial kernel function.
[0166] Wherein, the polynomial kernel function is:
[0167]
[0168] in,
[0169] x i is the feature vector representing the i-th sample;
[0170] x j is the feature vector representing the jth sample;
[0171] <x i ,x j > represents a vector x i and x j The inner product of
[0172] d is the degree of the kernel function.
[0173] A polynomial kernel function with a lower degree d may not be sufficient to capture the complex relationships in the data, resulting in insufficient model fitting ability; a polynomial kernel function with a higher degree d provides a stronger fitting ability and can fit more complex data patterns, but usually results in a more complex model and requires careful adjustment to avoid overfitting. In this embodiment, d is selected as 3.
[0174] Regarding the data training module, such as Figure 1 As shown, it includes:
[0175] A base model component, wherein the base model component includes more than two base training models, each base training model is trained on the processed training set to obtain more than two base prediction models;
[0176] An integrated model component, wherein the integrated model component integrates more than two base prediction models into a final prediction model, and the final prediction model can obtain a final prediction result based on the prediction results of more than two base prediction models.
[0177] In this embodiment, the training sets are trained separately by multiple base training models to obtain different base prediction models, so that the relationship between features and labels can be established from different dimensions. After that, the base prediction models are integrated through the integrated model component. Therefore, when making predictions, predictions are actually made jointly through different dimensions to increase the accuracy of the prediction results.
[0178] The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model. The final prediction model can obtain a final prediction result according to the prediction results of two or more base prediction models: the two or more base prediction models independently predict to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is used as the final prediction result.
[0179] In this embodiment, the above four base training models are used for training at the same time, wherein the linear regression model attempts to learn the linear relationship between the data, the gradient boosting regression model captures the nonlinear relationship in the data, the support vector regression model emphasizes the model's prediction accuracy for certain specific areas, and the decision tree regression model adjusts the model's focus on extreme values or important areas. The prediction models trained by these base training models participate in the prediction process together. When making predictions, each model will independently give its own prediction results, and the integrated model component calculates the final prediction output based on these results. Thus, a very good prediction effect can be achieved for a single VOC mass concentration (a very good prediction effect can be achieved for a single VOC mass concentration (especially benzene compounds)).
[0180] Regarding the model evaluation module, such as Figure 1 As shown, it includes:
[0181] A prediction component, which can use the final prediction model to obtain a predicted value of the standard equipment detection data from the processed test set (specifically, the test feature standardized data);
[0182] An evaluation component compares the predicted value of the standard equipment detection data with the true value of the standard equipment detection data (test label data) to obtain an evaluation result.
[0183] Specifically, the evaluation component compares the predicted value of the standard equipment detection data with the true value of the standard equipment detection data to obtain an evaluation result, which means that the evaluation component calculates at least one of the mean square error (MSE), determination coefficient (R² score), root mean square error (RMSE) and / or mean absolute error (MAE) between the model prediction value and the true value of the test set as the evaluation result.
[0184] In this embodiment, the above four evaluation indicators are used simultaneously to obtain the evaluation results, providing model performance evaluation from different angles, which can help understand the prediction ability of the model. Among them, MSE and RMSE can tell the degree to which the predicted value deviates from the true value; the R² score can tell how much data variability the model explains; MAE provides a robust measure of the size of the prediction error; MAPE gives the ratio of the prediction error to the true value. Therefore, it is helpful to obtain a better performance final prediction model to increase the accuracy of the prediction results.
[0185] Exemplary Methods
[0186] Figure 2 A machine learning-based model training method for volatile organic compound prediction is illustrated.
[0187] like Figure 2 As shown, this embodiment provides an online warning method for volatile organic compounds based on machine learning, which includes:
[0188] A data acquisition step, acquiring detection data of target detection points as training sets and test sets; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data;
[0189] A data preprocessing step, preprocessing the detection data to obtain a processed training set and a processed test set;
[0190] A data training step, training the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data;
[0191] The model evaluation step compares the predicted value of the standard equipment test data obtained by using the final prediction model to predict the test set with the actual value of the standard equipment test data to obtain an evaluation result.
[0192] The collection time, quantity and meaning of standard equipment of the test data have been introduced above and will not be repeated here.
[0193] Regarding meteorological data, one or more of wind direction, wind speed, temperature, humidity, and air pressure can be selected. In this embodiment, the meteorological data selected by the inventors include wind direction, wind speed, temperature, humidity, and air pressure of the target monitoring point, which can achieve very good prediction results for a single VOC mass concentration (for a single VOC mass concentration (especially benzene compounds) can have a good prediction effect). Benzene compounds are organic compounds containing a benzene ring structure, also known as aromatic compounds, for VOCs, for example, toluene, xylene, etc.
[0194] Through the above-mentioned machine learning-based model training method for volatile organic compound prediction of the present application (hereinafter referred to as "the method of the present application"), the detection data (including training set and test set) of the target monitoring point can be obtained through the data acquisition step for subsequent processing; the detection data (including training set and test set) is preprocessed (such as standardization, etc.) through the data preprocessing step for subsequent machine learning; through data training, etc., the PID sensor detection data and meteorological data in the processed training set are regressed with the standard equipment detection data to construct a final prediction model that predicts the standard equipment detection data from the PID sensor detection data and meteorological data; model evaluation, etc. can evaluate the final prediction model. Thereby, the qualified final prediction model can be used to predict the predicted value (μg / m 3 PID detection has the advantages of shorter monitoring frequency (minutes or even seconds), small size, low price, and easy deployment of large numbers of points. In addition, the predicted value of the standard equipment detection data is the mass concentration of each substance (μg / m 3 ), which can be used to determine the specific detection value of each substance for alarm, thereby solving the problems of the above-mentioned prior art.
[0195] In addition, the above-mentioned PID sensors are set at multiple points in the industrial park, and the prediction results of VOCs at multiple points can be predicted separately through the method of the present application, so as to facilitate timely and comprehensive understanding of the pollution status and trends in the entire park.
[0196] Regarding data preprocessing steps, such as Figure 2 As shown, it includes:
[0197] a data separation sub-step, dividing the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and dividing the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively;
[0198] A feature standardization sub-step of standardizing the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data respectively;
[0199] In the dimension upgrading sub-step, the dimension of the training feature standardized data and the test feature standardized data are upgraded to obtain training feature upgraded dimension data and test feature upgraded dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upgraded dimension data, and a processed test set consisting of test label data and test feature upgraded dimension data, respectively.
[0200] Before subsequent training or testing of the detection data used as training sets and test sets, they need to be processed first to facilitate better training and testing results. Specifically, in this solution, in order to first divide the training set and the test set, the training set is divided into training label data (corresponding to standard equipment detection data) and training feature data (corresponding to PID sensor detection data and meteorological data) through a data separation step, and the test set is divided into test label data (corresponding to standard equipment detection data) and test feature data (corresponding to PID sensor detection data and meteorological data) to facilitate the construction of a model for predicting standard equipment detection data from PID sensor detection data and meteorological data; in addition, in order to achieve better training results, the training feature data and the test feature data are further standardized through a feature standardization step to accelerate the convergence speed of many machine learning algorithms and improve the performance of the model; and the standardized data (training feature standardized data and test feature standardized data) need to be further dimensioned, and the sample points are projected to a higher dimension to separate the mixed sample points, thereby preventing underfitting and improving the accuracy of the model.
[0201] In this embodiment, respectively standardizing the training feature data and the test feature data specifically means: making the training feature data and the test feature data have zero mean and unit variance, specifically:
[0202] For each feature j, first calculate the mean µ of all samples j and standard deviation σ j :
[0203]
[0204]
[0205] in:
[0206] m is the number of samples;
[0207] x ij is the original value of the i-th sample on the j-th feature;
[0208] µ j is the mean of the jth feature;
[0209] σ j is the standard deviation of the jth feature.
[0210] Use the mean and standard deviation of the features to transform the original data to standardize it:
[0211]
[0212] Where: z ij is the normalized value of the i-th sample on the j-th feature.
[0213] In addition, respectively increasing the dimension of the training feature standardized data and the test feature standardized data specifically refers to: using a polynomial kernel function to map the training feature standardized data and the test feature standardized data to a high-dimensional space.
[0214] Wherein, the polynomial kernel function is:
[0215]
[0216] in,
[0217] x i is the feature vector representing the i-th sample;
[0218] x j is the feature vector representing the jth sample;
[0219] <x i ,x j > represents a vector x i and x j The inner product of
[0220] d is the degree of the kernel function.
[0221] A polynomial kernel function with a lower degree d may not be sufficient to capture the complex relationships in the data, resulting in insufficient model fitting ability; a polynomial kernel function with a higher degree d provides a stronger fitting ability and can fit more complex data patterns, but usually results in a more complex model and requires careful adjustment to avoid overfitting. In this embodiment, d is 3.
[0222] Regarding the data training steps, such as Figure 2 As shown, it includes:
[0223] A base model sub-step, training the processed training set by two or more base training models respectively to obtain two or more base prediction models;
[0224] In the integrated model sub-step, the final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models.
[0225] In this embodiment, the training sets are trained separately by multiple base training models to obtain different base prediction models, so that the relationship between features and labels can be established from different dimensions. After that, the base prediction models are integrated through the integrated model component. Therefore, when making predictions, predictions are actually made jointly through different dimensions to increase the accuracy of the prediction results.
[0226] The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model. The final prediction model can obtain a final prediction result according to the prediction results of two or more base prediction models: the two or more base prediction models independently predict to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is used as the final prediction result.
[0227] In this embodiment, the above four base training models are used for training at the same time, wherein the linear regression model attempts to learn the linear relationship between data, the gradient boosting regression model captures the nonlinear relationship in the data, the support vector regression model emphasizes the prediction accuracy of the model for certain specific areas, and the decision tree regression model adjusts the model's attention to extreme values or important areas. The prediction models trained by these base training models participate in the prediction process together. When making predictions, each model will independently give its own prediction results, and the final prediction output is calculated based on these results. Thus, a very good prediction effect can be achieved for a single VOC mass concentration (a very good prediction effect can be achieved for a single VOC mass concentration (especially benzene compounds)).
[0228] Regarding the model evaluation step, such as Figure 2 As shown, including:
[0229] The prediction sub-step uses the final prediction model to obtain the predicted value of the standard equipment detection data from the processed test set (specifically, the test feature standardized data);
[0230] An evaluation sub-step is to compare the predicted value of the standard equipment detection data with the true value of the standard equipment detection data (test label data) to obtain an evaluation result.
[0231] Specifically, the evaluation step of comparing the predicted value of the standard equipment detection data with the true value of the standard equipment detection data to obtain an evaluation result means that the evaluation step calculates at least one of the mean square error (MSE), determination coefficient (R² score), root mean square error (RMSE) and / or mean absolute error (MAE) between the model prediction value and the true value of the test set as the evaluation result.
[0232] In this embodiment, the above four evaluation indicators are used simultaneously to obtain the evaluation results, providing model performance evaluation from different angles, which can help understand the prediction ability of the model. Among them, MSE and RMSE can tell the degree to which the predicted value deviates from the true value; the R² score can tell how much data variability the model explains; MAE provides a robust measure of the size of the prediction error; MAPE gives the ratio of the prediction error to the true value. Therefore, it is helpful to obtain a better performance final prediction model to increase the accuracy of the prediction results.
[0233] Exemplary Device
[0234] Figure 3 The diagram shows an online early warning device for volatile organic compounds based on machine learning.
[0235] like Figure 3 As shown, this embodiment provides a volatile organic compound online warning device based on machine learning, which includes:
[0236] More than one (e.g., 2, 3, 4, 5, 10, 20, 30, 40, 50 or more) detection modules, different detection modules all include PID sensors, and each detection module detects the concentration of volatile organic compounds at a corresponding target monitoring point to obtain PID sensor detection data;
[0237] A data acquisition module, wherein the data acquisition module acquires the detection data of the PID sensor and the meteorological data of the target monitoring point as prediction data;
[0238] A prediction module, wherein the prediction module is deployed with a final prediction model, and the prediction module obtains a prediction result of a corresponding target monitoring point from the processed prediction data based on the final prediction model;
[0239] Wherein, the final prediction model is obtained by the above-mentioned model training system or the above-mentioned model training method.
[0240] Among them, if the scale of the park is small and the meteorological conditions are simple, only one set of meteorological monitoring equipment can be used. If the scale of the park is large, a meteorological detection device can be set in each detection module, and the meteorological detection device detects the meteorological data of the target monitoring point. Regarding the meteorological data, it can be one or more selected from wind direction, wind speed, temperature, humidity, and air pressure. Preferably, wind direction, wind speed, temperature, humidity, and air pressure are included at the same time.
[0241] Test example
[0242] The inventors of this application adopt the system of the above embodiment (see Figure 1 ) and methods (see Figure 2 ) to obtain the final prediction model, and deploy the final prediction model on the above-mentioned online early warning device (see Figure 3 ), the meteorological data also include wind direction, wind speed, temperature, humidity, and air pressure. The applicant uses standard equipment to monitor the concentration (FID detection concentration) when the standard equipment is used to give an early warning in the industrial park under study. The PID sensor deployed in the online early warning device obtains the detection data and meteorological data, and further predicts the predicted concentration of the standard setting detection value. The specific statistics are shown in Table 1.
[0243] Among them, the standard equipment is the SynspecGC955 series 615 / 815 toxic and hazardous volatile organic compound analyzer (referred to as "GC-FID equipment") from SYNSPEC of the Netherlands; the PID sensor is the MiniPID 2 from ION Science of the UK.
[0244] Table 1: Statistics of standard equipment warning status and online warning device prediction status.
[0245]
[0246] As can be seen from Table 1, in one month, the standard equipment gave five alarms for a single VOC. Among them, the online early warning device of this application can be used for alarm when the predicted concentration reaches / exceeds the species alarm concentration in the first four times; even if the species alarm concentration is not reached in the fifth time, the predicted concentration is extremely close to the standard equipment. The staff can know that the corresponding single VOC concentration of the target monitoring point corresponding to the PID sensor is at a high level based on the predicted concentration, which provides a reliable basis for the staff to handle it later.
[0247] Although the embodiments of the present application are described above, the present application is not limited to the above specific embodiments and application fields, and the above specific embodiments are merely illustrative and instructive, rather than restrictive. A person of ordinary skill in the art can make many forms under the guidance of this specification and without departing from the scope of protection of the claims of the present application, all of which belong to the scope of protection claimed in the present application.
Claims
1. A model training system for volatile organic compound prediction based on machine learning, wherein: include: A data acquisition module, wherein the data acquisition module acquires detection data of target monitoring points as training sets and test sets; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data; A data preprocessing module, which preprocesses the detection data to obtain a processed training set and a processed test set; A data training module, wherein the data training module trains the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data; A model evaluation module compares the predicted value of the standard equipment detection data obtained by using the final prediction model to predict the test set and the true value of the standard equipment detection data to obtain an evaluation result.
2. The model training system according to claim 1, wherein: The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
3. The model training system according to claim 1, wherein: The data preprocessing module comprises: a data separation component, wherein the data separation component divides the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and divides the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively; A feature standardization component, wherein the feature standardization component standardizes the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data; A dimension-upgrading component is used to upscale the training feature standardized data and the test feature standardized data to obtain training feature upscaled dimension data and test feature upscaled dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upscaled dimension data, and a processed test set consisting of test label data and test feature upscaled dimension data, respectively.
4. The model training system according to claim 3, wherein: The training feature data and the test feature data are respectively standardized as follows: The training feature data and the test feature data are made to have zero mean and unit variance.
5. The model training system according to claim 4, wherein: The training feature data and the test feature data are made to have zero mean and unit variance: For each feature j , First calculate the mean µ of all samples j and standard deviation σ j : ; ; in: m is the number of samples; x ij is the original value of the i-th sample on the j-th feature; µ j is the mean of the jth feature; σ j is the standard deviation of the jth feature; Use the mean and standard deviation of the features to transform the original data to standardize it: ; Where: z ij is the normalized value of the i-th sample on the j-th feature.
6. The model training system according to claim 3, wherein: The training feature standardized data and the test feature standardized data are respectively upgraded to: The training feature standardized data and the test feature standardized data are mapped to a high-dimensional space using a polynomial kernel function.
7. The model training system according to claim 6, wherein: The polynomial kernel function is: ; in, x i is the feature vector representing the i-th sample; x j is the feature vector representing the jth sample; <x i ,x j > represents a vector x i and x j The inner product of d is the degree of the kernel function.
8. The model training system according to claim 1, wherein: The data training module includes: A base model component, wherein the base model component includes more than two base training models, each base training model is trained on the processed training set to obtain more than two base prediction models; An integrated model component, wherein the integrated model component integrates more than two base prediction models into a final prediction model, and the final prediction model can obtain a final prediction result based on the prediction results of more than two base prediction models.
9. The model training system according to claim 8, wherein: The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model.
10. The model training system according to claim 8, wherein: The final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models: Two or more base prediction models predict independently to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is taken as the final prediction result.
11. The model training system according to claim 1, wherein: The model evaluation module includes: A prediction component, which can use the final prediction model to obtain a predicted value of the standard equipment detection data from the processed test set; An evaluation component compares the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result.
12. The model training system according to claim 11, wherein: The evaluation component compares the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result: The evaluation component calculates the mean square error, determination coefficient, root mean square error and / or mean absolute error between the predicted value of the standard equipment detection data and the true value of the standard equipment detection data as an evaluation result.
13. A model training method for volatile organic compound prediction based on machine learning, wherein: include: A data acquisition step, acquiring detection data of target monitoring points and dividing them into a training set and a test set; wherein the detection data includes standard equipment detection data of volatile organic compound concentration, PID sensor detection data of volatile organic compound concentration, and meteorological data; A data preprocessing step, preprocessing the detection data to obtain a processed training set and a processed test set; A data training step, training the processed training set to obtain a final prediction model for predicting standard equipment detection data based on PID sensor detection data and meteorological data; The model evaluation step compares the predicted value of the standard equipment test data obtained by using the final prediction model to predict the test set with the actual value of the standard equipment test data to obtain an evaluation result.
14. The model training method according to claim 13, wherein: The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
15. The model training method according to claim 13, wherein: The data preprocessing step includes: a data separation sub-step, dividing the standard equipment detection data of the training set and the standard equipment detection data of the test set into training label data and test label data, respectively, and dividing the PID sensor detection data and the meteorological data of the training set and the PID sensor detection data and the meteorological data of the test set into training feature data and test feature data, respectively; A feature standardization sub-step of standardizing the training feature data and the test feature data to obtain training feature standardized data and test feature standardized data respectively; In the dimension upgrading sub-step, the dimension of the training feature standardized data and the test feature standardized data are upgraded to obtain training feature upgraded dimension data and test feature upgraded dimension data, respectively, so as to obtain a processed training set consisting of training label data and training feature upgraded dimension data, and a processed test set consisting of test label data and test feature upgraded dimension data, respectively.
16. The model training method according to claim 15, wherein: The training feature data and the test feature data are respectively standardized as follows: The training feature data and the test feature data are made to have zero mean and unit variance.
17. The model training method according to claim 16, wherein: The training feature data and the test feature data are made to have zero mean and unit variance: For each feature j , First calculate the mean µ of all samples j and standard deviation σ j : ; ; in: m is the number of samples; x ij is the original value of the i-th sample on the j-th feature; µ j is the mean of the jth feature; σ j is the standard deviation of the jth feature; Use the mean and standard deviation of the features to transform the original data to standardize it: ; Where: z ij is the normalized value of the i-th sample on the j-th feature.
18. The model training method according to claim 15, wherein: The training feature standardized data and the test feature standardized data are respectively upgraded to: The training feature standardized data and the test feature standardized data are mapped to a high-dimensional space using a polynomial kernel function.
19. The model training method according to claim 18, wherein: The polynomial kernel function is: ; in, x i is the feature vector representing the i-th sample; x j is the feature vector representing the jth sample; <x i ,x j > represents a vector x i and x j The inner product of d is the degree of the kernel function.
20. The model training method according to claim 13, wherein: The data training step comprises: A base model sub-step, training the processed training set by two or more base training models respectively to obtain two or more base prediction models; In the integrated model sub-step, the final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models.
21. The model training method according to claim 20, wherein: The base training model is selected from two or more models of a linear regression model, a gradient boosting regression model, a support vector regression model, and a decision tree regression model.
22. The model training method according to claim 20, wherein: The final prediction model can obtain the final prediction result according to the prediction results of more than two base prediction models: Two or more base prediction models predict independently to obtain prediction results respectively, and the majority voting result or the average prediction result of the prediction results of each base prediction model is taken as the final prediction result.
23. The model training method according to claim 13, wherein: The model evaluation step includes: In the prediction sub-step, the final prediction model is used to obtain the predicted value of the standard equipment detection data from the processed test set; An evaluation sub-step is to compare the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain an evaluation result.
24. The model training method according to claim 23, wherein: Compare the predicted value of the standard equipment detection data with the actual value of the standard equipment detection data to obtain the evaluation result: The mean square error, determination coefficient, mean absolute error and / or mean absolute percentage error between the predicted value of the standard equipment test data and the true value of the standard equipment test data are calculated as the evaluation result.
25. A volatile organic compound online early warning device based on machine learning, wherein: include: More than one detection module, different detection modules all include PID sensors, and each detection module detects the concentration of volatile organic compounds at a corresponding target monitoring point to obtain PID sensor detection data; A data acquisition module, wherein the data acquisition module acquires the detection data of the PID sensor and the meteorological data of the target monitoring point as prediction data; A prediction module, wherein the prediction module is deployed with a final prediction model, and the prediction module obtains a prediction result of a corresponding target monitoring point from the prediction data based on the final prediction model; Wherein, the final prediction model is obtained by the model training system described in any one of claims 1 to 12 or the model training method described in any one of claims 13 to 24.
26. The volatile organic compound online early warning device based on machine learning as claimed in claim 25, wherein: The detection module also includes a meteorological detection device, which detects the meteorological data of the target monitoring point.
27. The volatile organic compound online early warning device based on machine learning according to claim 25 or 26, wherein: The meteorological data is selected from one or more of wind direction, wind speed, temperature, humidity and air pressure.
Citation Information
Patent Citations
Method for detecting toxic and harmful gas in chemical industrial park
CN112229952A
Volatile organic compound on-line monitoring optimization system and method thereof
CN116256457A
Construction method and device of equipment operation condition prediction model and storage medium
CN117076984A
Indoor TVOC concentration real-time monitoring system based on wireless sensor network
CN117129556A
Construction method of ultra-fine tailing paste thixotropic rheological parameter prediction model
CN118335266A