Method for predicting combustion emission of solid fuel at different altitudes
By constructing a random forest model and using solid fuel combustion emission data at altitudes, the problem in the prior art that solid fuel combustion emissions at different altitudes cannot be accurately predicted, achieving higher prediction accuracy and reliability.
Patent Information
- Application Number
- CN202510207193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to accurately predict solid fuel combustion emissions at different altitudes, and fails to fully consider the impact of altitude and its associated environmental factors on the combustion process.
By collecting and preprocessing solid fuel combustion emission data at different altitudes, key factors are determined and characteristic parameters are extracted, random forest models are constructed for parameter tuning and model training, model performance is verified and prediction is made.
It realizes accurate prediction of solid fuel combustion emissions at different altitudes, improves the accuracy and reliability of the predictions, and provides technical support for energy utilization and air quality improvement in different altitudes around the world.
Smart Images

Figure CN120046806A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of environmental science and air pollution control, and particularly relates to a method for predicting solid fuel combustion emissions at different altitudes. Background Art
[0002] Solid fuels mainly include coal (such as lump coal and honeycomb coal) and biomass (such as firewood and crop straw, etc.), and are the main living energy sources for rural residents in China for cooking and heating activities. Globally, there are still about 3 billion people relying on solid fuels for cooking and heating. However, the combustion efficiency of solid fuels is generally low, resulting in a large amount of pollutants being generated during the combustion process. These pollutants not only have a serious impact on outdoor air quality, leading to the formation of haze, but may also cause health problems such as cardiovascular and respiratory diseases, seriously threatening people's lives and health.
[0003] Traditional methods for predicting solid fuel combustion emissions are mostly constructed based on combustion theoretical models or empirical formulas under ideal conditions. These methods often assume that the combustion environment is under idealized conditions such as standard atmospheric pressure, normal temperature and humidity, etc., and do not fully consider the profound impact of altitude and associated complex environmental factor changes on the combustion process, and cannot meet the demand for accurate prediction of combustion emissions under diverse environmental conditions in practical applications. Summary of the Invention
[0004] In order to solve the technical problems existing in the background art, the present invention aims to provide a method for predicting solid fuel combustion emissions at different altitudes.
[0005] In order to solve the technical problems, the technical solution of the present invention is:
[0006] A method for predicting solid fuel combustion emissions at different altitudes, the method comprising:
[0007] S1: Collect the original emission data during the combustion process of different types of solid fuels, and perform preprocessing, including data cleaning and data standardization;
[0008] S2: Based on the preprocessed data, determine the key factors affecting solid fuel combustion emissions, and extract effective characteristic parameters, that is, a refined characteristic subset, as the input of the model;
[0009] S3: Construct a random forest model for parameter tuning and model training to obtain a trained model;
[0010] S4: Based on the preprocessed data, use the test set in the feature subset to verify the trained model, evaluate the prediction performance and generalization ability of the model, and according to the verification results, select the parameters with the best performance as the parameters of the final solid fuel combustion emission prediction model; use the trained model to predict the solid fuel combustion emissions at different altitudes.
[0011] Furthermore, the emission data includes: data related to solid fuel combustion at different altitudes, fuel types, fuel characteristics, and stove types.
[0012] Furthermore, the preprocessing of the data includes:
[0013] Data cleaning, removing outliers and filling in missing values; data standardization, standardizing the cleaned data, and the standardization formula is as follows:
[0014]
[0015] where x is the original data, μ is the mean of the data, σ is the standard deviation of the data, and the standardized data follows a standard normal distribution with a mean of 0 and a variance of 1.
[0016] Furthermore, the extraction of the feature parameters includes:
[0017] Feature selection strategy: Combine correlation analysis, principal component analysis PCA and domain knowledge to screen key features; generate interaction features to capture non-linear relationships;
[0018] Automated feature selection: Use LASSO regression or recursive feature elimination RFE to extract important features; analyze the importance of features and eliminate redundant features;
[0019] Verification of feature selection effect: Compare the prediction performance before and after feature selection in a simple model to verify the effectiveness of the selection.
[0020] Furthermore, select and construct a machine learning model, and perform parameter tuning and model training on the selected machine learning algorithm to obtain a trained model, including:
[0021] Model training: Use the random forest algorithm to train the model and fit the reduced feature subset;
[0022] Hyperparameter optimization: Use grid search or Bayesian optimization to find the optimal hyperparameters; use cross-validation to evaluate the stability of the parameter combination;
[0023] Early stopping mechanism: Introduce early stopping during training to prevent overfitting;
[0024] Among them, select the random forest algorithm as the machine learning model and set the basic parameters of the random forest:
[0025] n_estimators: The number of trees;
[0026] max_depth: The maximum depth of the tree. Branches exceeding the maximum depth will be pruned;
[0027] min_samples_split: The minimum number of samples required for internal node splitting;
[0028] min_samples_leaf: The minimum number of samples in a leaf node;
[0029] max_features: The number of features considered for the best split.
[0030] Common parameter adjustment ranges:
[0031] n_estimators: The number of trees, ranging from 50 to 200, with a step size of 1;
[0032] max_depth: The maximum depth of the tree, ranging from 5 to 30, with a step size of 5;
[0033] min_samples_split: The minimum number of samples required for internal node splitting, set between 2 and 20;
[0034] min_samples_leaf: The minimum number of samples in a leaf node, set between 1 and 5;
[0035] max_features: The number of features considered for the best split, set to'sqrt', 'log2', or a specific number of features.
[0036] Furthermore, the model evaluation and validation specifically include:
[0037] Build a model using the optimal parameters and make predictions on the test set. Compare the predicted values with the true values and calculate various evaluation metrics to comprehensively evaluate the trained model, including: mean squared error, mean absolute error, and coefficient of determination. The calculation formulas are as follows:
[0038]
[0039] where n represents the number of samples in the test set, is the predicted value, yi is the true value, is the average of the true values.
[0040] Furthermore, use the trained model to predict the emissions of solid fuel combustion at different altitudes, specifically including:
[0041] Prediction framework: Taking fuel characteristics, stove type, and environmental conditions as inputs, and combining with a trained random forest model to predict emissions;
[0042] Prediction confidence interval: Calculating the confidence interval of the predicted value to improve the reliability of the prediction result;
[0043] Visualization output: Constructing an interactive visualization interface to display the emission change trends under different conditions;
[0044] Real-time prediction system: Combining Internet of Things technology to achieve online monitoring and real-time prediction, and outputting the emission concentrations of PM 2.5 、CO、CO 2 pollutants.
[0045] Compared with the prior art, the advantages of the present invention are as follows:
[0046] Based on a large amount of measured data in different altitude regions, and relying on the excellent highly non-linear fitting ability of the machine learning model, the present invention accurately deconstructs the internal logical chain between multi-factors such as altitude, fuel type, and stove type and solid fuel combustion emissions, and constructs a set of solid fuel combustion emission prediction systems with strong universality and excellent accuracy, laying a solid foundation for the efficient utilization of energy, the improvement of air quality, and the ecological environment protection in different altitude regions around the world. Brief Description of the Drawings
[0047] Figure 1 、The main flowchart of a method for predicting solid fuel combustion emissions at different altitudes according to the present invention;
[0048] Figure 2 、Prediction result data graph;
[0049] Figure 3 、Original dataset graph;
[0050] Figure 4 、Preprocessed dataset graph. Detailed Embodiments
[0051] The following describes the specific embodiments of the present invention in conjunction with the embodiments:
[0052] It should be noted that the structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0053] Meanwhile, terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear description and do not limit the scope of implementation of the present invention. Changes or adjustments to their relative relationships shall be regarded as the scope of implementation of the present invention when there is no substantial change in the technical content.
[0054] Example 1:
[0055] As Figure 1 shown:
[0056] A method for predicting solid fuel combustion emissions at different altitudes based on machine learning, which specifically includes the following steps:
[0057] S1. Data collection: Collect emission data during the combustion of different types of solid fuels.
[0058] S2. Data preprocessing: Preprocess the collected data, including data cleaning and data standardization.
[0059] S3. Feature selection: Deeply analyze the key factors affecting solid fuel combustion emissions, and extract effective feature parameters as the input of the model.
[0060] S4. Selection and construction of machine learning model: Comprehensively investigate various machine learning algorithms applicable to this problem, and analyze their advantages, disadvantages, and applicability in dealing with complex non-linear relationships and small sample data, etc.
[0061] S5. Model training and optimization: Optimize the parameters of the selected machine learning algorithm and train the model.
[0062] S6. Model evaluation and verification: Use the test set to verify the trained model, evaluate the prediction performance and generalization ability of the model, and select the parameters with the best performance as the parameters of the final solid fuel combustion emission prediction model according to the verification results.
[0063] S7. Emission prediction: Use the trained model to achieve rapid and accurate prediction of solid fuel combustion emissions at different altitudes.
[0064] The data collection described in S1 specifically includes the following steps:
[0065] S11. Collect data related to solid fuel combustion emissions, including different altitudes, fuel types, fuel characteristics (such as the content of carbon, hydrogen, sulfur, and nitrogen), and stove types, etc.
[0066] Including relevant data such as different altitudes, fuel types, fuel characteristics (such as the contents of carbon, hydrogen, sulfur, and nitrogen), stove types, etc.; adding emission data under extreme conditions (such as high altitude and low oxygen, high humidity, etc.); dynamically collecting parameters during the combustion process (such as temperature, combustion time) to capture the dynamic laws of combustion; including emission characteristics in different geographical regions (such as plateaus, plains, and hills).
[0067] The specific data preprocessing in S2 includes the following steps:
[0068] S21. Data cleaning: For outliers, according to the 3σ rule, for continuous numerical features, calculate their mean and standard deviation, mark the data points that deviate from the mean by more than 3 times the standard deviation as outliers, and delete or correct them using a reasonable interpolation method in combination with the actual situation. For missing values, first count the proportion of missing values in each feature. For features with a missing value proportion less than 5%, fill them with the mean or median; if the proportion of missing values is relatively large, then use the random forest for prediction filling in combination with the correlation between this feature and other relevant features.
[0069] S22. Data standardization: Standardize the cleaned data, and the standardization formula is as follows:
[0070]
[0071] Among them, x is the original data, μ is the mean of the data, and σ is the standard deviation of the data. The standardized data follows a standard normal distribution with a mean of 0 and a variance of 1.
[0072] The specific feature selection in S3 includes the following steps:
[0073] S31. Feature selection strategy: Use correlation analysis to calculate the Pearson correlation coefficient between each feature and the solid fuel combustion emission target variable, and initially determine that features with an absolute value greater than 0.5 have strong correlations. At the same time, use the principal component analysis (PCA) technique to project the original high-dimensional feature space onto a low-dimensional principal component space, and extract the key principal components as new features on the premise of retaining most of the data variance (such as more than 90%). Subsequently, combine the knowledge of domain experts to conduct a secondary screening of the features selected through correlation analysis and PCA. Finally, determine a set of concise and efficient feature subsets for model training.
[0074] S32. Feature selection effect verification: Before and after feature selection, use the same simple machine learning model (such as a linear regression model) to train the data, and compare the performance of the model on the test set. If the model after feature selection shows significant improvement (such as a reduction of more than 10%) in evaluation metrics such as mean squared error (MSE) and mean absolute error (MAE), it proves that feature selection is effective, reducing the computational complexity and overfitting risk of the model, while improving the prediction accuracy of the model.
[0075] The selection and construction of the machine learning model described in S4 specifically include the following steps:
[0076] S41. Machine learning model selection: By researching and comparing various machine learning algorithms, analyzing their advantages, disadvantages, and applicability in dealing with complex non-linear relationships and small sample data, it is found that the random forest algorithm is suitable for various data scales, especially in cases where the data may have noise, missing values, or contain multiple types of features (such as continuous and discrete). The collected emission data usually contains various types of parameters, and the data may be disturbed by various environmental factors. Therefore, for the problem of solid fuel combustion emission prediction, the random forest is a good choice.
[0077] S42. Set the basic parameters of the random forest:
[0078] n_estimators: The number of trees;
[0079] max_depth: The maximum depth of the tree, and the branches exceeding the maximum depth will be cut off;
[0080] min_samples_split: The minimum number of samples required for internal node splitting;
[0081] min_samples_leaf: The minimum number of samples in the leaf node;
[0082] The model training and optimization described in S5 specifically include the following steps:
[0083] S51. Divide the data set into a training set and a test set according to a ratio of 8:2. To ensure the randomness and repeatability of the division, set the random seed (random_state = 42). Use the training set to train the random forest model;
[0084] S52. Use grid search combined with 5-fold cross-validation to find the optimal parameter combination:
[0085] S53. Common parameter adjustment ranges:
[0086] n_estimators: The number of trees, from 50 to 200, with a step size of 1;
[0087] max_depth: The maximum depth of the tree, ranging from 5 to 30, with a step size of 5;
[0088] min_samples_split: The minimum number of samples required for splitting an internal node, set between 2 and 20;
[0089] min_samples_leaf: The minimum number of samples in a leaf node, set between 1 and 5;
[0090] The model evaluation and verification described in S6 specifically include the following steps:
[0091] S61. Construct a model using the optimal parameters and make predictions on the test set. Compare the predicted values with the true values and calculate various evaluation metrics to comprehensively evaluate the trained model, such as Mean Square Error (MSE), Mean Absolute Error (MAE), and Coefficient of determination (R 2 ), and the calculation formulas are as follows:
[0092]
[0093] where n represents the number of samples in the test set, is the predicted value, y i is the true value, is the average value of the true values.
[0094] The emission prediction described in S7 specifically includes the following steps:
[0095] S71. Data input: Use the re - collected data as the dataset and divide it into a training set and a test set;
[0096] S72. Model training: Use the training set data to train a random forest regression model;
[0097] S73. Model evaluation: Use the test set data to evaluate the model performance;
[0098] S74. Model prediction and result output: Input the solid fuel combustion emission data to be predicted into the trained and optimized model. The model, based on the learned complex feature relationships and patterns, quickly outputs the corresponding prediction results, such as PM 2.5 emission concentration, other pollutant emissions, etc.
[0099] Example 2:
[0100] This example is a specific experimental illustration of Example 1, where the original dataset, such as Figure 3As shown; the preprocessed dataset, such as Figure 4 shown.
[0101] Model training:
[0102] Optimal random forest parameters:
[0103] 'max_depth': None
[0104] 'min_samples_leaf': 2
[0105] 'min_samples_split': 2
[0106] 'n_estimators': 100
[0107] Prediction results, such as Figure 2 shown.
[0108] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0109] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in multiple blocks.
[0112] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.
[0113] Many other changes and modifications can be made without departing from the concept and scope of the present invention. It should be understood that the present invention is not limited to the specific embodiments, and the scope of the present invention is defined by the appended claims.
Claims
1. A method for predicting solid fuel combustion emissions at different altitudes, characterized in that: The method comprises: S1: Collect raw emission data from the combustion of different types of solid fuels and perform preprocessing, including data cleaning and data standardization; S2: Based on the preprocessed data, determine the key factors affecting solid fuel combustion emissions and extract effective feature parameters, i.e., a simplified feature subset, as the input of the model; S3: Build a random forest model to perform parameter tuning and model training to obtain a trained model; S4: Based on the preprocessed data, the trained model is verified using the test set in the feature subset to evaluate the model's prediction performance and generalization ability. According to the verification results, the parameters with the best performance are selected as the final solid fuel combustion emission prediction model parameters; the trained model is used to predict solid fuel combustion emissions at different altitudes.
2. A method for predicting solid fuel combustion emissions at different altitudes according to claim 1, characterized in that: The emission data include data related to solid fuel combustion emissions at different altitudes, fuel types, fuel characteristics and stove types.
3. The method for predicting solid fuel combustion emissions at different altitudes according to claim 1, characterized in that: The data is preprocessed, including: Data cleaning, remove outliers and fill missing values; data standardization, standardize the cleaned data, the standardization formula is as follows: Among them, x is the original data, μ is the mean of the data, σ is the standard deviation of the data, and the standardized data obeys the standard normal distribution with a mean of 0 and a variance of 1.
4. The method for predicting solid fuel combustion emissions at different altitudes according to claim 1, characterized in that: The feature parameter extraction includes: Feature selection strategy: Combine correlation analysis, principal component analysis (PCA) and domain knowledge to screen key features; generate interactive features to capture nonlinear relationships; Automated feature selection: extract important features using LASSO regression or recursive feature elimination (RFE); analyze feature importance and eliminate redundant features; Verification of feature selection effect: Compare the prediction performance before and after feature selection in a simple model to verify the effectiveness of the selection.
5. The method for predicting solid fuel combustion emissions at different altitudes according to claim 1, characterized in that: The step of selecting and constructing a machine learning model, and performing parameter tuning and model training on the selected machine learning algorithm to obtain a trained model includes: Model training: Use the random forest algorithm to train the model and fit a reduced feature subset; Hyperparameter optimization: Use grid search or Bayesian optimization to find the optimal hyperparameters; use cross-validation to evaluate the stability of parameter combinations; Early stopping mechanism: early stopping is introduced during training to prevent overfitting; Among them, select the random forest algorithm as the machine learning model and set the basic parameters of the random forest: n_estimators: the number of trees; max_depth: The maximum depth of the tree. Branches exceeding the maximum depth will be pruned. min_samples_split: the minimum number of samples required for internal node splitting; min_samples_leaf: minimum number of samples for leaf nodes; max_features: The number of features to consider for the best split. Common parameter adjustment range: n_estimators: the number of trees, from 50 to 200, with a step size of 1; max_depth: the maximum depth of the tree, from 5 to 30, with a step size of 5; min_samples_split: The minimum number of samples required for internal node splitting, set between 2 and 20; min_samples_leaf: minimum number of samples for leaf nodes, set between 1 and 5; max_features: The number of features to consider for the best split, set to 'sqrt', 'log2', or a specific number of features.
6. The method for predicting solid fuel combustion emissions at different altitudes based on machine learning according to claim 5, characterized in that: Model evaluation and verification specifically include: Use the optimal parameters to build a model and make predictions on the test set. Compare the predicted values with the true values, and calculate various evaluation indicators to comprehensively evaluate the trained model, including: mean square error, mean absolute error, and determination coefficient. The calculation formula is as follows: Where n represents the number of test set samples. is the predicted value, y i is the true value, is the average of the true values.
7. The method for predicting solid fuel combustion emissions at different altitudes based on machine learning according to claim 1, characterized in that: Use the trained model to predict emissions from solid fuel combustion at different altitudes, including: Prediction framework: Fuel characteristics, stove type, and environmental conditions are used as inputs and combined with the trained random forest model to predict emissions; Prediction credible interval: Calculate the credible interval of the predicted value to improve the reliability of the prediction results; Visualization output: Build an interactive visualization interface to show emission trends under different conditions; Real-time prediction system: Combined with Internet of Things technology, it realizes online monitoring and real-time prediction, and outputs PM 2.5 , CO, and CO2 pollutant emission concentrations.
Citation Information
Cited By
Altitude-based energy determination method
CN120911857A