Sewage treatment plant carbon accounting optimization prediction method based on machine learning
Through machine learning methods, using XGBoost regressor and scientific data processing, the accurate accounting of carbon emissions in sewage treatment plants is solved, and the accurate prediction and optimization management of carbon emissions are achieved.
Patent Information
- Application Number
- CN202510502600.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technology cannot accurately calculate the carbon emissions of sewage treatment plants, ignore the impact of incoming water quality characteristics, sewage treatment process differences and equipment parameter changes on carbon emissions, and the prediction method cannot effectively deal with complex nonlinear relationships.
Using machine learning methods, through data collection, preprocessing, feature analysis and model training, XGBoost regressor is used to predict carbon emissions, combined with dynamic and static data, reasonable hyperparameters are set, and model performance is evaluated through mean square error and average absolute error.
It has achieved accurate prediction of carbon emissions of sewage treatment plants, provided scientific and highly adaptable carbon accounting optimization solutions, and supported the low-carbon operation of sewage treatment plants.
Smart Images

Figure CN120494126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon accounting optimization and prediction, and in particular to a carbon accounting optimization and prediction method for sewage treatment plants based on machine learning. Background Art
[0002] As global attention to climate change grows, the accurate accounting and effective control of carbon emissions have become crucial issues for all industries. As a key component of urban infrastructure, sewage treatment plants involve significant energy consumption and chemical reactions during the wastewater purification process, inevitably generating various greenhouse gas emissions. Their carbon emissions cannot be ignored.
[0003] Traditionally, sewage treatment plants have focused primarily on ensuring that sewage meets treatment standards, and have not paid sufficient attention to the accounting and management of carbon emissions. Early carbon accounting methods were often crude, simply considering some obvious sources of carbon emissions, such as carbon dioxide emissions from electricity consumption, but ignoring many complex and critical factors in the sewage treatment process. For example, different influent water quality characteristics, such as fluctuations in influent flow, and differences in COD (chemical oxygen demand), TN (total nitrogen), and TP (total phosphorus) concentrations in sewage, will significantly affect resource consumption and gas generation during the treatment process, but previous accounting methods have failed to accurately consider these dynamically changing factors.
[0004] Furthermore, wastewater treatment processes are diverse, ranging from activated sludge to biofilm processes. Each process exhibits significant differences in treatment processes and microbial metabolic mechanisms, leading to widely varying carbon emissions. Traditional accounting methods struggle to distinguish and accurately assess the specific characteristics of these processes. Furthermore, variations in equipment power within wastewater treatment plants and changes in the operating parameters of energy-saving and emission-reduction facilities such as methane recovery units have a significant impact on overall carbon emissions, yet these factors have not been fully reflected in previous accounting methods.
[0005] Early prediction methods, often based on simple empirical formulas or linear regression models, were unable to effectively address the complex nonlinear relationships and interactions between numerous variables in the wastewater treatment process. With the continued expansion of wastewater treatment plants, increasing treatment requirements, and increasingly stringent environmental protection policies, there is an urgent need for a more scientific, accurate, and adaptable approach to optimize carbon accounting predictions for wastewater treatment plants. The rise of machine learning technology has provided a new approach to addressing this challenge. Summary of the Invention
[0006] The purpose of the present invention is to provide a carbon accounting optimization prediction method for sewage treatment plants based on machine learning, which solves the technical problems raised in the background technology.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A method for optimizing and predicting carbon accounting in a sewage treatment plant based on machine learning includes the following steps:
[0009] Data collection: Collect the operation data of the sewage treatment plant during the observation period.
[0010] Preprocessing: Eliminate outliers in the operating data based on the 3σ principle;
[0011] Characterization analysis: Determine the carbon emissions of wastewater treatment plants based on operational data;
[0012] Model analysis: build and train an XGBoost regressor model;
[0013] Calculation and prediction: Input the preprocessed operating data into the trained XGBoost regressor and obtain the output carbon emission prediction value.
[0014] As a further solution of the present invention: the operation data is divided into dynamic data and static data; the dynamic data includes water flow, COD, TN, TP concentration, energy consumption, and chemical consumption;
[0015] Static data include treatment process type, equipment power, and methane recovery unit parameters.
[0016] As a further solution of the present invention: the pretreatment method is as follows:
[0017] First calculate the mean μ of various running data i and standard deviation σ i ;
[0018] Wherein, i=1, 2, ..., n, n represents the number of types of operating data;
[0019] When x i <μ i -3σ i or x i >μ i +3σ i When , the corresponding operating data is regarded as an abnormal value; where x i is the i-th running data.
[0020] As a further solution of the present invention: the feature analysis method is as follows:
[0021] Methane emissions calculation:
[0022] Where, E CH4 is the methane emission, Q is the water inlet flow, C COD is COD concentration, EF CH4 is the preset methane emission factor, HSCH4 is the preset methane recovery efficiency;
[0023] Calculation of Nitrous Oxide Emissions:
[0024] Where, E N2O is the nitrous oxide emission, C TN is the TN concentration, EF N2O is the nitrous oxide emission factor;
[0025] Calculation of carbon dioxide emissions from electricity consumption: E DL =∑(P j ×T j )×EF DW ;
[0026] Where, E DL is the carbon dioxide emissions from electricity consumption, P j is the equipment power of different power equipment; T j is the operating time, j represents different power equipment; EF DW is the preset grid emission factor;
[0027] Calculation of carbon dioxide emissions from pharmaceutical consumption: E YH =∑(M k ×EF k )
[0028] Where, E DL is the carbon dioxide emissions from electricity consumption, where M k is the dosage of the drug; EF k The emission factor for the preset pharmaceutical production; k represents different pharmaceuticals;
[0029] Total carbon emissions calculation:
[0030] Where, E1 total carbon emissions, GWP CH4 and GWP N2O are the global warming potentials of methane and nitrous oxide, respectively.
[0031] As a further solution of the present invention: Model analysis:
[0032] Step 1. Model selection:
[0033] Use XGBoost regressor for model training;
[0034] Step 2: Data partitioning:
[0035] Divide the historical specified period into several specified time periods, where the specified time period covers multiple standard time periods, and then obtain the operating data of multiple standard time periods within the same specified time period and obtain the total carbon emissions through feature analysis;
[0036] Then, the data sets are divided into multiple standard periods in the same specified time period in chronological order;
[0037] The first 70% of the time series data is used as the training set, the middle 15% as the validation set, and the last 15% as the test set;
[0038] Step 3, model training:
[0039] Based on the training set, the mean square error is used as the loss function;
[0040] The loss function calculation formula is:
[0041] Where MSE is the mean square error as the loss function, m is the number of data points in the training set, and F0 g is the total carbon emissions of the g-th data point in the training set, F1 g is the total carbon emissions predicted by the model for the g-th data point;
[0042] Step 4. Model evaluation:
[0043] Evaluate the trained XGBoost regressor based on the test set;
[0044] pass:
[0045] Calculate the root mean square error RMSE;
[0046] Also through:
[0047] Calculate the mean absolute error MAE;
[0048] Step L5, Model Optimization:
[0049] Compare the root mean square error (RMSE) and mean absolute error (MAE) with the preset error thresholds 1, RMSEy, and 2, MAEy, respectively:
[0050] When RMSE>RMSEy and MAE>MAEy, adjust the hyperparameters of XGBoost.
[0051] As a further solution of the present invention: during model selection, the initial learning rate of the XGBoost regressor is set to 0.1, the initial maximum depth is set to 6, and the initial subsample ratio is set to 0.8.
[0052] As a further solution of the present invention: in the model evaluation, the adjustment method is as follows:
[0053] For the learning rate:
[0054] If the model converges too slowly and the training time is too long, increase the learning rate appropriately and adjust it to 0.15;
[0055] If the model training is unstable, reduce the learning rate and adjust it to 0.05;
[0056] For maximum depth:
[0057] If the model is underfitting and the error on the validation set does not decrease significantly, increase the maximum depth to 8;
[0058] If the model is overfitting, reduce the maximum depth to 4;
[0059] For subsample proportions:
[0060] If the model's generalization ability is insufficient and its performance on the test set is poor, the subsample ratio should be appropriately reduced to 0.7;
[0061] If the model training efficiency is low, increase the subsample ratio and adjust it to 0.9.
[0062] As a further solution of the present invention: wherein, after each adjustment of the hyperparameters, the model is retrained using the training set, and the model performance is evaluated using the validation set, and whether to continue adjusting the hyperparameters is determined based on the validation set results.
[0063] Beneficial effects of the present invention:
[0064] Comprehensive Data Collection: When collecting sewage treatment plant operating data, we meticulously categorize it into dynamic data (such as influent flow, COD, TN, TP concentrations, energy consumption, and chemical consumption) and static data (such as treatment process type, equipment power, and methane recovery unit parameters). This comprehensive data collection provides a solid foundation for subsequent accurate feature analysis and model training, covering the multiple factors that influence sewage treatment plant carbon accounting, making carbon accounting more scientific and comprehensive.
[0065] Accurate feature analysis: The calculations of methane emissions, nitrous oxide emissions, CO2 emissions from electricity consumption, CO2 emissions from reagent consumption, and total carbon emissions comprehensively consider various relevant factors. For example, methane emissions are calculated using factors such as influent flow rate, COD concentration, methane emission factor, and methane recovery efficiency; while CO2 emissions from electricity consumption are calculated using factors such as equipment power, operating hours, and grid emission factors. This precise feature analysis enables more accurate calculation of various carbon emissions, resulting in a more reliable total carbon emissions figure, providing accurate data support for carbon accounting.
[0066] Reasonable model selection: The XGBoost regressor was selected for model training, and hyperparameters such as the initial learning rate, initial maximum depth, and initial subsample ratio were appropriately set. A smaller learning rate ensured more stable model training, limiting the maximum tree depth prevented overfitting, and randomly selecting a subset of training samples to enhance model generalization. These settings enabled the model to more effectively learn patterns and regularities in the data during training, improving both performance and accuracy.
[0067] Scientific data partitioning: Historical data is divided chronologically into training, validation, and test sets, with ratios of 70%, 15%, and 15%, respectively. This partitioning allows the model to fully learn data patterns, adjust hyperparameters during training to prevent overfitting, and accurately evaluate the model's predictive ability on unseen data, ensuring its reliability and practicality.
[0068] Effective model training: Using mean squared error as a loss function effectively measures the degree of deviation between the model's predictions and the true values. By minimizing the loss function and continuously optimizing the model, the model's prediction performance on the training set continues to improve, providing a strong guarantee for accurately predicting carbon emissions from sewage treatment plants.
[0069] Detailed model evaluation: Models are evaluated by calculating the root mean square error (RMSE) and mean absolute error (MAE), measuring the error between the model's predicted values and the true values from different perspectives. RMSE reflects the average magnitude of the error, while MAE more directly reflects the degree of prediction bias. Combining these two error evaluation methods provides a more comprehensive assessment of model performance.
[0070] Flexible hyperparameter adjustment: XGBoost hyperparameters are flexibly adjusted based on the comparison of RMSE and MAE with preset error thresholds. To address various issues that arise during model training (such as slow convergence, unstable training, underfitting, overfitting, insufficient generalization, and low training efficiency), hyperparameters such as learning rate, maximum depth, and subsample ratio are adjusted in a targeted manner, allowing the model to continuously optimize and adapt to different datasets and actual situations, thereby improving the model's prediction accuracy and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The present invention will be further described below with reference to the accompanying drawings.
[0072] Figure 1 It is a flow chart of a method for optimizing and predicting carbon accounting for sewage treatment plants based on machine learning according to the present invention.
[0073] Figure 2 This is a flow chart of feature analysis in a method for optimizing and predicting carbon accounting for sewage treatment plants based on machine learning according to the present invention.
[0074] Figure 3 It is a flow chart of model analysis in a method for optimizing and predicting carbon accounting for sewage treatment plants based on machine learning according to the present invention. DETAILED DESCRIPTION
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0076] Example 1
[0077] See also Figure 1 、 Figure 2 and Figure 3 As shown, the present invention is a method for optimizing and predicting carbon accounting in sewage treatment plants based on machine learning, comprising:
[0078] Step 1: Data Collection:
[0079] Collect the operation data of the sewage treatment plant during the observation period. The operation data is divided into dynamic data and static data;
[0080] Dynamic data include influent flow, COD, TN, TP concentration, energy consumption, and chemical consumption;
[0081] The influent flow reflects the load of the sewage treatment plant;
[0082] The concentrations of COD, TN, and TP reflect the degree of sewage pollution;
[0083] Energy consumption and drug consumption are directly related to resource consumption in the sewage treatment process;
[0084] Static data include treatment process type, equipment power, and methane recovery unit parameters;
[0085] The type of treatment process determines the manner and process of wastewater treatment, and different processes have different impacts on carbon emissions;
[0086] Equipment power affects electricity consumption;
[0087] The parameters of the methane recovery unit are related to the methane recovery efficiency;
[0088] Among them, COD is the amount of reducing substances that need to be oxidized in the water sample measured by chemical methods, TN refers to the total amount of various forms of inorganic and organic nitrogen in the water, and TP is the total amount of various forms of phosphorus in the water sample;
[0089] Step 2: Feature Analysis:
[0090] Methane emissions calculation:
[0091] pass:
[0092] Calculate the methane emissions E CH4 ;
[0093] Where Q is the water flow rate, C COD is COD concentration, EF CH4 is the methane emission factor, where the methane emission factor is predetermined based on the treatment process type, HS CH4 is the methane recovery efficiency; the methane recovery efficiency is predetermined based on the parameters of the methane recovery device;
[0094] Calculation of Nitrous Oxide Emissions:
[0095] pass:
[0096] Calculate the nitrous oxide emissions E N2O ;
[0097] Where C TN is the TN concentration, EF N2O is the nitrous oxide emission factor;
[0098] Calculation of carbon dioxide emissions from electricity consumption:
[0099] By: E DL =∑(P j ×T j )×EF DW
[0100] Calculate the carbon dioxide emissions E from electricity consumption DL ;
[0101] Where, P j is the equipment power of different power equipment, where different power equipment has different powers and consumes different amounts of electricity; T j is the operating time, where the longer the power equipment operates, the greater the power consumption, and j represents different power equipment; EF DWThe grid emission factor is related to factors such as the power generation structure of the local grid. For example, the grid with a high proportion of thermal power has a relatively large emission factor.
[0102] Calculation of carbon dioxide emissions from pharmaceutical consumption:
[0103] By: E YH =∑(M k ×EF k )
[0104] Calculate the carbon dioxide emissions E from electricity consumption DL ;
[0105] In the formula, M k is the dosage of the agent, where the more agents used, the more carbon emissions will be generated; EF k is the emission factor for pharmaceutical production, where the carbon emission intensity in the production process of different pharmaceuticals is different; k represents different pharmaceuticals;
[0106] Total carbon emissions calculation:
[0107] pass:
[0108] Calculate the total carbon emissions E1;
[0109] Where GWP CH4 and GWP N2O are the global warming potentials of methane and nitrous oxide, respectively. In this embodiment, the values are 25 and 298, respectively. This is because, although methane and nitrous oxide are present in lower concentrations in the atmosphere than carbon dioxide, their impact on global warming, measured on an equal basis, is 25 times and 298 times greater than that of carbon dioxide, respectively.
[0110] Step 3: Model analysis:
[0111] Step 1. Model selection:
[0112] Use XGBoost regressor for model training;
[0113] Set its initial learning rate to 0.1. The learning rate controls the step size of each model update. A smaller learning rate can make the model training more stable, but the training speed will be slower.
[0114] Set its initial maximum depth to 6, which limits the maximum depth of the tree in the tree model to prevent the model from being too complex and causing overfitting;
[0115] The initial subsample ratio is set to 0.8, that is, 80% of the samples are randomly selected for training each time. This helps prevent the model from overfitting and increases the generalization ability of the model.
[0116] Step 2: Data partitioning:
[0117] Divide the historical specified period into several specified time periods, where the specified time period covers multiple standard time periods, and then obtain the operating data of multiple standard time periods within the same specified time period and obtain the total carbon emissions through feature analysis;
[0118] Then, the data sets are divided into multiple standard periods in the same specified time period in chronological order;
[0119] The first 70% of the time series data is used as a training set to train the model and allow the model to learn the patterns and regularities in the data;
[0120] The middle 15% is used as a validation set to adjust hyperparameters during model training, evaluate the performance of the model under different parameter settings, and prevent model overfitting;
[0121] The last 15% is used as a test set to evaluate the trained XGBoost regressor, evaluate the performance of the model, and test the model's predictive ability on unseen data;
[0122] Step 3, model training:
[0123] Based on the training set, the mean square error is used as the loss function;
[0124] The loss function calculation formula is:
[0125] Where MSE is the mean square error as the loss function, m is the number of data points in the training set, and F0 g is the total carbon emissions of the g-th data point in the training set, F1 g is the total carbon emissions predicted by the model for the g-th data point;
[0126] In this embodiment, the loss function measures the degree of deviation between the model prediction value and the true value; the smaller the value, the better the prediction effect of the model on the training set; the larger the value, the larger the model prediction error, and further optimization and adjustment are needed;
[0127] Step 4. Model evaluation:
[0128] Evaluate the trained XGBoost regressor based on the test set;
[0129] pass:
[0130] Calculate the root mean square error RMSE;
[0131] In this embodiment, the root mean square error reflects the average magnitude of the error between the model prediction value and the true value. The smaller the value, the higher the model prediction accuracy.
[0132] Also through:
[0133] Calculate the mean absolute error MAE;
[0134] In this embodiment, the mean absolute error measures the average level of absolute errors between the predicted value and the true value, which can more intuitively reflect the degree of prediction deviation;
[0135] Step L5, Model Optimization:
[0136] Compare the root mean square error (RMSE) and mean absolute error (MAE) with the preset error thresholds 1, RMSEy, and 2, MAEy, respectively:
[0137] When RMSE>RMSEy and MAE>MAEy;
[0138] Then adjust the hyperparameters of XGBoost:
[0139] The specific adjustment methods are as follows:
[0140] For the learning rate:
[0141] If the model converges too slowly and the training time is too long, increase the learning rate appropriately and adjust it to 0.15. However, be careful to prevent the model from not converging or oscillating due to an excessively large learning rate.
[0142] If the model training is unstable and easily falls into the local optimal solution, reduce the learning rate and adjust it to 0.05;
[0143] For maximum depth:
[0144] If the model is underfitting, that is, the model has large errors on both the training set and the test set, and the error on the validation set has not decreased significantly, try increasing the maximum depth to 8 to allow the model to learn more complex patterns.
[0145] If the model is overfitting, that is, the error of the training set is small and the error of the test set is large, reduce the maximum depth to 4 to reduce the complexity of the model;
[0146] For subsample proportions:
[0147] If the model's generalization ability is insufficient and its performance on the test set is poor, the subsample ratio can be appropriately reduced to 0.7 to further enhance the model's generalization ability.
[0148] If the model training efficiency is low, the subsample ratio can be appropriately increased to 0.9 to speed up the training;
[0149] After each hyperparameter adjustment, the model is retrained using the training set and the model performance is evaluated using the validation set. The decision on whether to continue adjusting the hyperparameters is based on the validation set results.
[0150] Step 4: Calculation and Forecast:
[0151] The collected operational data from the sewage treatment plant during the observation period is fed into a trained and evaluated XGBoost regressor. The model then outputs a predicted carbon emission value for the current sewage treatment plant based on the relationship between the learned features and carbon emissions. Based on the predicted carbon emissions, sewage treatment plant managers can formulate corresponding energy conservation and emission reduction strategies.
[0152] Example 1 comprehensively collects dynamic and static operating data of the sewage treatment plant during the observation period, covering key information such as load, pollution level, resource consumption, process, equipment, etc., to provide a rich basis for subsequent analysis. In the feature analysis stage, accurate calculation models of different greenhouse gas emissions and total carbon emissions are constructed, and the impact of various factors on carbon emissions is fully considered. The XGBoost regressor is used for modeling, the initial hyperparameters are reasonably set, and through scientific data division, the training set, validation set and test set are used to respectively realize model training, hyperparameter adjustment and performance evaluation. In the model evaluation, the root mean square error and mean absolute error are compared with the preset thresholds, and the hyperparameters are adjusted accordingly to optimize the model performance. Ultimately, the trained model can predict carbon emissions based on the input operating data, help managers formulate energy-saving and emission reduction strategies, achieve optimized prediction of carbon accounting for sewage treatment plants, and provide strong support for low-carbon operations.
[0153] Example 2
[0154] See also Figure 1 、 Figure 2 and Figure 3 As shown, as the second embodiment of the present invention, when the present application is specifically implemented, compared with the first embodiment, the technical solution of this embodiment is different from that of the first embodiment only in that this embodiment further includes a preprocessing step:
[0155] Eliminate outliers in operating data based on the 3σ principle;
[0156] The specific method is as follows:
[0157] First calculate the mean μ of various running data i and standard deviation σ i ;
[0158] Wherein, i=1, 2, ..., n, n represents the number of types of operating data;
[0159] When x i <μ i -3σ i or x i >μ i +3σ i When , the corresponding operating data is regarded as an abnormal value; where x i is the i-th type of operation data;
[0160] In this embodiment, under the normal distribution assumption, the running data falls within μ i +3σ i The probability of being within the range is about 99.7%. Data outside this range is likely to be erroneous data or extreme values under special circumstances, which will interfere with subsequent analysis;
[0161] Example 2 adds a preprocessing step based on the 3σ principle to Example 1. By calculating the mean and standard deviation of various types of operating data, and based on the rule that data falling outside the range of the mean plus or minus three times the standard deviation is an outlier, abnormal data can be effectively identified and eliminated. Under the assumption of normal distribution, high-probability errors or extreme interference data can be excluded, ensuring that the data used for subsequent feature analysis and model construction is more accurate and reliable, improving data quality, and making the carbon accounting optimization prediction model constructed based on this data more stable and accurate, laying a solid foundation for the accurate prediction of carbon emissions from sewage treatment plants, and enhancing the guiding value of the prediction results for actual operational decisions.
[0162] Example 3
[0163] See also Figure 1 、 Figure 2 and Figure 3 As shown, as the third embodiment of the present invention, when this application is specifically implemented, compared with the first and second embodiments, the technical solution of this embodiment is to combine the solutions of the above-mentioned first and second embodiments for implementation.
[0164] Example 3 combines the solutions of Example 1 and Example 2. It not only has the advantages of Example 1 in comprehensively collecting data, scientifically constructing feature analysis models, and using XGBoost regressors for effective modeling and evaluation optimization to achieve carbon emission prediction and guide energy conservation and emission reduction; it also integrates the characteristics of Example 2 in eliminating outliers in operating data based on the 3σ principle to ensure data quality. By combining the two, accuracy is guaranteed from the source of the data, and optimized predictions are achieved during the model construction and application process, which comprehensively improves the reliability and practicality of carbon accounting predictions for sewage treatment plants, providing more complete and powerful technical support for sewage treatment plants to efficiently carry out carbon management work and achieve sustainable low-carbon development.
[0165] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field according to actual conditions.
[0166] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for optimizing and predicting carbon accounting in sewage treatment plants based on machine learning, characterized in that: The following steps are involved: Data collection: Collect the operation data of the sewage treatment plant during the observation period. Preprocessing: Eliminate outliers in the operating data based on the 3σ principle; Characterization analysis: Determine the carbon emissions of the wastewater treatment plant based on post-treatment operational data; Model analysis: build and train an XGBoost regressor model; Calculation and prediction: Input the preprocessed operating data into the trained XGBoost regressor and obtain the output carbon emission prediction value.
2. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 1 is characterized in that: Operation data is divided into dynamic data and static data; dynamic data includes influent flow, COD, TN, TP concentration, energy consumption, and chemical consumption; static data includes treatment process type, equipment power, and methane recovery device parameters.
3. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 1 is characterized in that: The preprocessing method is as follows: First calculate the mean μ of various running data i and standard deviation σ i ; Wherein, i=1, 2, ..., n, n represents the number of types of operating data; When x i <μ i -3σ i or x i >μ i +3σ i When , the corresponding operating data is regarded as an abnormal value; where x i is the i-th running data.
4. The method for optimizing and predicting carbon accounting in sewage treatment plants based on machine learning according to claim 2 is characterized in that: The feature analysis method is as follows: Methane emissions calculation: Where, E CH4 is the methane emission, Q is the water inlet flow, C COD is COD concentration, EF CH4 is the preset methane emission factor, HS CH4 is the preset methane recovery efficiency; Calculation of Nitrous Oxide Emissions: Where, E N2O is the nitrous oxide emission, C TN is the TN concentration, EF N2O is the nitrous oxide emission factor; Calculation of carbon dioxide emissions from electricity consumption: E DL =∑(P j ×T j )×EF DW ; Where, E DL is the carbon dioxide emissions from electricity consumption, P j is the equipment power of different power equipment; T j is the operating time, j represents different power equipment; EF DW is the preset grid emission factor; Calculation of carbon dioxide emissions from pharmaceutical consumption: E YH =∑(M k ×EF k ) Where, E DL is the carbon dioxide emissions from electricity consumption, where M k is the dosage of the drug; EF k The emission factor for the preset pharmaceutical production; k represents different pharmaceuticals; Total carbon emissions calculation: Where, E1 total carbon emissions, GWP CH4 and GWP N2O are the assumed global warming potential values for methane and nitrous oxide, respectively.
5. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 4 is characterized in that: in, GWP CH4 The value is 25, and GWP N2O The value is 298.
6. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 4 is characterized in that: Model analysis: Step 1. Model selection: Use XGBoost regressor for model training; Step 2: Data partitioning: Divide the historical specified period into several specified time periods, where the specified time period covers multiple standard time periods, and then obtain the operating data of multiple standard time periods within the same specified time period and obtain the total carbon emissions through feature analysis; Then, the data sets are divided into multiple standard periods in the same specified time period in chronological order; The first 70% of the time series data is used as the training set, the middle 15% as the validation set, and the last 15% as the test set; Step 3, model training: Based on the training set, the mean square error is used as the loss function; Step 4. Model evaluation: Evaluate the trained XGBoost regressor based on the test set; Step L5, Model Optimization: Compare the root mean square error (RMSE) and mean absolute error (MAE) with the preset error thresholds 1, RMSEy, and 2, MAEy, respectively: When RMSE>RMSEy and MAE>MAEy, adjust the hyperparameters of XGBoost.
7. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 6 is characterized in that: The loss function calculation formula is: Where MSE is the mean square error as the loss function, m is the number of data points in the training set, and F0 g is the total carbon emissions of the g-th data point in the training set, F1 g is the total carbon emissions predicted by the model for the g-th data point; The formula in model evaluation is: and Where RMSE is the root mean square error and MAE is the mean absolute error.
8. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 6 is characterized in that: During model selection, the initial learning rate of the XGBoost regressor was set to 0.1, the initial maximum depth was set to 6, and the initial subsample ratio was set to 0.
8.
9. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 6 is characterized in that: During model evaluation, the adjustments are as follows: If the model converges slowly, adjust the learning rate to 0.15; If the model training is unstable, reduce the learning rate and adjust it to 0.05; If the model is underfitting, adjust the maximum depth to 8; If the model is overfitting, adjust the maximum depth to 4; If the model generalization ability is insufficient, the subsample ratio is adjusted to 0.7; If the model training efficiency is low, the sub-sample ratio is adjusted to 0.
9.
10. The method for optimizing and predicting carbon accounting of sewage treatment plants based on machine learning according to claim 6, characterized in that: in, After each hyperparameter adjustment, the model is retrained using the training set and the model performance is evaluated using the validation set. The decision on whether to continue adjusting the hyperparameters is based on the validation set results.