A short-range denitrification-anaerobic ammonium oxidation system and its intelligent control method
By constructing a machine learning model and particle swarm optimization algorithm, key influencing factors were identified and precise regulation of the short-range denitrification-anaerobic ammonium oxidation system was achieved, solving the problem of unstable denitrification performance of the system and improving the total nitrogen removal rate and system stability.
Patent Information
- Application Number
- CN202510299764.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The denitrification performance of the short-range denitrification-anaerobic ammonium oxidation system is unstable due to changes in the microbial community during operation, and existing technologies make it difficult to achieve precise control and efficient denitrification.
A machine learning model is constructed to identify key influencing factors through feature importance analysis, and the optimal control strategy is designed in combination with the particle swarm optimization algorithm to achieve precise regulation of the system and efficient denitrification.
The total nitrogen removal rate of the system was improved, the stability and adaptability of the system were enhanced, the operating costs were reduced, and the generalization ability of the model and the adaptability of the control strategy were improved.
Smart Images

Figure CN119828482B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent regulation of sewage treatment, and more specifically, to a short-range denitrification-anaerobic ammonia oxidation system and an intelligent regulation method thereof. Background Art
[0002] Among the many processes in wastewater treatment, biological denitrification has always occupied a pivotal position. However, with rising environmental awareness and increasingly stringent energy conservation and emission reduction requirements, traditional biological denitrification processes face the dual challenges of low energy consumption and low carbon emissions. Against this backdrop, the discovery of the anaerobic ammonium oxidation (ANAMMOX) biological denitrification pathway has revolutionized the field of biological denitrification in wastewater. ANAMMOX technology, with its high efficiency and low energy consumption, has rapidly become a research hotspot in wastewater treatment. The ANAMMOX process relies on a stable supply of nitrite as a substrate to achieve efficient ammonia nitrogen removal. However, ensuring a continuous and stable supply of nitrite has become a key constraint to the widespread application of ANAMMOX technology. In recent years, researchers have discovered through in-depth research that the short-cut denitrification process has extremely high nitrite accumulation efficiency, providing an ideal nitrite source for ANAMMOX technology. Combining short-cut denitrification with ANAMMOX has resulted in a novel biological denitrification process: the short-cut denitrification-ANAMMOX process. This coupled system not only improves the denitrification efficiency, but also significantly reduces energy consumption and carbon emissions, meeting the current high requirements of the environmental protection field for sewage treatment processes. However, the complexity of the short-cut denitrification-anaerobic ammonium oxidation coupling system also brings considerable challenges. In actual operation, the denitrification performance of the system is affected by a variety of factors, including the accumulation efficiency of nitrite in the short-cut denitrification process, the denitrification rate in the anaerobic ammonium oxidation process, and the synergistic effect between the two. The interaction between these factors makes it difficult to ensure the stability of the system. Once fluctuations or even collapses occur, it will seriously affect the denitrification effect and operational efficiency. In order to meet this challenge, researchers began to try to model the sewage treatment process to simulate and optimize the biological treatment process.
[0003] However, traditional modeling methods need to consider complex dynamics and a large number of chemical reaction parameters, which not only increases the difficulty of modeling, but also limits the versatility of the model in different application scenarios. In addition, the process of adjustment and parameter verification for different systems also increases the application cost of the model. In recent years, with the rapid development of big data and machine learning technologies, machine learning-based models have gradually become an effective tool for predicting the performance of biological denitrification of sewage. Compared with traditional modeling methods, machine learning models do not require an in-depth understanding of the complex biological denitrification mechanism, and can only rely on the system's existing operating parameters and denitrification performance data samples for prediction and optimization. Applying big data-driven machine learning to short-term denitrification-anaerobic ammonium oxidation systems can achieve accurate prediction and intelligent regulation of system performance, thereby improving the stability and denitrification efficiency of the system.
[0004] However, although machine learning models have shown great potential in the field of biological denitrification of sewage, how to carry out targeted intelligent regulation based on the characteristics of short-range denitrification-anaerobic ammonium oxidation systems remains a technical problem that needs to be solved urgently. In addition, there are obvious differences in the key functional bacteria in different systems, and their requirements for optimal environmental conditions are also different. If regulation is only carried out from the perspective of operating parameters, the differences between the microbial communities in each system may be ignored, resulting in poor results of the regulation strategy. In addition, different actual operating environments may also lead to differences in adjustable parameters, further reducing the adaptability of the optimization regulation strategy.
[0005] In related technologies, for example, Chinese patent CN114790039A provides a method and system for intelligent denitrification control of aquaculture wastewater. This method includes obtaining different water quality indicators for different wastewater treatment units based on the wastewater treatment process; using big data computing and deep learning algorithms to predict the control parameters in wastewater treatment based on the obtained water quality indicators based on a microbial dynamics model. The control parameters include carbon source amount, internal reflow ratio, and external reflow ratio; based on the prediction results and the set four-level early warning and control measures, early warning and adjustment are carried out for instability, and the dosage and aeration volume are finely controlled to ensure that the effluent is stable and meets the standards. This solution establishes a microbial dynamics model, utilizes big data computing and platform deep learning to achieve automatic (semi-automatic) remote control of the sewage treatment plant. However, the solution has the disadvantage of not eliminating the impact of dynamic changes in the microbial community on system stability.
[0006] From the above, it can be seen that the relevant technology does not provide any technical inspiration on how to solve the problems of performance fluctuation and collapse that may be encountered during the operation of the short-range denitrification-anaerobic ammonium oxidation system. Summary of the Invention
[0007] 1. Technical problems to be solved
[0008] In response to the problem in the prior art that fluctuations in system operating conditions cause changes in the microbial community in the short-range denitrification-anaerobic ammonium oxidation system, thereby leading to unstable system denitrification performance, the present invention provides a short-range denitrification-anaerobic ammonium oxidation system and an intelligent control method thereof, which can construct and screen an optimal model for predicting the system's denitrification performance, and identify key influencing factors for efficient denitrification based on feature importance analysis, thereby using the particle swarm optimization algorithm to design the optimal control strategy for a specific system, thereby achieving precise control of the system and efficient denitrification.
[0009] 2. Technical solution
[0010] The purpose of the present invention is achieved through the following technical solutions.
[0011] The content of this application is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this application is not intended to identify key features or essential features of the technical solution for which protection is sought, nor is it intended to limit the scope of the technical solution for which protection is sought.
[0012] Some embodiments of the present application propose a short-range denitrification-anaerobic ammonium oxidation system and an intelligent control method thereof to solve the technical problems mentioned in the above background technology section.
[0013] As a first aspect of the present application, some embodiments of the present application provide an intelligent control method for a short-term denitrification-anaerobic ammonium oxidation system, comprising the following steps: obtaining sample point data of the short-term denitrification-anaerobic ammonium oxidation system, including operation data, operating parameters, and microbial community structure data, and constructing a data set for machine learning; performing outlier processing, category information conversion, data standardization, correlation analysis, and data set splitting on the data set to obtain a training set and a test set; based on the training set, establishing a prediction model through a machine learning method, and evaluating the prediction performance of each prediction model on the test set to select the optimal prediction model; using the data in the data set as feature data, performing feature data importance analysis, and identifying key parameters that affect the total nitrogen removal rate of the system; using the key parameters as the optimization control parameters of the short-term denitrification-anaerobic ammonium oxidation system, and constructing an optimal control strategy through a particle swarm optimization algorithm.
[0014] Furthermore, outlier processing includes calculating the mean and standard deviation of the feature data of all sample points and making judgments on each feature data. The judgment process is as follows:
[0015] Take the data in the dataset as feature data, calculate the mean and standard deviation of each feature data of all sample points; take the mean value of the feature data of the current sample point minus the standard deviation as the minimum value, and take the mean value of the feature data of the current sample point plus the standard deviation as the maximum value to define the current range; determine whether the feature data of the current sample point is within the current range. If not, it is regarded as an outlier and replaced with the mean value of the feature data of the current sample point.
[0016] Furthermore, data standardization includes using the Standardscaler standard method to normalize the sample point data. The process is expressed as follows:
[0017] ;
[0018] in, Represents sample feature data, is the average value of each feature data, is the standard deviation of each feature data.
[0019] Furthermore, the process of splitting the dataset is as follows: according to the five-fold cross-validation method in the machine learning model evaluation method, the dataset is randomly split into two mutually exclusive datasets in a ratio of 8:2, of which 80% of the data is used as the training set for model training; and 20% of the data is used as the test set for model testing.
[0020] Furthermore, the operating data include ammonia nitrogen concentration, nitrate nitrogen concentration, nitrite concentration, COD concentration, pH, carbon-nitrogen ratio, and nitrate-ammonia nitrogen ratio; the operating parameters include system type, system volume, recirculation ratio, hydraulic retention time, temperature, inoculum sludge type, and carbon source type;
[0021] The microbial community structure data are presented in matrix form, with each column representing a species, each row representing a sample, and the element value being the relative abundance of the corresponding species.
[0022] Furthermore, based on the operational data and operating parameters in the training set, a machine learning method is selected to build a prediction model;
[0023] Using the coefficient of determination , mean square error and mean absolute error The prediction performance of the prediction models is compared by using the indicators and the optimal prediction model is selected.
[0024] Furthermore, the process of comparing the prediction performance of the prediction models is expressed as:
[0025] ;
[0026] ;
[0027] ;
[0028] in, The sample index in the test set is The actual total nitrogen removal rate value, The sample index in the test set is The total nitrogen removal rate value predicted by the prediction model is is the average of the actual total nitrogen removal rates of all samples in the test set, is the sample size.
[0029] Furthermore, the process of feature data importance analysis includes using Python's SHAP package to calculate the SHAP value of each feature data; analyzing the SHAP value of the feature data, and taking features with high SHAP values as key parameters.
[0030] Furthermore, the process of constructing the optimal control strategy includes: selecting one or more key parameters as control parameters, using the particle swarm optimization algorithm to find the optimization parameters, and constructing the optimal control strategy for a specific scenario.
[0031] As a second aspect of the present application, some embodiments of the present application provide a short-cut denitrification-anaerobic ammonium oxidation system, comprising a data module for obtaining sample point data of the short-cut denitrification-anaerobic ammonium oxidation system, including operation data, operating parameters, and microbial community structure data, and constructing a machine learning data set;
[0032] Processing module: performs outlier processing, category information conversion, data standardization, correlation analysis, and data set splitting on the data set to obtain training sets and test sets;
[0033] Model module: Based on the training set, a prediction model is established through machine learning methods, and the prediction performance of each prediction model on the test set is evaluated to select the optimal prediction model;
[0034] Analysis module: Use the data in the dataset as feature data, perform feature data importance analysis, and identify key parameters that affect the total nitrogen removal rate of the system;
[0035] Building module: Using key parameters as the optimization and control parameters of the short-range denitrification-anaerobic ammonium oxidation system, the optimal control strategy is constructed through the particle swarm optimization algorithm.
[0036] 3. Beneficial effects
[0037] Compared with the prior art, the advantages of the present invention are:
[0038] (1) The technical solution of the present invention makes full use of a large amount of multidimensional data such as the operating parameters and microbial community structure of the short-range denitrification-anaerobic ammonium oxidation system, uses machine learning methods to construct and optimize the system denitrification performance prediction model, identifies the key factors affecting the system performance, and the prediction model can fully reflect the operating characteristics of the system;
[0039] (2) The technical solution of the present invention can simultaneously consider the key factors identified based on machine learning and the unique operating scenarios of different systems, and design a unique optimal control strategy for each scenario, thereby improving the high adaptability of the strategy;
[0040] (3) By constructing a multidimensional data set that integrates operating parameters, operation parameters and microbial community structure, a high-precision prediction model was established using the gradient boosting algorithm (the test set R² reached 0.958). The key regulatory factors such as temperature, pH, and carbon-nitrogen ratio were accurately identified by combining the SHAP feature importance analysis. A scenario-adaptive control strategy was generated based on the particle swarm optimization algorithm, which increased the total nitrogen removal rate of the short-range denitrification-anaerobic ammonia oxidation system from 80.7% to 93.5%. The fluctuation range of the removal rate was reduced under the condition of a ±50% fluctuation in the total nitrogen concentration of the influent. At the same time, the operating cost was reduced by optimizing the carbon source addition amount. This effectively solved the problem of system instability caused by ignoring the dynamic changes of the microbial community in the traditional method. The system has the technical advantages of high multi-source data fusion, strong model generalization ability and excellent adaptability of the control strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Schematic diagram of a process for an intelligent control method of a short-range denitrification-anaerobic ammonium oxidation system according to one embodiment of the present invention;
[0042] Figure 2 is a correlation coefficient diagram between features in one embodiment of the present invention;
[0043] Figure 3 A schematic diagram of feature importance ranking in one embodiment of the present invention;
[0044] Figure 4 Schematic diagram for verifying the intelligent control method in one embodiment of the present invention; (a) is the total nitrogen removal rate of the system before intelligent control and its predicted value; (b) is the total nitrogen removal rate of the system after intelligent control for scenario one and its predicted value; (c) is the total nitrogen removal rate of the system after intelligent control for scenario two and its predicted value; (d) is the average value of the total nitrogen removal rate and its predicted value before and after intelligent control. DETAILED DESCRIPTION
[0045] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Combine Figures 1 to 4The present invention provides a smart control method for a short-range denitrification-anaerobic ammonium oxidation system, comprising the following steps:
[0047] Obtain sample point data from the short-range denitrification-anaerobic ammonium oxidation system and construct a machine learning dataset;
[0048] The dataset is processed for outlier processing, category information conversion, data standardization, correlation analysis, and dataset splitting to obtain training and test sets;
[0049] Based on the training set, a prediction model is established through machine learning methods, and the prediction performance of each prediction model on the test set is evaluated to select the optimal prediction model;
[0050] The data in the dataset are used as feature data, and feature data importance analysis is performed to identify the key parameters that affect the total nitrogen removal rate of the system;
[0051] The operation data, operating parameters and community structure data of the short-cut denitrification-anaerobic ammonium oxidation system were obtained, and the key parameters were used as the optimization control parameters of the short-cut denitrification-anaerobic ammonium oxidation system. The optimal control strategy was constructed using the particle swarm optimization algorithm.
[0052] In a specific embodiment, the specific process of the intelligent control method of the short-cut denitrification-anaerobic ammonium oxidation system is as follows:
[0053] S1. Build a dataset
[0054] Obtain data from a short-range denitrification-anaerobic ammonium oxidation system and construct a machine learning dataset by fusing multi-source data.
[0055] Specifically, WebPlotDigidizer is used to obtain the short-range denitrification-anaerobic ammonium oxidation system data, which consists of all sample point data in the short-range denitrification-anaerobic ammonium oxidation system.
[0056] More specifically, the sample point data includes operational data, operating parameters, and microbial community structure data.
[0057] Among them, the operating data include ammonia nitrogen concentration, nitrate nitrogen concentration, nitrite concentration, COD concentration, pH, carbon-nitrogen ratio and nitrate-ammonia nitrogen ratio.
[0058] Operating parameters include system type, system volume, recirculation ratio, hydraulic retention time, temperature, inoculum sludge type, and carbon source type.
[0059] Microbial community structure data includes matrix data of the relative abundance of different genera of microorganisms at each sample site. More precisely, in the form of microbial community structure data, it is presented in a matrix format, with each column representing a genus (species), each row representing a sample, and the element value being the relative abundance of the corresponding species.
[0060] All system data obtained through the above process are merged according to the sample point names to form a matrix containing all necessary information. This matrix will serve as the dataset required to build the machine learning model.
[0061] S2. Data Preprocessing
[0062] The data preprocessing process includes outlier processing, category information conversion, data standardization, correlation analysis, and data set splitting.
[0063] (1) Outlier processing:
[0064] First, before conducting machine learning model analysis, the operating data, operating parameters and community structure data in the above dataset are used as feature data, the mean and standard deviation of each feature data of all sample points are calculated, and the feature data of each sample point are judged.
[0065] Specifically, the judgment process is as follows: the current minimum value is taken as the average value of the current sample point feature data minus the standard deviation, the current maximum value is taken as the average value of the current sample point feature data plus the standard deviation, and the range defined by the current minimum value and the current maximum value as the boundary values is taken as the current range;
[0066] Then, it is determined whether the current sample point feature data is within the current range. If it is not within the current range, the current sample point feature data is an outlier, and the average value of the current sample point feature data is used to replace the outlier.
[0067] (2) Category information conversion:
[0068] In one specific embodiment, since the system type, sludge type, and carbon source type in the operating parameters are all string-type categorical information, they cannot be directly used in the subsequent machine learning model construction. Therefore, the get_dummies function in the pandas package is used to perform one-hot encoding on the string-type categorical information in the operating parameters, such as the system type, sludge type, and carbon source type. A new binary (0 or 1) column is created for each possible category. If a sample belongs to that category, the corresponding column is 1, otherwise it is 0. This converts the string categorical information into numerical information. In this embodiment, the dataset includes a total of 4,863 sample point data.
[0069] get_dummies is a function in the Pandas library that converts categorical data into one-hot-encoded dummy variables. Categorical data is often encountered in machine learning and data processing, but it cannot be directly used for training most models because it is non-numeric. One-hot encoding can be used to convert this categorical data into a numeric form that the model can process.
[0070] (3) Data standardization:
[0071] After the category information is converted, data standardization is performed.
[0072] Specifically, the Standardscaler standard method is used to normalize the sample point data, and the calculation formula is:
[0073]
[0074] in, Represents sample feature data, is the average value of each feature data, is the standard deviation of each feature data.
[0075] (4) Correlation analysis:
[0076] The SciPy package in Python was used to calculate the Pearson correlation between different feature data. Pearson correlation is a statistical method to measure the linear relationship between two continuous variables.
[0077] In this example, Pearson correlation coefficients were calculated between 14 features in the operational data and operating parameters.
[0078] The Pearson correlation coefficient is used to measure the strength and direction of the linear relationship between two continuous variables. The value range of the Pearson correlation coefficient is between -1 and 1. When the Pearson correlation coefficient is equal to 1, it means that the two variables are completely positively correlated, that is, when one variable increases, the other variable also increases, showing a completely linear positive correlation. When the Pearson correlation coefficient is equal to -1, it means that the two variables are completely negatively correlated, that is, when one variable increases, the other variable decreases, showing a completely linear negative correlation. When the Pearson correlation coefficient is equal to 0, it means that there is no linear relationship between the two variables, that is, there is no linear correlation between the two.
[0079] like Figure 2As shown in the figure, the correlation between the characteristic data is judged according to the correlation coefficient. The result shows that the correlation coefficient between the influent ammonia nitrogen concentration and the influent total nitrogen concentration is 0.93, and its absolute value exceeds the significant correlation threshold of 0.8. Therefore, there is a significant positive correlation between the two characteristic data.
[0080] In this example, since the influent ammonia nitrogen and total nitrogen concentrations directly determine the important control indicators of the short-cut denitrification-anaerobic ammonium oxidation system's operating performance, the two key characteristic data, influent ammonia nitrogen concentration and influent total nitrogen concentration, were retained. There was no significant correlation between the remaining characteristic data.
[0081] (5) Dataset splitting:
[0082] According to the five-fold cross-validation method in the machine learning model evaluation method, the dataset was randomly split into two mutually exclusive datasets in a ratio of 8:2, of which 80% of the data was used as the training set for model training; and 20% of the data was used as the test set for model testing.
[0083] The data preprocessing process described above is an important step before building a machine learning model, aiming to improve data quality and ensure the accuracy and effectiveness of model training. Among them, outlier processing can prevent the impact of extreme values on model training and improve the robustness of the model; category information conversion enables string-type category information to be correctly processed by the machine learning model; data standardization helps eliminate dimensional differences between different feature data, improving the convergence speed and performance of the model; correlation analysis helps identify key feature data, simplify the model structure, and improve model interpretability; dataset splitting is the basis for model evaluation. By dividing the training set into a test set, the generalization ability of the model can be objectively evaluated.
[0084] S3. Build a prediction model
[0085] The process of building a predictive model includes selecting a machine learning method, establishing a predictive model, model evaluation and selection, and optimizing the model by combining community structure data.
[0086] Specifically, based on the operating data and operating parameters in the training set divided as above, the program package Sklearn was used and various machine learning methods were used to establish prediction models for total nitrogen removal rate. The five-fold cross-validation method was used to compare and evaluate the prediction performance of each prediction model on the test set, and the optimal prediction model was selected.
[0087] (1) Selecting a machine learning method:
[0088] Based on the operating data and operating parameters in the training set divided above, a variety of machine learning methods are selected to establish a prediction model.
[0089] In this embodiment, nine machine learning methods are used, including linear regression, K-nearest neighbor, random forest, gradient boosting, XGBoost, support vector machine, AdaBoost, LightBoost and artificial neural network.
[0090] (2) Establish a prediction model:
[0091] The program package Sklearn was used to establish prediction models for total nitrogen removal rate based on various machine learning methods.
[0092] Sklearn is an open source machine learning library based on Python. Sklearn supports a variety of machine learning algorithms, including linear regression, logistic regression, decision trees, random forests, support vector machines, and neural networks, and provides a variety of evaluation metrics and cross-validation methods to measure model performance. In this embodiment, Sklearn can be used to partition data, train models, and evaluate them.
[0093] (3) Model evaluation and selection:
[0094] The five-fold cross-validation method was used to compare and evaluate the prediction performance of each prediction model on the test set and select the optimal prediction model.
[0095] Specifically, the coefficient of determination , mean square error and mean absolute error Three indicators are used to compare the prediction performance of the prediction models.
[0096] Using the coefficient of determination Measures the degree of fit of the model to the data. The closer the value is to 1, the better the model performance is. Measures the deviation between the predicted value and the actual value. The smaller the value, the more accurate the model prediction. The mean absolute error is used Measures the average value of the prediction error. The smaller the value, the smaller the model prediction error. The specific calculation process is as follows:
[0097] ;
[0098] ;
[0099] ;
[0100] in, The sample index in the test set is The actual total nitrogen removal rate value, The sample index in the test set is The total nitrogen removal rate value predicted by the prediction model is is the average of the actual total nitrogen removal rates of all samples in the test set, is the sample size.
[0101] Specifically, the optimal prediction model is selected based on the results of the three evaluation indicators. Table 1 shows the prediction performance of each prediction model on the test set.
[0102] Table 1 Prediction performance of each prediction model on the test set
[0103]
[0104] In this example, based on the results of the three evaluation indicators, it can be found that among the various models, gradient boosting and XGBoost have the best prediction performance. Therefore, gradient boosting and XGBoost were selected to further build a total nitrogen removal rate prediction model based on operation data, operating parameters and community structure data.
[0105] (4) Optimizing the model by combining community structure data:
[0106] We further considered incorporating microbial community structure data into the prediction model to assess its impact on model performance. We compared the performance of models using only operational data and operating parameters with those using operational data, operating parameters, and microbial community structure data. Based on the evaluation results, we selected the model with the best performance as the final prediction model.
[0107] As shown in Table 2, the prediction performance of the two selected prediction models on the test set is shown.
[0108] Table 2 Prediction performance of the two prediction models on the test set
[0109]
[0110] The results of the three evaluation indicators show that, combined with the microbial community structure data, gradient boosting has the best performance in predicting the total nitrogen removal rate of the short-range denitrification-anaerobic ammonium oxidation system. Therefore, gradient boosting was selected as the optimal prediction model.
[0111] By constructing and evaluating multiple prediction models, this step identified gradient boosting as the optimal prediction model for the total nitrogen removal rate of the short-range denitrification-anaerobic ammonium oxidation system. This model not only takes into account the operating data and operating parameters, but also combines the microbial community structure data, thereby improving the accuracy and reliability of the prediction.
[0112] This solution leverages a wealth of multidimensional data, including operational parameters and microbial community structures, from short-range denitrification-anaerobic ammonium oxidation systems. It utilizes machine learning to construct and optimize a prediction model for system denitrification performance, identifying key factors influencing system performance and fully reflecting the system's operational characteristics. This solution simultaneously considers key factors identified through machine learning and the unique operational scenarios of different systems, designing a unique optimal control strategy for each scenario, thereby enhancing the strategy's adaptability.
[0113] S4. Importance analysis of feature data
[0114] The purpose of this step is to perform characteristic data importance analysis and identify key parameters that affect the total nitrogen removal rate of the system.
[0115] Based on the gradient boosting prediction model constructed in step S3 above, the Shapley additive interpretation method is used to perform feature data importance analysis. The SHAP package in Python is used to calculate the SHAP value of each feature data, which is used to represent the average contribution of each feature data to the predicted value of the prediction model. The SHAP values of the feature data are analyzed, and features with high SHAP values are used as key parameters.
[0116] The Shapley Additive Explanations method (SHAP) is a game-theoretic method for interpreting feature importance. It explains the average contribution of each feature to the model's predictions by interpreting the predictions of a machine learning model as the weighted sum of the contributions of each feature. The SHAP value measures the contribution of each feature to the model's predictions, averaging the contributions across all possible feature combinations to provide a global importance assessment. The SHAP package in Python implements the SHAP method, providing a series of functions and classes for calculating SHAP values, visualizing feature importance, and interpreting machine learning model predictions.
[0117] By analyzing the SHAP value of each feature data, the feature data with high SHAP values are identified as key parameters. These feature data with higher SHAP values than other feature data have a significant impact on the prediction model's results and are factors that require special attention when regulating system performance.
[0118] like Figure 3 As shown in the figure, these key parameters include temperature, inlet pH, inlet total nitrogen load, inlet nitrate-ammonia ratio, and inlet carbon-nitrogen ratio. The results of the characteristic data importance analysis provide an important basis for intelligent control. In actual operation, the high SHAP values of these key parameters can be used to adjust system operating conditions to optimize total nitrogen removal efficiency.
[0119] For example, if temperature is a key factor affecting system performance, denitrification can be optimized by adjusting the system temperature. Similarly, if influent pH has a significant impact on system performance, denitrification efficiency can be improved by adjusting the influent pH.
[0120] S5. Use particle swarm optimization algorithm to build the optimal control strategy
[0121] The process of constructing an optimal control strategy includes collecting and processing data, selecting control parameters, applying the particle swarm optimization algorithm, and constructing the optimal control strategy.
[0122] (1) Data collection and processing:
[0123] For the running short-cut denitrification-anaerobic ammonium oxidation system, the system inlet and outlet water are collected, and the inlet and outlet water quality is tested using national standard methods to form the specific operating data and detailed operating parameters of the short-cut denitrification-anaerobic ammonium oxidation system.
[0124] Sludge samples in the system were collected and subjected to 16S rRNA high-throughput sequencing. QIIME2 software was used to analyze the genus-level community structure matrix data.
[0125] (2) Select control parameters:
[0126] Based on the optimization and control requirements of the short-cut denitrification-anaerobic ammonium oxidation system in a specific scenario, one or more key parameters are selected as control parameters. In this method, key parameters include temperature, influent pH, influent total nitrogen load, influent nitrate-ammonia nitrogen ratio, and influent carbon-nitrogen ratio. These key parameters have been identified in step S4 as important factors affecting the efficient denitrification of the system.
[0127] The other operating data, detailed operating parameters and microbial community structure data obtained above were used as input data for subsequent optimization calculations.
[0128] (3) Application of particle swarm optimization algorithm:
[0129] Based on the above control parameters and input data, the Scikit-opt package in Python is used to perform calculations based on the particle swarm optimization algorithm (PSO) to find the optimal values of the above control parameters as the optimization parameters of the optimal control strategy in a specific scenario.
[0130] Scikit-opt is a Python library for solving optimization problems. It integrates many classic intelligent optimization algorithms, such as genetic algorithms, particle swarm optimization (PSO), and simulated annealing. In Scikit-opt, the particle swarm optimization (PSO) algorithm simulates the foraging behavior of a flock of birds to search for the optimal solution in the solution space. Each particle represents a potential solution, and the optimal solution is approached by iteratively updating its position and velocity.
[0131] In this step, the control parameter is used as the position of the particle, the total nitrogen removal rate of the system or other performance indicators are used as the fitness function, and the PSO algorithm is used to find the optimization parameters that maximize the fitness function.
[0132] (4) Constructing the optimal control strategy:
[0133] According to the optimization parameters calculated by the PSO algorithm, the optimal control strategy for a specific scenario is constructed, thereby realizing intelligent control of the system's efficient denitrification based on the optimal control strategy.
[0134] These optimal control strategies will serve as guidance for system operation and be used to optimize the total nitrogen removal rate of the system or maintain long-term stable operation of the system under fluctuating influent conditions.
[0135] In a specific example, optimal control strategies were designed for two specific scenarios: improving the total nitrogen removal efficiency of the short-cut ANAMMOX system and maintaining long-term stable operation of the short-cut ANAMMOX system under fluctuating influent conditions. The short-cut ANAMMOX system had a volume of 1 L, a recirculation ratio of 2, and an influent total nitrogen concentration of 60 mg / L. Other parameters were as shown in Table 3, Stage I.
[0136] Table 3 Experimental plan of intelligent control
[0137]
[0138] As shown in Table 3, in Scenario 1, pH, temperature, nitrogen source substrate ratio, and carbon-nitrogen ratio characteristics were selected as control parameters. The optimal values of these four control parameters were calculated using the particle swarm optimization algorithm and used as the optimization parameters for the optimal control strategy (Phase II in Table 3). The remaining parameters remained the same as in Phase I. Based on the obtained optimal control strategy, the short-cut denitrification-anaerobic ammonium oxidation system was optimized and regulated in Phase II, achieving an increase in the average total nitrogen removal efficiency of the short-cut denitrification-anaerobic ammonium oxidation system from 80.7% in Phase I to 92.7% in Phase II.
[0139] As shown in Table 3, in Scenario 2, simulating fluctuating influent conditions (total nitrogen concentration increasing from 60 mg / L to 90 mg / L), pH, temperature, nitrogen source substrate ratio, carbon-nitrogen ratio, and hydraulic retention time were selected as control parameters. The particle swarm optimization algorithm was used to obtain the optimal values of these five control parameters and used them as the optimization parameters for the optimal control strategy (Phase III in Table 3). The remaining parameters remained the same as in Phase II. Based on the obtained optimal control strategy, the short-cut denitrification-anaerobic ammonium oxidation system was optimized and controlled in Phase III, achieving an increase in the average total nitrogen removal efficiency of the short-cut denitrification-anaerobic ammonium oxidation system from 92.7% in Phase II to 93.5% in Phase III.
[0140] By adjusting control parameters such as pH, temperature, nitrogen source matrix ratio, carbon-nitrogen ratio and hydraulic retention time, the total nitrogen removal rate of the system was significantly improved.
[0141] like Figure 4 As shown, the total nitrogen removal rate and its predicted value before and after the system intelligent regulation are displayed. Figure 4 Middle: (a) is the total nitrogen removal rate of the system before intelligent regulation and its predicted value; (b) is the total nitrogen removal rate of the system after intelligent regulation for scenario one and its predicted value; (c) is the total nitrogen removal rate of the system after intelligent regulation for scenario two and its predicted value; (d) is the average of the total nitrogen removal rate and its predicted value before and after intelligent regulation.
[0142] The present invention constructs a multidimensional data set that integrates operating parameters, operational parameters, and microbial community structure, uses a gradient boosting algorithm to establish a high-precision prediction model (test set R² reaches 0.958), combines SHAP feature importance analysis to accurately identify key regulatory factors such as temperature, pH, and carbon-nitrogen ratio, and generates a scenario-adaptive control strategy based on the particle swarm optimization algorithm. The total nitrogen removal rate of the short-range denitrification-anaerobic ammonia oxidation system is increased from 80.7% to 93.5%, and the removal rate fluctuation range is reduced under the condition of a ±50% fluctuation in the total nitrogen concentration of the influent. At the same time, the operating cost is reduced by optimizing the carbon source addition amount, effectively solving the system instability problem caused by ignoring the dynamic changes of the microbial community in traditional methods. The present invention has the technical advantages of high multi-source data fusion, strong model generalization ability, and excellent adaptability of the control strategy.
[0143] This step optimizes the performance of the short-range denitrification-anaerobic ammonium oxidation system by collecting system data, selecting key control parameters, and applying the particle swarm optimization algorithm to construct an optimal control strategy. This strategy is then implemented and verified in a real-world system. This step provides a specific implementation path and method for intelligent control, helping to improve the system's total nitrogen removal rate and operational stability.
[0144] In a specific embodiment, a short-cut denitrification-anaerobic ammonium oxidation system includes a data module, a processing module, a model module, an analysis module, and a construction module.
[0145] Specifically, the data module is used to obtain sample point data of the short-range denitrification-anaerobic ammonium oxidation system, including operation data, operating parameters, and microbial community structure data, and to construct a data set for machine learning;
[0146] The processing module is used to process outliers, convert category information, standardize data, perform correlation analysis, and split the data set to obtain training and test sets.
[0147] The model module is used to establish a prediction model based on the training set through machine learning methods, evaluate the prediction performance of each prediction model on the test set, and select the optimal prediction model;
[0148] The analysis module is used to use the data in the dataset as feature data, perform feature data importance analysis, and identify key parameters that affect the total nitrogen removal rate of the system;
[0149] The construction module is used to use key parameters as the optimization and control parameters of the short-range denitrification-anaerobic ammonium oxidation system, and to construct the optimal control strategy through the particle swarm optimization algorithm.
[0150] The technical solution of the present invention reveals the microbial community structure in sludge samples through the comprehensive use of 16S rRNA high-throughput sequencing technology, uses QIIME2 software for in-depth analysis to obtain genus-level community structure matrix data, combines the gradient boosting prediction model and the SHAP method to perform feature importance analysis to optimize the prediction performance, and uses the particle swarm optimization algorithm (PSO) in the Scikit-opt package for calculation, which effectively improves the accuracy and efficiency of sewage treatment process parameter optimization, provides a scientific basis and intelligent means for understanding the function of microbial communities and optimizing sewage treatment processes, and realizes the full-chain technology integration and innovation from data collection and analysis to model optimization and application.
[0151] The above schematically describes the invention and its implementation methods. This description is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited to this. Any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it and designs a structural method and embodiment similar to the technical solution without creativity without departing from the purpose of the invention, they should all fall within the scope of protection of this patent. In addition, the word "including" does not exclude other elements or steps, and the word "one" before an element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. Words such as first and second are used to indicate names and do not indicate any specific order.
Claims
1. A smart control method for a short-range denitrification-anaerobic ammonium oxidation system, characterized by: The following steps are involved: Obtaining sample point data of a shortcut denitrification-anaerobic ammonium oxidation system to construct a machine learning dataset, wherein the sample point data includes operating data, operating parameters, and microbial community structure data, wherein the microbial community structure data is presented in matrix form, with each column representing a species, each row representing a sample, and the element value representing the relative abundance of the corresponding species. The operating data includes ammonia nitrogen concentration, nitrate nitrogen concentration, nitrite concentration, COD concentration, pH, carbon-nitrogen ratio, and nitrate-ammonia nitrogen ratio. The operating parameters include system type, system volume, reflux ratio, hydraulic retention time, temperature, inoculated sludge type, and carbon source type. The dataset was processed for outlier processing, category information conversion, data standardization, correlation analysis, and dataset splitting. Outlier processing used a dynamic range judgment method based on the mean and standard deviation of feature data. Data standardization used the Standardscaler method for standard normalization. The dataset was split using a five-fold cross-validation method to split the training set into a test set with an 8:2 ratio. A prediction model was established by integrating running parameters, operating parameters, and microbial community structure data based on the gradient boosting algorithm, and key microbial features with SHAP values above the threshold were identified through SHAP value analysis. A multi-objective optimization model was constructed, and the particle swarm optimization algorithm was used to dynamically adjust the temperature, pH, carbon-nitrogen ratio and hydraulic retention time parameters to generate the optimal control strategy adapted to different influent conditions. The optimization objectives included a total nitrogen removal rate ≥ 93.5%, a removal rate fluctuation range ≤ ± 5%, and a reduction in carbon source dosage by 20%-30%.
2. The intelligent control method for the short-cut denitrification-anaerobic ammonium oxidation system according to claim 1, characterized in that: Outlier processing includes calculating the mean and standard deviation of the feature data of all sample points and making judgments on each feature data. The judgment process is as follows: Take the data in the dataset as feature data and calculate the mean and standard deviation of each feature data of all sample points; The current range is defined by taking the mean value of the feature data of the current sample point minus the standard deviation as the minimum value and taking the mean value of the feature data of the current sample point plus the standard deviation as the maximum value; Determine whether the feature data of the current sample point is within the current range. If not, it is considered an outlier and replaced with the average value of the feature data of the current sample point.
3. The intelligent control method for the short-cut denitrification-anaerobic ammonium oxidation system according to claim 1, characterized in that: Data standardization includes using the Standardscaler standard method to perform standard normalization on the sample point data. The process is expressed as: ; Represents sample feature data, is the average value of each feature data, is the standard deviation of each feature data.
4. The intelligent control method for the short-cut denitrification-anaerobic ammonium oxidation system according to claim 1, characterized in that: Based on the operating data and operating parameters in the training set, a machine learning method is selected to build a prediction model; Using the coefficient of determination , mean square error and mean absolute error The prediction performance of the prediction models is compared by using the indicators and the optimal prediction model is selected.
5. The intelligent control method for the short-cut denitrification-anaerobic ammonium oxidation system according to claim 3, characterized in that: The process of comparing the predictive performance of the prediction models is expressed as: ; ; ; The sample index in the test set is The actual total nitrogen removal rate value, The sample index in the test set is The total nitrogen removal rate value predicted by the prediction model is is the average of the actual total nitrogen removal rates of all samples in the test set, is the sample size.
6. The intelligent control method for a short-cut denitrification-anaerobic ammonium oxidation system according to claim 1, characterized in that: The process of feature data importance analysis includes using Python's SHAP package to calculate the SHAP value of each feature data; analyzing the SHAP value of the feature data, and taking features with high SHAP values as key parameters.
7. The intelligent control method for a short-cut denitrification-anaerobic ammonium oxidation system according to claim 1, characterized in that: The process of constructing the optimal control strategy includes: selecting one or more key parameters as control parameters, using the particle swarm optimization algorithm to find the optimization parameters, and constructing the optimal control strategy for a specific scenario.
8. A short-cut denitrification-anaerobic ammonium oxidation system, characterized by: It includes a data module: obtaining sample point data of a short-range denitrification-anaerobic ammonium oxidation system to construct a machine learning data set, wherein the sample point data includes operation data, operation parameters and microbial community structure data, wherein the microbial community structure data is presented in a matrix form, each column represents a species, each row represents a sample, and the element value is the relative abundance of the corresponding species. The operation data includes ammonia nitrogen concentration, nitrate nitrogen concentration, nitrite concentration, COD concentration, pH, carbon-nitrogen ratio and nitrate-ammonia nitrogen ratio. The operation parameters include system type, system volume, reflux ratio, hydraulic retention time, temperature, inoculated sludge type and carbon source type; Processing module: performs outlier processing, category information conversion, data standardization, correlation analysis, and data set splitting on the dataset. Outlier processing uses a dynamic range judgment method based on the mean and standard deviation of feature data. Data standardization uses the Standardscaler method for standard normalization. The dataset is split using a 5-fold cross-validation method to split the training set and test set in an 8:2 ratio. Model module: A prediction model is established based on the gradient boosting algorithm, integrating operating parameters, operational parameters, and microbial community structure data. SHAP value analysis is used to identify key microbial features with SHAP values above the threshold. Analysis module: Construct a multi-objective optimization model and use the particle swarm optimization algorithm to dynamically adjust the temperature, pH, carbon-nitrogen ratio and hydraulic retention time parameters to generate the optimal control strategy adapted to different influent conditions. The optimization objectives include a total nitrogen removal rate ≥ 93.5%, a removal rate fluctuation range ≤ ± 5%, and a reduction in carbon source dosage by 20%-30%.
Citation Information
Patent Citations
Intelligent denitrification regulation and control method and regulation and control system for aquaculture wastewater
CN114790039A
Microflora-based sewage treatment aeration system control method
CN114573096A
Optimization control method for sewage treatment process
CN116360366A