Method for closed-loop regulation of nano zero-valent iron coupled biological system for enhancing organic wastewater treatment by integrating kinetics fitting and machine learning ensemble tree model
By combining kinetic fitting with machine learning ensemble tree model, the kinetic coefficient k is extracted and the dosage of nano-zero valent iron is inverted, which solves the problem of inaccurate dosage control in nano-zero valent iron coupled biological system, and realizes efficient and stable removal of organic pollutants and optimization of operating costs.
Patent Information
- Application Number
- CN202610089082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-06-05
AI Technical Summary
Existing nano-zero-valent iron coupled biological systems for treating nitrobenzene wastewater suffer from poor precision in controlling the dosage of nano-zero-valent iron, leading to unstable degradation efficiency. Reliance on empirical experiments results in high dosage costs and makes it difficult to develop an executable control strategy. Furthermore, existing methods lack a framework that integrates kinetic curve information with machine learning, making it difficult to achieve intelligent and refined operation.
By combining kinetic fitting with machine learning integrated tree model, the kinetic coefficient k is extracted as a process characterization variable through kinetic fitting. Machine learning is used to predict the dosage of nano-zero valent iron and to invert the minimum dosage under a given target degradation rate. Combined with pH and carbon source regulation, a closed-loop control is formed to realize the intelligent management of nano-zero valent iron coupled biological system.
It improves the removal efficiency and stability of organic pollutants, reduces reagent consumption and operational errors, and realizes intelligent and refined operation and management of the nano-zero-valent iron coupled biological system.
Smart Images

Figure CN122144926A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water pollution control and wastewater treatment technology, specifically involving a method for enhancing organic wastewater treatment by combining dynamic fitting and machine learning integrated tree model closed-loop regulation of a nano-zero-valent iron coupled biological system. Background Technology
[0002] Among numerous technologies for treating nitroaromatic pollutants, the coupled process of nano-zero-valent iron reduction and biotransformation has become an important direction for the treatment of recalcitrant organic pollutants such as nitrobenzene due to its advantages such as fast reaction rate, wide applicability, and ability to achieve deep removal. Nitrobenzene, as a typical nitroaromatic pollutant, is highly toxic and has low direct biodegradation efficiency. In engineering, nano-zero-valent iron is typically added to provide electrons and a reducing environment, converting nitrobenzene into intermediates or products that are more easily degraded by microorganisms, thereby improving overall removal efficiency and stability. However, due to the large fluctuations in actual wastewater quality and the complexity of the system processes, traditional operation methods that rely on experience-based addition often struggle to maintain stable and efficient treatment results in the long term.
[0003] The degradation process of the nano-zero-valent iron coupled with biological systems is influenced by multiple factors, including pH, initial nitrobenzene concentration, nano-zero-valent iron particle size and dosage, available carbon source level, and reaction time. Among these, the dosage of nano-zero-valent iron is a key control factor affecting the system's reduction capacity and reaction rate: insufficient dosage leads to insufficient electron supply and difficulty in achieving the target removal rate; excessive dosage may cause particle aggregation and surface passivation, reducing effective reaction sites and potentially causing local pH changes and iron ion accumulation, thus affecting microbial activity and system stability. Simultaneously, pH and carbon source can significantly alter the electron release from nano-zero-valent iron corrosion and the metabolic state of microorganisms within certain ranges, indirectly affecting the nano-zero-valent iron dosage required to achieve the target removal rate. Due to the nonlinear interaction and threshold effect of these factors, control strategies using fixed dosage or manual experience adjustment are difficult to respond to influent fluctuations in real time, easily leading to unstable treatment efficiency, high reagent consumption, and increased operational risks.
[0004] While existing technologies include methods for predicting removal rates based on statistical regression or machine learning, most of these methods directly use the degradation rate at a fixed reaction time point as the prediction target, which has the following shortcomings: First, nitrobenzene concentration is usually obtained through timed sampling and offline analysis such as HPLC. Differences in the sampling time points and curve morphology can lead to significant label noise, making it difficult to obtain stable generalization performance when directly predicting the degradation rate. Second, directly predicting the degradation rate makes it difficult to ensure physical consistency over time, and it is difficult to directly convert the prediction results into executable control quantities such as "the minimum amount of nano-zero valent iron required to achieve the standard". Third, the core requirement of engineering operation is to form a closed-loop control around the amount of nano-zero valent iron added, and, when necessary, combine the synergistic adjustment of pH and carbon source to reduce the demand for nano-zero valent iron or operating costs. However, existing methods lack an integrated framework that connects "kinetic curve information - prediction - inversion - actuator control", thus making it difficult to meet the requirements of intelligent, low-error, and feasible operation and control on site. Summary of the Invention
[0005] To address the problems of poor precision in controlling the dosage of nano-zero-valent iron in existing nano-zero-valent iron coupled biological systems, unstable degradation efficiency of nitrobenzene under fluctuating operating conditions, high dosage costs due to reliance on empirical experiments, and difficulty in forming an executable control strategy, this invention provides a method for enhancing the treatment of organic wastewater by combining kinetic fitting and machine learning integrated tree model for closed-loop control of nano-zero-valent iron coupled biological systems. This method employs an integrated framework of "kinetic fitting—machine learning prediction—target inversion—on-site closed-loop control." It uses the kinetic coefficient k obtained from fitting the organic pollutant concentration-time curve as a process characterization variable. A target model is used to accurately predict k under different operating conditions. Given a target degradation rate, the minimum dosage of nano-zero-valent iron is inverted, thus achieving closed-loop regulation of the nano-zero-valent iron dosing pump dosage. Simultaneously, when model inversion results indicate that adjusting pH or the concentration of available carbon sources is more economical or can significantly reduce the demand for nano-zero-valent iron, the corresponding recommended pH and carbon source settings are output as a synergistic control strategy. The ultimate goal is to improve the removal efficiency and stability of organic pollutants, while reducing reagent consumption and operational errors while meeting treatment requirements, thereby achieving intelligent and refined operation and management of the nano-zero-valent iron coupled biological system.
[0006] The technical solution of the present invention is as follows:
[0007] A method for enhancing organic wastewater treatment by combining kinetic fitting and machine learning ensemble tree model for closed-loop regulation of a nano-zero-valent iron coupled biological system includes the following steps:
[0008] S1. Data Collection and Cleaning:
[0009] Experimental or field data of the nano-zero-valent iron coupled biological system under different operating conditions were collected and recorded. The data included operating parameters and organic pollutant concentrations that varied over time. The operating parameters were scale-normalized to ensure that input features of different dimensions could be trained and compared on the same scale, thereby improving model training efficiency and enhancing prediction stability. The organic pollutant concentration was converted into the degradation rate D(t), and the calculation formula is as follows:
[0010] (1),
[0011] Where C0 refers to the initial concentration of organic pollutants in the system before the reaction begins (t=0), and the unit is mmol·L. -1 (Or, according to the equivalent concentration unit determined by the detection method), this concentration is obtained by experimental determination or calibration; C t The remaining concentration of organic pollutants in the system at time t is given by the reaction. The unit is the same as that of C0 and is obtained from the quantitative detection results at the corresponding time point. D(t) refers to the degradation rate of organic pollutants at time t. It is a dimensionless parameter used to characterize the degree of removal of organic pollutants relative to the initial state.
[0012] S2. Fitting of organic pollutant degradation kinetics and sample construction:
[0013] S2-1. Perform kinetic fitting on the degradation rate-time curve of organic pollutants under the same operating condition to obtain the kinetic parameter k corresponding to that condition. The kinetic fitting preferentially uses a first-order kinetic model with a plateau term to fit the kinetic parameter k and the plateau term d. max :
[0014] (2)
[0015] When the model with platform term fails to fit or d max When it becomes unrecognizable, it degenerates into a platformless first-order dynamic model or d max Setting it to 1 enhances the robustness of the fit, where the plateau-free first-order dynamic model is:
[0016] (3),
[0017] Obtain the dynamic parameters k and platform term d corresponding to each set of working conditions. max ;
[0018] S2-2. Construct a machine learning training sample set by using the working condition parameters to form the input feature vector X and the dynamic parameter k or log(k) as the output label y.
[0019] S3. Construct a model pool and filter target models:
[0020] S3-1. Construct a machine learning model pool based on the sample set. The model pool shall contain at least two or more of the following models: multiple linear regression, support vector regression, decision tree, random forest, gradient boosting regression tree, and XGBoost.
[0021] S3-2. K-fold cross-validation is used to train and evaluate each candidate model in the model pool; based on the results of each fold cross-validation, the coefficient of determination R0 is used... 2 The mean absolute error (MAE) and root mean square error (RMSE) are used as evaluation indicators to comprehensively compare the prediction performance of each candidate model and select the target model with the best prediction performance.
[0022] S4. Model Interpretation and Key Factor Identification:
[0023] SHAP interpretation analysis is performed on the trained target model to output the average absolute SHAP value of each input feature with respect to k or log(k) and obtain the importance ranking. At the same time, the positive and negative contribution directions of each input feature to the model output are obtained to identify the influence of operating condition parameters on dynamic parameters.
[0024] S5. Prediction-Inversion Application and On-site Closed-Loop Control:
[0025] S5-1. Input the current operating condition parameters from the field or experiment into the target model, and output the predicted dynamic parameter k. pred Or log(k) pred Based on the kinetic equation, the predicted degradation rate D at a given reaction time t is calculated. pred (t);
[0026] S5-2, Set the target degradation rate D target Calculate the target degradation kinetic coefficient k with reaction time t. target =-ln(1-D target Under fixed partial operating parameters, search for conditions satisfying k / t; pred ≥k target Minimum nano-zero valent iron dosage ZVI required And calculate the additional dosage ΔZVI=max(0, ZVI) required -ZVI current );
[0027] S5-3, Recommended strategy for simultaneous output of carbon source dosage and pH adjustment: Under the constraint conditions, make the predicted k pred The recommended values are the carbon source dosage that maximizes or minimizes ΔZVI and the pH value.
[0028] S5-4. Embed the target model into the field control system, and adjust the dosage of the nano-zero valent iron dosing pump, the dosage of the carbon source dosing pump, and the acid / alkali dosing unit according to the model output to achieve closed-loop control of the organic pollutant degradation process; and record the adjusted organic pollutant concentration and operation data back to form a closed-loop update.
[0029] Furthermore, in step S1, the organic pollutant is a common organic compound in the art, including but not limited to nitrobenzene.
[0030] Furthermore, in step S1, the operating parameters include at least: initial pH value, initial concentration of organic pollutants, nano-zero ferric iron particle size, nano-zero ferric iron dosage, available carbon source dosage, and reaction time; wherein the nano-zero ferric iron dosage is characterized by the dosage rate of the nano-zero ferric iron dosing pump or the cumulative dosage mass, the carbon source dosage is characterized by the dosage rate of the carbon source dosing pump or the cumulative dosage mass, and the pH is adjusted by the acid / alkali dosing unit.
[0031] Furthermore, in step S1, the concentration of organic pollutants is obtained by timed sampling and analysis using high-performance liquid chromatography (HPLC). The sampling time points include at least 0 h and multiple time points during the reaction process, preferably more than 5, to ensure the formation of a complete organic pollutant concentration-time curve. Even further, the sampling time interval is 6–24 h, more preferably 12 h, and the sampling time covers at least the kinetic range of 0–72 h to ensure the stability of the fitting of the kinetic parameter k.
[0032] Furthermore, in step S1, missing value processing, outlier identification and correction are performed on the collected data; the degradation rate is truncated to the [0,1] interval to suppress fitting failure caused by measurement noise; outlier identification adopts the box plot IQR rule or the threshold rule based on regression residuals; data points that are obviously recorded errors or measurement failures are removed, and reasonable but extreme data points are truncated or interpolated according to preset rules to obtain a training sample set with stable quality.
[0033] Furthermore, in step S1, Z-score standardization is used to normalize the operating parameters that are approximately normally distributed, and min-max normalization is used to map them to the [0,1] interval for normalization. Correlation analysis is performed on the operating parameters to identify highly correlated redundant variables. When the absolute value of the correlation coefficient between any two variables is higher than a preset threshold, variables with weak explanatory power or poor engineering controllability are removed to reduce redundancy and improve the interpretability and robustness of the model.
[0034] Furthermore, in step S3, Z-score standardization is uniformly adopted for the model pool screening stage to ensure the comparability of different algorithms.
[0035] Furthermore, in step S2-2, the output label y is log(k) to reduce the heteroscedasticity caused by the change of k across orders of magnitude and improve the model's generalization ability.
[0036] Furthermore, in step S5-2, the minimum nano-zero valent iron dosage ZVI required Solve using discrete scanning or monotonic search; when k pred When the amount of nano-zero valent iron added does not decrease monotonically, a binary search is preferred to determine ZVI. required .
[0037] Furthermore, in step S5-4, the main actuator of the on-site closed-loop control is the nano-zero-valent iron dosing pump, and the controlled quantity is the nano-zero-valent iron dosage. The acid / alkali dosing unit and the carbon source dosing pump serve as co-regulatory actuators. When the model inversion results show that adjusting the pH setpoint and / or the carbon source dosage under given constraints can reduce the nano-zero-valent iron dosage required to achieve the target degradation rate or reduce the overall operating cost, the co-regulatory actuators are set or adjusted according to the recommended pH setpoint and / or recommended carbon source dosage obtained from the model inversion.
[0038] Furthermore, step S5 also includes S5-5, introducing a safety margin factor m to obtain the recommended dosage ZVI of nano-zero valent iron. recommend =m·ZVI required m is set to 1.05~1.50, preferably 1.10~1.30, and the "minimum dosing curve" and "safety margin curve" are output within the recommended range.
[0039] Furthermore, step S5 also includes S5-6, continuously collecting organic pollutant concentration data and operating condition data during on-site operation, and retraining or incrementally updating the model according to a preset cycle to improve the predictive and control adaptability under different influent loads, pH fluctuations and carbon source changes. The update cycle is once a week, once a month or once a quarter.
[0040] This invention also provides a system for enhancing the treatment of organic wastewater by combining kinetic fitting and machine learning ensemble tree model for closed-loop regulation of a nano-zero-valent iron coupled biological system, comprising:
[0041] The data acquisition and analysis module is used to collect experimental or field data of the nano-zero-valent iron coupled biological system under different operating conditions. The data includes operating parameters and organic pollutant concentrations. The operating parameters are scaled and normalized, and the organic pollutant concentrations are converted into degradation rates.
[0042] The kinetics fitting module is used to perform kinetics fitting on the degradation rate-time curve of organic pollutants under the same operating conditions.
[0043] The machine learning modeling module is used to train and evaluate a sample set constructed from input feature vectors consisting of working condition parameters and output labels consisting of dynamic parameters k or log(k), and to select the target model with the best prediction performance.
[0044] The inversion decision module is used to input the current operating parameters of the field or experiment into the target model and output the recommended strategy of minimum nano-zero valent iron dosage or carbon source dosage and pH adjustment.
[0045] The dosing control module is used to adjust the dosage of the nano-zero valent iron dosing pump, the dosage of the carbon source dosing pump, and the acid / alkali dosing unit according to the model output, so as to achieve closed-loop control of the degradation process of organic pollutants.
[0046] Compared with the prior art, the present invention has the following advantages:
[0047] (1) Characterizing kinetic parameters improves label consistency and prediction stability:
[0048] Unlike methods that directly use the degradation rate of organic pollutants at a specific time point as the prediction target, this invention first performs kinetic fitting based on the organic pollutant concentration-time curve, extracting the kinetic coefficient k (or log(k)) as a process characterization variable, thus transforming the model's learning object from an outcome quantity to a process quantity. This approach can reduce the impact of sampling time point differences and measurement noise on the tags, and can calculate the degradation rate at any reaction time through the kinetic equation, thereby improving the consistency, physical rationality, and cross-condition generalization ability of the prediction results.
[0049] (2) Model pooling and XGBoost integration improve accuracy and robustness:
[0050] This invention employs a model pooling strategy to uniformly train and cross-validate models such as multiple linear regression, support vector regression, decision trees, random forests, and gradient boosting regression trees, and selects the model with the best performance as the target prediction model. It can effectively characterize the nonlinear coupling relationship between multiple factors such as the amount of nano-zero valent iron added, particle size, pH, carbon source and pollution load, and can still achieve high prediction accuracy and stronger robustness under limited data conditions.
[0051] (3) The inversion solution outputs the executable minimum dosage decision, which directly serves the engineering control:
[0052] This invention uses the k predicted by the constructed target model pred With the target degradation rate D target By correlating the kinetic equations, an inversion solution chain of "target-parameter-dosage" is constructed, which enables the minimum nano-zero valent iron dosage ZVI required to achieve the target under given reaction time and constraints to be obtained through inversion. requiredThe additional dosage ΔZVI is calculated. This inversion result can be directly converted into an engineering dosage strategy, avoiding the shortcomings of traditional methods that only provide predicted values but cannot form executable control quantities.
[0053] (4) The main closed-loop control focuses on the dosage of nano-zero valent iron, reducing operational errors and improving removal stability:
[0054] This invention uses the dosage of nano-zero valent iron as the core control variable in on-site closed-loop control. The nano-zero valent iron dosing pump is automatically adjusted through model inversion output, and rolling feedback is formed by combining subsequent sampling and HPLC detection results to achieve dynamic optimization of the organic pollutant degradation process. This closed-loop control can significantly reduce the errors and lags caused by human experience intervention, and improve the system's operational stability and compliance reliability under fluctuations in influent and changes in operating conditions.
[0055] (5) pH and carbon source are used as synergistic optional regulatory variables to achieve drug consumption optimization under more economical conditions:
[0056] This invention does not treat pH and carbon source as essential master variables, but rather as conditional variables and optional co-regulatory variables in the inversion solution: when model inversion shows that adjusting the pH setpoint and / or carbon source level within a feasible range can reduce ZVI. required When reducing overall operating costs, the system outputs corresponding recommended settings and can implement coordinated adjustments. This strategy, while ensuring processing effectiveness, provides an optimized path for engineering operations to achieve goals with lower consumption of nano-zero valent iron, improving both economic efficiency and operability.
[0057] (6) Multi-format visualization output supports deployment design, scaling up, and operation and maintenance decisions:
[0058] This invention can output feature screening and correlation maps, performance comparison results of different models, SHAP feature interpretation maps, as well as degradation rate visualization results under given operating conditions, additional dosage heatmaps, and minimum dosage design curves (including safety margin curves). These outputs can be used to interpret key control factors and their positive and negative effects, and also for engineering scale-up and operating parameter formulation, providing intuitive decision-making basis for the digital management of nano-zero-valent iron coupled biological systems. Attached Figure Description
[0059] Figure 1 This is a flowchart of the overall process of the method of the present invention.
[0060] Figure 2 The input feature distribution (box plot) and the Pearson correlation heatmap of the input variables are used.
[0061] Figure 3 This is a graph comparing the prediction performance of different machine learning models on the test set.
[0062] Figure 4 This is a schematic diagram illustrating the SHAP interpretation results of the target prediction model.
[0063] Figure 5 The operating surface for predicting the kinetic coefficient k obtained from the target prediction model under fixed conditions and the degradation efficiency calculated from k (for different reaction times).
[0064] Figure 6 Two-dimensional thermograms of the additional nano-zero-valent iron dosage required to achieve different target degradation efficiency thresholds.
[0065] Figure 7 Design curves and safety margin recommendations for minimum nano-zero valent iron addition under different target degradation efficiency thresholds. Detailed Implementation
[0066] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0067] Example 1
[0068] This embodiment takes the degradation of nitrobenzene pollutants by a nano-zero-valent iron coupled biological system as an example. It describes a method for enhancing organic wastewater treatment by combining kinetic fitting and a machine learning integrated tree model to close-loop control of the nano-zero-valent iron coupled biological system. The method is carried out in the following steps:
[0069] 1. Data collection and organization
[0070] (1) Setting the amount of nano-zero valent iron (ZVI) added to the reaction system concentration pH setpoint (pH), carbon source dosage (carbon) concentration ), and record the initial concentration of nitrobenzene (NB) concentration ), nano-zero valent iron particle size (ZVI) size The dosage of nano-zero valent iron is controlled by a nano-zero valent iron dosing pump, the pH is adjusted by an acid / base dosing unit, and the carbon source concentration is controlled by a carbon source dosing pump.
[0071] (2) p-nitrobenzene concentration C t The concentration-time curves were obtained by timed sampling and analysis by HPLC. The sampling time points t were 0 h, 12 h, 24 h, 36 h, 48 h, 60 h, and 72 h, in order to obtain complete concentration-time curves and ensure the stability of kinetic fitting.
[0072] (3) Unify the encoding and format of the sampled data, and construct a dataset grouped by working condition. Each data group includes: working condition number group, time, C t and the corresponding operating parameters ZVI concentration pH, carbonconcentration NB concentration ZVI size .
[0073] 2. Data cleaning and scaling
[0074] (1) Missing value screening and outlier identification were performed on the nitrobenzene concentration data obtained from sampling. Outlier identification adopted the box plot IQR rule or the 3-standard-deviation threshold method: when the observed value exceeded the range of [Q1-1.5×IQR, Q3+1.5×IQR] or deviated from the mean by more than 3 standard deviations, the observed value was judged as outlier. Data points that were obviously caused by recording errors or abnormal HPLC peaks were removed, and reasonable but extreme data points were truncated to the boundary value according to the preset rules. The distribution characteristics of each variable and the outlier identification results are as follows: Figure 2 As shown in (a).
[0075] (2) Convert the concentration sequence into a degradation rate sequence D(t) = 1 - C t / C0, and trim D(t) to the [0,1] interval to avoid fitting instability caused by noise; where C0 is the initial concentration of nitrobenzene at t=0.
[0076] (3) The working condition parameters used for subsequent modeling input features are scaled and normalized; Z-score standardization is used for the approximately normal distribution features to normalize the data to mean 0 and standard deviation 1; min-max normalization is used for the skewed distribution features to scale the data to the range of [0,1]; ensure that data of different dimensions are operated on the same scale to improve model training efficiency and enhance prediction stability.
[0077] (4) Calculate the correlation coefficients between the input features and draw a correlation heatmap to identify redundant features. Remove features with weak explanatory power or poor engineering controllability to obtain the final feature set for modeling, including ZVI. concentration pH, carbon concentration NB concentration and ZVI size The results of the correlation analysis are as follows: Figure 2 As shown in (b).
[0078] 3. The degradation kinetics of nitrobenzene were fitted to obtain k.
[0079] (1) Dynamic fitting is performed on the D(t) curve for each working condition parameter, using a first-order dynamic model with plateau terms:
[0080]
[0081] Where k is the dynamic coefficient, d max For platform items (maximum achievable degradation rate).
[0082] (2) Calculate the goodness-of-fit index for the fitting results, including the root mean square error (RMSE) and the coefficient of determination (R²). 2 When R 2 If the values are below a preset threshold or the fitting parameters do not meet the physical constraint k>0, the condition is marked as a fitting failure and removed from the model, and will not be included in subsequent machine learning modeling.
[0083] (3) Construct the input feature vector X=[ZVI] using the operating parameters. concentration carbon concentration ,NB concentration pH, ZVI size The output label y is constructed using the dynamic parameter k or log(k); log(k) is used as the modeling objective to reduce the heteroscedasticity caused by k varying across orders of magnitude, thereby improving the model's generalization performance.
[0084] This embodiment completes the extraction process from offline HPLC curves to kinetic coefficient k, and forms a sample set that can be used for machine learning training, providing a data foundation for subsequent modeling and inverse control.
[0085] Example 2
[0086] Based on the sample set obtained in Example 1, model pool training and screening were performed to determine the target model, and key influencing factors were interpreted and analyzed, following these steps:
[0087] 1. Model Pool Construction and Training
[0088] (1) The sample set is randomly divided into a training set and a test set, with the training set accounting for 80% and the test set accounting for 20%. To avoid randomness, k-fold cross-validation is used to evaluate the model performance, with k=5.
[0089] (2) Construct a machine learning model pool, which includes: multiple linear regression (MLR), support vector regression (SVR), decision tree regression (DT), random forest regression (RF), gradient boosting regression tree (GBRT) and extreme gradient boosting regression (XGBoost); each model performs hyperparameter optimization on the training set and completes training.
[0090] (3) Calculate the predictive performance index of each model on the target variable log(k) on the test set, including the coefficient of determination R. 2 Mean absolute error (MAE) and root mean square error (RMSE); based on "R 2 The optimal model is determined based on the principle of "higher accuracy and smaller error"; the comparison results of the prediction performance of different models on the test set are as follows: Figure 3 As shown.
[0091] 2. Select target models and define KiML-XGB
[0092] (1) Comparing the performance of each model in the model pool, the XGBoost model with the best test set index was selected as the target prediction model; the results and its test set fitting performance are as follows. Figure 3 As shown.
[0093] (2) The joint framework of "kinetics-fitting + model-pool selection + XGBoost" is defined as KiML-XGB, which is used to output the dynamic coefficient k under a given working condition. pred Or log(k) pred .
[0094] 3. SHAP Interpretation and Key Factor Identification
[0095] (1) Import the trained XGBoost model into the SHAP interpretation framework and calculate the marginal contribution of each input feature to the model output log(k);
[0096] (2) Using the mean absolute SHAP value as the global importance index, the importance ranking of features was obtained; and a SHAP scatter plot was drawn to identify the positive and negative contribution directions and nonlinear influence ranges of features such as the amount of nano-zero valent iron added, the initial concentration of nitrobenzene, particle size, pH, and carbon source to log(k); the global importance ranking and scatter plot results of SHAP are as follows: Figure 4 As shown.
[0097] (3) Based on the SHAP interpretation results, the priority control variable was determined to be the amount of nano-zero valent iron added; pH and carbon source were used as co-selectable control variables to assess the feasibility of "reducing ZVI demand or reducing costs" during the inversion stage; key influencing factors and their contribution directions were determined by... Figure 4 The SHAP result shown is confirmed.
[0098] Example 3
[0099] Based on the KiML-XGB model obtained in Example 2, the target degradation rate was inverted and the on-site dosing device was driven to form a closed-loop control, as follows:
[0100] 1. Prediction of kinetic parameters and calculation of degradation rate under given working conditions
[0101] (1) Obtain the current operating condition parameter vector: ZVI concentration carbon concentration NB concentration pH, ZVI size .
[0102] Input the above operating parameters into the KiML-XGB model and output log(k). pred And converted to: k pred =exp(log(k) pred ).
[0103] (2) Under a given reaction time t, k is calculated based on the kinetic model. pred Converted to predicted degradation rate:
[0104] When using a platformless first-order dynamics model:
[0105]
[0106] When using a first-order dynamic model with a plateau term:
[0107]
[0108] Where d max Take the d obtained from the dynamic fitting of this working condition in Example 1 max ; in d max If the platform item is missing or not enabled, take d. max =1.0.
[0109] The degradation rate D in this step pred (t) The dynamic coefficient k predicted by the model pred The results are obtained through calculation using kinetic equations, thus avoiding problems such as inconsistent time dimensions and label drift caused by changes in sampling points that may result from direct prediction of degradation rates, and allowing the prediction results to be extrapolated to different reaction times.
[0110] (3) By Figure 5 As shown, k is based on the output of KiML-XGB. pred By substituting these values into the kinetic equations, operating surfaces with varying reaction times can be obtained under fixed conditions such as pH, carbon source concentration, and particle size. Figure 5 (a) is the predicted degradation rate operating surface at (t=36h). Figure 5 (b) is the predicted degradation rate operating surface at (t=72h). Figure 5 (c) is the predicted degradation rate operating surface at t=108h; the operating surface uses the initial concentration of nitrobenzene and the amount of nano-zero valent iron added as the main independent variables, and the output is the D value at the corresponding reaction time. pred (t) is used to intuitively characterize the range of influence of addition strategies on removal efficiency and the nonlinear response characteristics under different pollution loads.
[0111] 2. Inversion of the minimum amount of nano-zero-valent iron required to achieve the target degradation rate
[0112] (1) Set the target degradation rate D target With the target reaction time t target Calculate the target dynamic coefficients based on the dynamic equations:
[0113] When using a platformless model:
[0114]
[0115] When using a model with platform terms, (D) is satisfied. target ≤ d max Constraints are applied, and the following formula is used for calculation:
[0116]
[0117] (2) In the fixed NB concentration ZVI size Under conditions where variables are uncontrollable or not adjusted for the time being, ZVI concentration Perform discrete scan, monotonic search, or binary search; for each candidate ZVI concentration This, along with other fixed operating conditions, forms the input feature vector (X), which is then input into KiML-XGB to obtain k. pred With k pred ≥k t Using `arget` as the criterion for compliance, find the minimum ZVI that satisfies the compliance condition. required .
[0118] (3) Calculate the additional dosage:
[0119]
[0120] It also outputs an "additional dosage heatmap" and a "minimum dosage design curve." Furthermore, for conservative engineering design, it provides a safety margin line:
[0121] ZVI required,safe =ZVI required (1+ )
[0122] in To preset the safety margin factor, select =5%.
[0123] (4) Figure 6 The following is given in t target Under 72h conditions, different D target Threshold corresponding Visualization of inversion results: Figure 6 (a) is a heatmap of the additional dosage when D(72h)≥0.85. Figure 6(b) is a heat map of the additional dosage when D(72h)≥0.90. Figure 6 (c) Heatmap of additional dosage when D(72h) ≥ 0.95. The heatmap is defined as "current nano-zero valent iron dosage (ZVI)". current ")" and "initial concentration of nitrobenzene (NB)" concentration Using ")" as the coordinate axis, the output ΔZVI required to meet the standard is used to quickly determine whether additional addition is needed under different pollution loads and the current addition level, as well as the order of magnitude of the addition.
[0124] (5) Figure 7 The minimum nano-zero-valent iron dosage required to achieve the target threshold at different initial nitrobenzene concentrations is presented, along with a safety margin curve: Figure 7 The solid line represents the minimum acceptable curve ZVI. required The dashed / dotted line represents the safety margin curve ZVI. required,safe This curve can be directly used as the basis for engineering design: given NB concentration With target D target The threshold value can be used to read the minimum dosage required to meet the standards, and at the same time, it provides a recommended value for a conservative safety margin.
[0125] 3. On-site closed-loop control execution and recommendation of cooperating variables
[0126] (1) The ZVI obtained by inversion required or The control commands are converted into control commands for the nano-zero valent iron dosing pump. The control actuators include only the nano-zero valent iron dosing pump, the carbon source dosing pump, and the acid / alkali dosing unit. The control quantities are the nano-zero valent iron dosage, the carbon source dosage, and the pH setpoint, respectively. Among them, the nano-zero valent iron dosage is executed as the main closed-loop control quantity.
[0127] (2) When the model inversion results show that adjusting the pH setpoint and / or carbon source dosage within a feasible range can significantly reduce ZVI required To reduce overall operating costs, the system outputs the corresponding recommended pH and carbon source settings; these settings are then coordinated with the acid / alkali dosing unit and the carbon source dosing pump; and these adjustments are used as optimization suggestions for subsequent operating strategy development.
[0128] (3) After completing one addition adjustment, the reaction system is sampled periodically according to the preset sampling cycle, and the nitrobenzene concentration is determined by HPLC. The C200 concentration is then updated. t With D(t), new data is backfilled into the dataset for rolling correction and model updates, thus forming a closed-loop optimization process of "prediction-inversion-addition-verification-relearning" to improve the stability and adaptability of long-term operation.
[0129] (4) In summary, Figures 5-7Results that can be directly applied to engineering decisions are presented from three levels: "prediction (operational aspect), inversion (heat map of additional dosage), and design (minimum dosage curve and safety margin)". Figure 5 Used to demonstrate the achievable degradation efficiency range and sensitive areas at different reaction times; Figure 6 Used for quickly querying supplementary requirements at a given target threshold; Figure 7 Used to generate injection design curves and conservative margin lines that can be submitted to operators. Figures 5-7 As can be seen, the present invention can output an executable dosing decision under a given operating condition and provide a stable dosing strategy to meet the standards under different pollution load conditions, thereby improving the degradation efficiency of nitrobenzene and reducing the dosing cost, and has the feasibility of on-site closed-loop control.
Claims
1. A method for enhancing organic wastewater treatment by combining kinetic fitting and machine learning ensemble tree model to achieve closed-loop regulation of a nano-zero-valent iron coupled biological system, characterized in that... Includes the following steps: S1. Data Collection and Cleaning: Experimental or field data of the nano-zero-valent iron coupled biological system under different operating conditions were collected and recorded. The data included operating parameters and organic pollutant concentrations that varied over time. The operating parameters were scale-normalized to ensure that input features of different dimensions could be trained and compared on the same scale, thereby improving model training efficiency and enhancing prediction stability. The organic pollutant concentration was converted into the degradation rate D(t), and the calculation formula is as follows: (1), Where C0 refers to the initial concentration of organic pollutants in the system at t=0; C t The remaining concentration of organic pollutants in the system at time t indicates the reaction progress; D(t) refers to the degradation rate of organic pollutants at time t, which is a dimensionless parameter. S2. Fitting of organic pollutant degradation kinetics and sample construction: S2-1. Perform kinetic fitting on the degradation rate-time curve of organic pollutants under the same operating condition to obtain the kinetic parameter k corresponding to that condition. The kinetic fitting preferentially uses a first-order kinetic model with a plateau term to fit the kinetic parameter k and the plateau term d. max : (2) When the model with platform term fails to fit or d max When it becomes unrecognizable, it degenerates into a platformless first-order dynamic model or d max Setting it to 1 enhances the robustness of the fit, where the plateau-free first-order dynamic model is: (3), Obtain the dynamic parameters k and platform term d corresponding to each set of working conditions. max ; S2-2. Construct a machine learning training sample set by using the working condition parameters to form the input feature vector X and the dynamic parameter k or log(k) as the output label y. S3. Construct a model pool and filter target models: S3-1. Construct a machine learning model pool based on the sample set. The model pool shall contain at least two or more of the following models: multiple linear regression, support vector regression, decision tree, random forest, gradient boosting regression tree, and XGBoost. S3-2. K-fold cross-validation is used to train and evaluate each candidate model in the model pool; based on the results of each fold cross-validation, the coefficient of determination R0 is used... 2 The mean absolute error (MAE) and root mean square error (RMSE) are used as evaluation indicators to comprehensively compare the prediction performance of each candidate model and select the target model with the best prediction performance. S4. Model Interpretation and Key Factor Identification: SHAP interpretation analysis is performed on the trained target model to output the average absolute SHAP value of each input feature with respect to k or log(k) and obtain the importance ranking. At the same time, the positive and negative contribution directions of each input feature to the model output are obtained to identify the influence of operating condition parameters on dynamic parameters. S5. Prediction-Inversion Application and On-site Closed-Loop Control: S5-1. Input the current operating condition parameters from the field or experiment into the target model, and output the predicted dynamic parameter k. pred Or log(k) pred Based on the kinetic equation, the predicted degradation rate D at a given reaction time t is calculated. pred (t); S5-2, Set the target degradation rate D target Calculate the target degradation kinetic coefficient k with reaction time t. target =-ln(1-D target Under fixed partial operating parameters, search for conditions satisfying k / t; pred ≥k target Minimum nano-zero valent iron dosage ZVI required And calculate the additional dosage ΔZVI=max(0, ZVI) required -ZVI current ); S5-3, Recommended strategy for simultaneous output of carbon source dosage and pH adjustment: Under the constraint conditions, make the predicted k pred The recommended values are the carbon source dosage that maximizes or minimizes ΔZVI and the pH value. S5-4. Embed the target model into the field control system, and adjust the dosage of the nano-zero valent iron dosing pump, the dosage of the carbon source dosing pump, and the acid / alkali dosing unit according to the model output to achieve closed-loop control of the organic pollutant degradation process; and record the adjusted organic pollutant concentration and operation data back to form a closed-loop update.
2. The method according to claim 1, characterized in that, In step S1, the organic pollutant is nitrobenzene; the operating parameters include at least: initial pH value, initial concentration of organic pollutant, nano-zero ferric iron particle size, nano-zero ferric iron dosage, available carbon source dosage, and reaction time; the concentration of organic pollutant is obtained by timed sampling and analysis by HPLC, and the sampling time points include at least 0 h and multiple time points during the reaction process, with a total of more than 5 sampling time points.
3. The method according to claim 2, characterized in that, The sampling time interval is 6 to 24 hours, preferably 12 hours, and the sampling time covers at least the kinetic range of 0 to 72 hours.
4. The method according to claim 1, characterized in that, In step S1, missing value processing, outlier identification and correction are performed on the collected data; the degradation rate is truncated to the [0,1] interval to suppress fitting failure caused by measurement noise; outlier identification adopts the box plot IQR rule or the threshold rule based on regression residuals; data points that are obviously recorded errors or measurement failures are removed, and reasonable but extreme data points are truncated or interpolated according to preset rules to obtain a training sample set with stable quality.
5. The method according to claim 1, characterized in that, In step S1, Z-score standardization is used to normalize the operating parameters that are approximately normally distributed, and min-max normalization is used to map them to the [0,1] interval for normalization. Correlation analysis is performed on the operating parameters to identify highly correlated redundant variables. When the absolute value of the correlation coefficient between any two variables is higher than a preset threshold, variables with weak explanatory power or poor engineering controllability are removed. In step S3, Z-score standardization is uniformly used in the model pool screening stage to ensure the comparability of different algorithms.
6. The method according to claim 1, characterized in that, In step S5-2, the minimum nano-zero valent iron dosage ZVI required Solve using discrete scanning or monotonic search; when k pred When the amount of nano-zero valent iron added does not decrease monotonically, a binary search is preferred to determine ZVI. required .
7. The method according to claim 1, characterized in that, In step S5-4, the main actuator of the on-site closed-loop control is the nano-zero-valent iron dosing pump, and the controlled quantity is the nano-zero-valent iron dosage. The acid / alkali dosing unit and the carbon source dosing pump are used as co-regulatory actuators. When the model inversion results show that adjusting the pH setpoint and / or the carbon source dosage under given constraints can reduce the amount of nano-zero-valent iron required to achieve the target degradation rate or reduce the overall operating cost, the co-regulatory actuators are set or adjusted according to the recommended pH setpoint and / or recommended carbon source dosage obtained from the model inversion.
8. The method according to claim 1, characterized in that, Step S5 also includes S5-5, introducing a safety margin factor m to obtain the recommended dosage ZVI of nano-zero valent iron. recommend =m·ZVI required m is taken as 1.05~1.50, preferably 1.10~1.
30.
9. The method according to claim 1, characterized in that, Step S5 also includes S5-6, continuously collecting organic pollutant concentration data and operating condition data during on-site operation, and retraining or incrementally updating the model according to a preset cycle, which is once a week, once a month, or once a quarter.
10. A system for enhancing organic wastewater treatment by combining kinetic fitting and machine learning ensemble tree model for closed-loop regulation of a nano-zero-valent iron coupled biological system, characterized in that: include: The data acquisition and analysis module is used to collect experimental or field data of the nano-zero-valent iron coupled biological system under different operating conditions. The data includes operating parameters and organic pollutant concentrations. The operating parameters are scaled and normalized, and the organic pollutant concentrations are converted into degradation rates. The kinetics fitting module is used to perform kinetics fitting on the degradation rate-time curve of organic pollutants under the same operating conditions. The machine learning modeling module is used to train and evaluate a sample set constructed from input feature vectors consisting of working condition parameters and output labels consisting of dynamic parameters k or log(k), and to select the target model with the best prediction performance. The inversion decision module is used to input the current operating parameters of the field or experiment into the target model and output the recommended strategy of minimum nano-zero valent iron dosage or carbon source dosage and pH adjustment. The dosing control module is used to adjust the dosage of the nano-zero valent iron dosing pump, the dosage of the carbon source dosing pump, and the acid / alkali dosing unit according to the model output, so as to achieve closed-loop control of the degradation process of organic pollutants.