A Method for Estimating Parameters of a Mathematical Model for Wastewater Treatment Process Based on Random Forest-Guided Genetic Algorithm
Patent Information
- Application Number
- CN202611000960.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-11
AI Technical Summary
[0007]为了克服现有技术的上述缺点与不足,本发明的目的在于提供一种基于随机森林引导遗传算法的污水处理工艺数学模型参数估计方法,以解决现有技术中污水处理工艺数学模型参数估计过程中存在的计算成本高、寻优效率低以及易受局部最优影响等问题
[0045]1、提高参数寻优效率:本发明在遗传算法迭代过程中引入随机森林反向代理模型,利用历史真实计算样本学习参数组合与目标函数值之间的映射关系,能够在大量候选参数组合中快速筛选出更有潜力的高适应度候选参数组合,从而减少盲目搜索,提高参数寻优效率;
Smart Images

Figure CN122735489A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wastewater treatment process modeling, parameter identification and intelligent optimization technology, specifically involving a method for estimating parameters of a mathematical model of wastewater treatment process based on a random forest-guided genetic algorithm. Background Technology
[0002] Wastewater treatment plant operation involves complex physical, chemical, and biological reactions. To simulate operation, predict performance, optimize operation, and make control decisions for wastewater treatment processes, it is typically necessary to establish corresponding mathematical models of the wastewater treatment process. These mathematical models can describe the transformation patterns of pollutants in each treatment unit and are important tools for the analysis and optimization of wastewater treatment systems.
[0003] In practical applications, wastewater treatment process mathematical models typically include multiple parameters, including but not limited to activated sludge kinetic parameters, stoichiometric parameters, and parameters related to hydraulic processes. Due to significant fluctuations in wastewater quality, complex process operating conditions, and coupling effects between treatment units, parameters can vary considerably between different wastewater treatment plants, and even within the same plant at different operational stages. Therefore, to ensure that the mathematical model accurately reflects the actual operating characteristics of the target wastewater treatment system, it is usually necessary to estimate or calibrate the model parameters using measured operational data.
[0004] In existing technologies, the main methods for estimating parameters of mathematical models for wastewater treatment processes include manual adjustment methods, local optimization methods, and intelligent optimization methods. Manual adjustment methods rely on the experience of technical personnel, resulting in low efficiency, strong subjectivity, and poor repeatability. While local optimization methods can achieve relatively fast convergence under certain conditions, they are prone to getting trapped in local optima and have weak adaptability to complex nonlinear and multi-parameter coupled problems. Intelligent optimization methods, such as genetic algorithms, possess strong global search capabilities and have been applied to parameter estimation of mathematical models for wastewater treatment processes. However, when there are many parameters to be calibrated, the model structure is complex, and a single model simulation takes a long time, traditional genetic algorithms still require a large amount of computation on the real model, resulting in high computational cost and low iterative efficiency.
[0005] On the other hand, surrogate model methods can utilize existing samples to learn the mapping relationship between parameters and the objective function, enabling rapid evaluation of candidate solutions with lower computational cost. Among these, the random forest model possesses strong nonlinear fitting capabilities, good robustness to sample noise, and relatively simple implementation, making it suitable for approximate modeling in complex engineering problems. However, if a surrogate model is used solely to replace the true mechanistic model for parameter evaluation, prediction errors may cause search bias, affecting the accuracy and reliability of the final parameter estimation results.
[0006] Therefore, how to reduce the number of calculations in the real model and improve the efficiency of parameter optimization by making full use of surrogate models while ensuring the accuracy of parameter estimation has become an urgent technical problem to be solved in the field of parameter estimation of mathematical models for wastewater treatment processes. Summary of the Invention
[0007] In order to overcome the above-mentioned shortcomings and deficiencies of the prior art, the purpose of this invention is to provide a method for estimating parameters of a mathematical model of a wastewater treatment process based on a random forest-guided genetic algorithm, so as to solve the problems of high computational cost, low optimization efficiency and susceptibility to local optima in the parameter estimation process of a mathematical model of a wastewater treatment process in the prior art.
[0008] The objective of this invention is achieved through the following technical solution:
[0009] This invention provides a method for estimating parameters of a mathematical model for wastewater treatment processes based on a random forest-guided genetic algorithm, comprising the following steps:
[0010] Obtain historical operating data of the wastewater treatment system, and construct a sample dataset for parameter calibration based on the historical operating data;
[0011] A mathematical model of the wastewater treatment process is established, and the set of parameters to be calibrated for the mathematical model of the wastewater treatment process and their value range are determined.
[0012] Using WNSE, which characterizes the fit between the model's calculated values and the measured values, as the objective function, and taking the maximization of the objective function as the parameter estimation objective, the parameters of the genetic algorithm are set, an initial population is generated, and iterative optimization is performed.
[0013] During the iterative optimization process, a random forest reverse proxy model is constructed based on the historical sample set of the genetic algorithm, and the random forest reverse proxy model is used to predict candidate optimal parameter combinations; the random forest reverse proxy model takes the objective function as input and the parameters as output;
[0014] The candidate optimal parameter combination is used to update the population of the genetic algorithm, and the iterative process is repeated until the termination condition is met. The optimal parameter combination is then output as the parameter estimation result of the mathematical model of the wastewater treatment process.
[0015] In some embodiments, the specific steps of the parameter estimation method for the mathematical model of wastewater treatment process based on random forest-guided genetic algorithm are as follows:
[0016] Step 1: Obtain historical operating data of the wastewater treatment system. The historical operating data includes at least influent water quality data, operating condition data, and effluent monitoring data. Based on the historical operating data, construct a sample dataset for parameter calibration.
[0017] Step 2: Establish a mathematical model for the wastewater treatment process. The mathematical model for the wastewater treatment process includes a water quality mechanism model for describing the pollutant transformation process and a hydraulic model for describing the hydraulic process of the treatment structure.
[0018] Step 3: Determine the set of parameters to be calibrated, and set the value range of each parameter, the population size of the genetic algorithm, the crossover probability, the mutation probability, the maximum number of iterations, and the fitness evaluation index. Then, determine the simulation target based on the actual test data.
[0019] Step 4: Use WNSE as the objective function and estimate the objective using the maximization of the objective function as a parameter;
[0020] Step 5: Latin hypercube sampling is used within the range of values of each parameter to be estimated to generate a parameter sample set, and the parameter sample set is encoded as the initial population of the genetic algorithm.
[0021] Step 6: Calculate the objective function value for each individual in the initial population using the mathematical model of the wastewater treatment process, and determine the individual fitness using the objective function value;
[0022] Step 7: Perform genetic operations on the population based on individual fitness to generate the next generation population;
[0023] Step 8: Based on the historical sample set obtained during the iteration of the genetic algorithm, construct a random forest reverse proxy model, train it, and predict the candidate optimal parameter combination corresponding to the maximum objective function value.
[0024] Step 9: Replace the low-fit individuals in the next generation of the genetic algorithm with the candidate optimal parameter combinations generated in Step 8 to update the population;
[0025] Step 10: Repeat steps 7 to 9 until the preset termination condition is met, and output the optimal parameter combination as the parameter estimation result of the mathematical model of the wastewater treatment process.
[0026] In some embodiments, in step 1, the influent water quality data includes one or more of chemical oxygen demand (COD), ammonia nitrogen, total nitrogen, total phosphorus, and suspended solids; the operating condition data includes one or more of flow rate, dissolved oxygen, sludge age, return ratio, aeration intensity, temperature, and pH; and the effluent monitoring data includes one or more of COD, ammonia nitrogen, total nitrogen, total phosphorus, and suspended solids.
[0027] In some embodiments, in step 2, the water quality mechanism model includes at least one of an activated sludge model and an anaerobic digestion model; the hydraulic model adopts a series model of multiple completely mixed reactors and / or a plug flow model.
[0028] In some embodiments, in step 3, the set of parameters to be calibrated includes one or more of the kinetic parameters, chemometric parameters, and operation-related parameters in the wastewater treatment mathematical model.
[0029] Preferably, in step 4, the WNSE is defined as:
[0030]
[0031] in:
[0032] n represents the number of simulated targets;
[0033] m represents the number of samples corresponding to the i-th simulated target;
[0034] Let represent the weight of the i-th simulated target, and ≥0 and ;
[0035] This represents the measured value of the i-th simulated target at the j-th sampling time;
[0036] This represents the model-calculated value of the i-th simulated target at the j-th sampling time.
[0037] This represents the average value of the i-th simulated measured value.
[0038] In some embodiments, in step 5, when generating the initial population using the Latin hypercube sampling method, each parameter to be calibrated is uniformly sampled in layers within its corresponding value range to improve the coverage of the parameter space by the initial population; the encoding method is real number encoding or binary encoding.
[0039] In some embodiments, in step 8, the historical sample set includes at least the parameter combinations of individuals in each generation of the population and their corresponding objective function values; the step of using the random forest reverse proxy model to predict candidate optimal parameter combinations specifically involves:
[0040] Using the random forest reverse proxy model, candidate samples are predicted based on the objective function value. That is, the parameter combination corresponding to the theoretical maximum value of the predicted objective function is selected as the candidate optimal parameter combination.
[0041] In some embodiments, in step 10, the termination condition includes at least one of the following conditions: reaching the maximum number of iterations, the improvement of the optimal value of the objective function within a consecutive preset number of algebras being less than a preset threshold, and the objective function value reaching a preset target threshold.
[0042] In some embodiments, the genetic operation includes at least two of selection, crossover, and mutation.
[0043] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned method for estimating parameters of a mathematical model of a wastewater treatment process based on a random forest-guided genetic algorithm.
[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0045] 1. Improve parameter optimization efficiency: This invention introduces a random forest reverse proxy model in the genetic algorithm iteration process. It uses historical real calculation samples to learn the mapping relationship between parameter combinations and objective function values. It can quickly screen out more promising high-fitness candidate parameter combinations from a large number of candidate parameter combinations, thereby reducing blind search and improving parameter optimization efficiency.
[0046] 2. Reduce computational costs of real models: Mathematical models for wastewater treatment processes are often complex and time-consuming for each simulation. This invention utilizes a random forest reverse proxy model to pre-screen candidate parameter combinations, performing real model calculations only on a subset of high-fitness candidate parameter combinations. This reduces the number of real model calls while maintaining optimization effectiveness, thereby lowering computational costs.
[0047] 3. Balancing Search Efficiency and Result Reliability: This invention does not directly replace the results of the real mechanism model with the results of the surrogate model. Instead, it uses a random forest reverse surrogate model as a guiding tool to predict and screen candidate parameter combinations. Subsequently, it still performs real calculations using the mathematical model of the wastewater treatment process, and uses the real objective function value as the basis for population updates and parameter quality judgment. This improves search efficiency while avoiding optimal solution shifts caused by surrogate model errors, thus enhancing the reliability of parameter estimation results.
[0048] 4. Enhanced global search capability: Genetic algorithms have strong global search capabilities. This invention combines the random forest reverse surrogate model with genetic algorithms, which can enhance the search guidance for high-quality solution regions while maintaining global search capabilities, and helps to improve the parameter estimation effect under complex nonlinear multi-parameter coupling problems.
[0049] 5. Solving the problem of initial sample bias: This invention uses Latin hypercube sampling to generate the initial population, which can improve the coverage uniformity of the initial parameter samples in the high-dimensional parameter space and reduce the population bias problem caused by random initialization.
[0050] 6. Strong applicability: This invention can be applied to parameter estimation problems with multiple parameters, strong nonlinearity, and strong coupling in various wastewater treatment processes. While ensuring or improving fitting accuracy, it can effectively shorten parameter calibration time and improve convergence speed, and has good engineering application value. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the embodiments of the present invention are described below. It should be understood that the drawings described below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a flowchart illustrating a method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm, as provided in one embodiment of the present invention.
[0053] Figure 2 for Figure 1 A flowchart illustrating the framework of a parameter estimation method for a mathematical model of wastewater treatment process based on a random forest-guided genetic algorithm.
[0054] Figure 3 for Figure 2 ASM1 biochemical model matrix diagram in the water quality mechanism model;
[0055] Figure 4 for Figure 2 A process flow diagram of multiple fully mixed reactors connected in series in the hydraulic model. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0057] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0058] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0059] Please see Figure 1 One embodiment of the present invention provides a method for estimating parameters of a mathematical model of a wastewater treatment process based on a random forest-guided genetic algorithm, comprising the following steps:
[0060] Obtain historical operating data of the wastewater treatment system, and construct a sample dataset for parameter calibration based on the historical operating data;
[0061] A mathematical model of the wastewater treatment process is established, and the set of parameters to be calibrated for the mathematical model of the wastewater treatment process and their value range are determined.
[0062] Using WNSE, which characterizes the fit between the model's calculated values and the measured values, as the objective function, and taking the maximization of the objective function as the parameter estimation objective, the parameters of the genetic algorithm are set, an initial population is generated, and iterative optimization is performed.
[0063] During the iterative optimization process, a random forest reverse proxy model is constructed based on the historical sample set of the genetic algorithm, and the random forest reverse proxy model is used to predict candidate optimal parameter combinations; the random forest reverse proxy model takes the objective function as input and the parameters as output;
[0064] The candidate optimal parameter combination is used to update the population of the genetic algorithm, and the iterative process is repeated until the termination condition is met. The optimal parameter combination is then output as the parameter estimate of the mathematical model of the wastewater treatment process.
[0065] In some embodiments, the specific implementation steps of the parameter estimation method for the mathematical model of wastewater treatment process based on random forest-guided genetic algorithm are as follows:
[0066] Step 1: Obtain historical operating data of the wastewater treatment system. This historical operating data includes at least influent water quality data, operating condition data, and effluent monitoring data. Based on this historical operating data, construct a sample dataset for parameter calibration. In this embodiment, the influent water quality data includes one or more of chemical oxygen demand (COD), ammonia nitrogen, total nitrogen (TNI), total phosphorus (TP), and suspended solids (SSD). The operating condition data includes one or more of flow rate, dissolved oxygen, sludge age, return ratio, aeration intensity, temperature, and pH. The effluent monitoring data includes one or more of COD, ammonia nitrogen, TNI, TNI, and SSD.
[0067] Step 2: Establish a mathematical model for the wastewater treatment process. This mathematical model includes a water quality mechanism model describing the pollutant transformation process and a hydraulic model describing the hydraulic processes of the treatment structures. In this embodiment, the water quality mechanism model is any one or more of the Activated Sludge Models (ASMs) series and / or an Anaerobic Digestion Model (ADM). Depending on the needs, the water quality mechanism model can also be other matrix models with similar structures. The hydraulic model can also employ other common wastewater treatment process reactor models. The hydraulic model uses a series model of multiple completely mixed reactors and / or a plug flow model. Depending on the needs, the hydraulic model can also employ other common wastewater treatment process reactor models.
[0068] Step 3: Determine the set of parameters to be calibrated, and set the value range of each parameter, the population size of the genetic algorithm, the crossover probability, the mutation probability, the maximum number of iterations, and the fitness evaluation index. Then, determine the simulation target based on the measured operational data. In this embodiment, the set of parameters to be calibrated includes one or more of the following: kinetic parameters, chemometric parameters, and operationally relevant parameters from the wastewater treatment mathematical model.
[0069] Step 4: Use WNSE (Weighted Nash-Sutcliffe efficiency coefficient, WNSE) as the objective function to characterize the fit between the model-calculated values and the measured values, and maximize the objective function as the parameter estimation objective. In this embodiment, the objective function adopts a weighted error index, which is constructed based on the difference between the model simulation output value and the measured value, and different pollutant indicators correspond to different weights to reflect the importance of different effluent indicators in parameter calibration. The WNSE is defined as follows:
[0070]
[0071] in:
[0072] n represents the number of simulated targets;
[0073] m represents the number of samples corresponding to the i-th simulated target;
[0074] Let represent the weight of the i-th simulated target, and ≥0 and ;
[0075] This represents the measured value of the i-th simulated target at the j-th sampling time;
[0076] This represents the model-calculated value of the i-th simulated target at the j-th sampling time.
[0077] This represents the average value of the i-th simulated measured value.
[0078] Step 5: Latin Hypercube Sampling (LHS) is used within the value range of each parameter to be estimated to generate a parameter sample set, and this parameter sample set is encoded into the initial population of the genetic algorithm. In this embodiment, when generating the initial population using the Latin Hypercube Sampling method, each parameter to be calibrated is sampled evenly and hierarchically within its corresponding value range to improve the coverage of the parameter space by the initial population; the encoding method is real number encoding or binary encoding.
[0079] Step 6: Calculate the objective function value for each individual in the initial population using the mathematical model of the wastewater treatment process, and determine the individual fitness using the objective function value.
[0080] Step 7: Perform genetic operations on the population based on individual fitness to generate the next generation population. In this embodiment, the genetic operations include at least two of selection, crossover, and mutation.
[0081] Step 8: Based on the historical sample set obtained during the iteration of the genetic algorithm, a random forest reverse surrogate model is constructed. The reverse surrogate model is trained using the output of the genetic algorithm as input and the input of the genetic algorithm as output, and then predicts the candidate optimal parameter combination corresponding to the maximum objective function value. In this embodiment, the historical sample set includes at least the parameter combinations of individuals in each generation of the population and their corresponding objective function values. Specifically, the step of using the random forest reverse surrogate model to screen or predict candidate parameter combinations corresponding to high objective function values involves: using the random forest reverse surrogate model to predict candidate samples based on the objective function value, that is, selecting the parameter combination corresponding to the predicted objective function value being the theoretical maximum value (the theoretical maximum value is 1) as the candidate optimal parameter combination;
[0082] Step 9: Replace the low-fit individuals in the next generation of the genetic algorithm with the candidate optimal parameter combinations generated in Step 8 to update the population.
[0083] Step 10: Repeat steps 7 to 9 until a preset termination condition is met, and output the optimal parameter combination as the calibration result of the wastewater treatment mathematical model. In this embodiment, the termination condition includes any one or more of the following: reaching the maximum number of iterations, the improvement of the optimal value of the objective function within a consecutive preset number of algebras being less than a preset threshold, or the objective function value reaching a preset target threshold.
[0084] In some embodiments, the simulation targets include, but are not limited to, one or more pollutants selected from chemical oxygen demand (COD), ammonia nitrogen, total nitrogen, nitrate nitrogen, nitrite nitrogen, dissolved oxygen, total phosphorus, phosphate, and sludge concentration. The wastewater treatment process includes, but is not limited to, AAO, SBR, oxidation ditch, MBR, or anaerobic-aerobic combined wastewater treatment processes. The parameter estimation results are used for wastewater treatment process model calibration, operational status prediction, control strategy optimization, or digital twin simulation.
[0085] Please see also Figure 2 The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm is as follows, and it is used for calibrating the parameters of a certain wastewater treatment process mathematical model.
[0086] Constructing a mathematical model for wastewater treatment processes:
[0087] Based on the process flow of the wastewater treatment system to be simulated, a mathematical model of the wastewater treatment process is established. The mathematical model of the wastewater treatment process includes two parts: a water quality mechanism model and a hydraulic model.
[0088] Among them, the water quality mechanism model is used to describe the biochemical processes such as organic matter degradation, nitrification, denitrification, phosphorus removal, sludge growth and decay; the hydraulic model is used to describe the hydraulic behaviors such as mixing, plug flow, recirculation and residence time distribution in different treatment units.
[0089] In a preferred embodiment, the water quality mechanism model can be the ASM1 model, such as... Figure 3 As shown. The hydraulic model can be constructed using multiple fully mixed reactors connected in series, such as... Figure 4 As shown.
[0090] Determine the parameters to be estimated and the simulation objective:
[0091] Based on the calibration requirements of the wastewater treatment process model, a set of parameters to be estimated is selected. These parameters may include, but are not limited to: maximum specific growth rate of heterotrophic bacteria, maximum specific growth rate of nitrifying bacteria, decay coefficient, ammonia nitrogen half-saturation constant, dissolved oxygen half-saturation constant, yield coefficient, and related hydraulic parameters.
[0092] Further simulation targets were selected based on actual operational monitoring data from the wastewater treatment plant. These simulation targets may include effluent COD and effluent NH4. + -N, effluent TN, effluent NO3 - -N, reaction tank DO, MLSS, etc., one or more of these.
[0093] Upper and lower boundaries are set for each parameter to be estimated to form a parameter search space. The boundaries can be determined based on empirical literature, model recommended values, historical operating data, or process design data.
[0094] Establish the objective function:
[0095] Based on the measured data and model calculation results of the simulated target, a parameter estimation objective function is established. In this embodiment, WNSE is used as the objective function, and maximizing WNSE is the optimization objective.
[0096] WNSE can be calculated using the following formula:
[0097]
[0098] The meanings of the symbols are the same as those defined above.
[0099] In one example, if the simulation targets include COD and NH4 + -N and TN can be set with corresponding weights to reflect the importance of different indicators in model calibration.
[0100] The initial population was generated using LHS:
[0101] Within the range of all parameters to be estimated, the LHS (Lower-Side Parameters) is used to generate an initial parameter sample set. Since the LHS can obtain a relatively uniform sample distribution in the multidimensional parameter space, it can improve the quality of the initial population.
[0102] The generated parameter sample set is encoded into the initial population for the genetic algorithm. The encoding method can be either real-number encoding or binary encoding. For parameter estimation problems in wastewater treatment models with a large number of continuous parameters, binary encoding is preferred.
[0103] Calculate the initial population fitness:
[0104] The parameter set corresponding to each individual in the initial population is input into the mathematical model of the wastewater treatment process for calculation, and the corresponding simulation results are obtained. The WNSE value of each individual is then calculated based on the measured data. The WNSE value is used as the fitness value or a fitness function is constructed based on the WNSE value.
[0105] Perform genetic operations to generate the next generation population:
[0106] Based on the fitness values of each individual in the current population, genetic operations are performed on the population, including selection, crossover, and mutation. Selection is used to retain individuals with high fitness; crossover is used to combine the parameter characteristics of different superior individuals; and mutation is used to maintain population diversity and prevent the algorithm from converging prematurely.
[0107] The next generation of the population is generated through the above genetic operations.
[0108] Training a random forest reverse proxy model based on historical samples:
[0109] During the iteration process of the genetic algorithm, the combination of individual parameters that have been calculated in the real model and their corresponding WNSE values in each generation of the population are summarized to form a historical sample set.
[0110] Using WNSE values as input features and parameter combinations as output targets, a random forest model is trained to obtain a random forest reverse proxy model between the objective function value and the parameter combinations.
[0111] In this embodiment, the random forest reverse proxy model does not replace the mathematical model of the wastewater treatment process itself, but is used to learn high-fitness regional features from historical samples, thereby providing guidance for subsequent iterations.
[0112] Screening for high-fitness candidate parameter combinations and updating the population:
[0113] Based on the current parameter boundaries, a batch of candidate parameter combinations is generated within the parameter space; or candidate parameter combinations are generated centered on the current best individual and its neighborhood. The random forest reverse surrogate model trained in step 8 is used to predict the candidate parameter combinations, and one or more candidate parameter combinations with higher predicted WNSE values are selected as high-fitness candidate parameter combinations.
[0114] High-fitness candidate parameter combinations are injected into the next-generation population generated in step 6, replacing one or more individuals with the lowest fitness. This allows the next-generation population to retain the global search capability of the genetic algorithm while increasing its ability to target potentially high-fitness regions.
[0115] Iterate through the loop and output the optimal parameters:
[0116] The process continues with genetic operations to generate the next generation of the population, followed by screening for high-fitness candidate parameter combinations and updating the population. When the termination condition is met, the iteration ends, and the parameter combination with the optimal WNSE value is output as the parameter estimation result for the wastewater treatment process mathematical model.
[0117] The termination conditions can be: reaching the maximum number of iterations, the improvement of the optimal WNSE for several consecutive generations being less than a threshold, or the optimal WNSE reaching a set target value.
[0118] In an exemplary experiment, the method of this invention and the traditional genetic algorithm were used to perform 400 calculations each. The average WNSE was 0.77, the minimum WNSE was 0.74, and the maximum WNSE was 0.81. The average computation time decreased from 156s to 135s (a reduction of 13%), the minimum computation time decreased from 48s to 32s (a reduction of 33%), and the maximum computation time decreased from 535s to 438s (a reduction of 18%). The average number of iterations decreased from 748 to 594 (a reduction of 21%), the minimum number of iterations decreased from 229 to 138 (a reduction of 40%), and the maximum number of iterations decreased from 2324 to 1747 (a reduction of 25%). The average WNSE / computation time ratio increased from 0.00575 / s to 0.007077 / s (an improvement of 23%), and the minimum WNSE / computation time ratio increased from 0.001514 / s to 0.001797 / s (an improvement of 19%). The maximum WNSE / computation time increased from 0.01580 / s to 0.02314 / s (an improvement of 46%); the average WNSE / iterations increased from 0.001188 / iteration to 0.001604 / iteration (an improvement of 35%); the minimum WNSE / iterations increased from 0.0003485 / iteration to 0.00045 / iteration (an improvement of 29%); and the maximum WNSE / iterations increased from 0.003334 / iteration to 0.005329 / iteration (an improvement of 60%). This indicates that the method of the present invention can effectively improve the efficiency of parameter estimation for mathematical models of wastewater treatment processes. Furthermore, under the same number of iterations or the same solution accuracy requirements, the method of the present invention can also shorten the solution time and improve the convergence efficiency.
[0119] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0120] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0121] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for estimating parameters of a mathematical model for wastewater treatment processes based on a random forest-guided genetic algorithm, characterized in that, Includes the following steps: Obtain historical operating data of the wastewater treatment system, and construct a sample dataset for parameter calibration based on the historical operating data; A mathematical model of the wastewater treatment process is established, and the set of parameters to be calibrated for the mathematical model of the wastewater treatment process and their value range are determined. Using WNSE, which characterizes the fit between the model's calculated values and the measured values, as the objective function, and taking the maximization of the objective function as the parameter estimation objective, the parameters of the genetic algorithm are set, an initial population is generated, and iterative optimization is performed. During the iterative optimization process, a random forest reverse proxy model is constructed based on the historical sample set of the genetic algorithm, and the random forest reverse proxy model is used to predict candidate optimal parameter combinations; the random forest reverse proxy model takes the objective function as input and the parameters as output; The candidate optimal parameter combination is used to update the population of the genetic algorithm, and the iterative process is repeated until the termination condition is met. The optimal parameter combination is then output as the parameter estimation result of the mathematical model of the wastewater treatment process.
2. The parameter estimation method for the mathematical model of wastewater treatment process based on random forest-guided genetic algorithm according to claim 1, characterized in that, The specific steps are as follows: Step 1: Obtain historical operating data of the wastewater treatment system. The historical operating data includes at least influent water quality data, operating condition data, and effluent monitoring data. Based on the historical operating data, construct a sample dataset for parameter calibration. Step 2: Establish a mathematical model for the wastewater treatment process. The mathematical model for the wastewater treatment process includes a water quality mechanism model for describing the pollutant transformation process and a hydraulic model for describing the hydraulic process of the treatment structure. Step 3: Determine the set of parameters to be calibrated, and set the value range of each parameter, the population size of the genetic algorithm, the crossover probability, the mutation probability, the maximum number of iterations, and the fitness evaluation index. Then, determine the simulation target based on the actual test data. Step 4: Use WNSE as the objective function and estimate the objective using the maximization of the objective function as a parameter; Step 5: Latin hypercube sampling is used within the range of values of each parameter to be estimated to generate a parameter sample set, and the parameter sample set is encoded as the initial population of the genetic algorithm. Step 6: Calculate the objective function value for each individual in the initial population using the mathematical model of the wastewater treatment process, and determine the individual fitness using the objective function value; Step 7: Perform genetic operations on the population based on individual fitness to generate the next generation population; Step 8: Based on the historical sample set obtained during the iteration of the genetic algorithm, construct a random forest reverse proxy model, train it, and predict the candidate optimal parameter combination corresponding to the maximum objective function value. Step 9: Replace the low-fit individuals in the next generation of the genetic algorithm with the candidate optimal parameter combinations generated in Step 8 to update the population; Step 10: Repeat steps 7 to 9 until the preset termination condition is met, and output the optimal parameter combination as the parameter estimation result of the mathematical model of the wastewater treatment process.
3. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 1, the influent water quality data includes one or more of chemical oxygen demand (COD), ammonia nitrogen, total nitrogen, total phosphorus, and suspended solids; the operating condition data includes one or more of flow rate, dissolved oxygen, sludge age, return ratio, aeration intensity, temperature, and pH; and the effluent monitoring data includes one or more of chemical oxygen demand (COD), ammonia nitrogen, total nitrogen, total phosphorus, and suspended solids.
4. The parameter estimation method for the mathematical model of wastewater treatment process based on random forest-guided genetic algorithm according to claim 2, characterized in that, In step 2, the water quality mechanism model includes at least one of the activated sludge model and the anaerobic digestion model; the hydraulic model adopts a series model of multiple completely mixed reactors and / or a plug flow model.
5. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 3, the set of parameters to be calibrated includes one or more of the following: kinetic parameters, stoichiometric parameters, and operation-related parameters from the wastewater treatment mathematical model.
6. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 4, the WNSE is defined as: in: n represents the number of simulated targets; m represents the number of samples corresponding to the i-th simulated target; Let represent the weight of the i-th simulated target, and ≥0 and ; This represents the measured value of the i-th simulated target at the j-th sampling time; This represents the model-calculated value of the i-th simulated target at the j-th sampling time. This represents the average value of the i-th simulated measured value.
7. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 5, when generating the initial population using the Latin hypercube sampling method, each parameter to be calibrated is sampled evenly in layers within its corresponding value range to improve the coverage of the parameter space by the initial population; the encoding method is real number encoding or binary encoding.
8. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 8, the historical sample set includes at least the parameter combinations of individuals in each generation of the population and their corresponding objective function values; the step of using the random forest reverse surrogate model to predict candidate optimal parameter combinations specifically involves: Using the random forest reverse proxy model, candidate samples are predicted based on the objective function value. That is, the parameter combination corresponding to the theoretical maximum value of the predicted objective function is selected as the candidate optimal parameter combination.
9. The method for estimating parameters of a wastewater treatment process mathematical model based on a random forest-guided genetic algorithm according to claim 2, characterized in that, In step 10, the termination condition includes at least one of the following conditions: reaching the maximum number of iterations, the improvement of the optimal value of the objective function within a consecutive preset number of algebras being less than a preset threshold, and the objective function value reaching a preset target threshold.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the parameter estimation method for the mathematical model of wastewater treatment process based on random forest-guided genetic algorithm as described in any one of claims 1 to 9.