A two-stage activated sludge system optimization method and system based on microbial community prediction
By constructing a microbial community prediction model and a pollutant removal efficiency prediction model, and integrating them with intelligent algorithms, the problem of regulating the complexity and dynamism of the microbial community structure in the activated sludge system was solved. This achieved precise system regulation and improved pollutant removal rate, and features intelligent, automated, and sustainable development characteristics.
Patent Information
- Application Number
- CN202411928483.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The complexity and dynamism of the microbial community structure in existing activated sludge technology make it difficult to achieve effective quantitative control, resulting in unstable treatment efficiency and difficulty in meeting environmental protection requirements.
A two-stage activated sludge system optimization method based on microbial community prediction is constructed, including a microbial community succession prediction model and a pollutant removal efficiency prediction model, which are integrated with intelligent algorithms to achieve precise control of the system by adjusting the control indicators.
It improves pollutant removal rate, reduces treatment costs, realizes intelligent and automated control of the system, enhances operational stability and reliability, has good scalability and adaptability, and meets the requirements of energy conservation, emission reduction and sustainable development.
Smart Images

Figure CN119722415B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of water treatment technology and relates to a two-stage activated sludge system optimization method and system based on microbial community prediction. Background Technology
[0002] Activated sludge technology, a widely used wastewater treatment method, relies on the flocculent sludge formed by microorganisms under aerobic conditions to effectively remove pollutants such as organic matter, nitrogen, and phosphorus from wastewater. Due to its high efficiency, economy, and environmental friendliness, it is widely used globally for the treatment of municipal and industrial wastewater.
[0003] Despite the numerous advantages of activated sludge technology, its stability and efficiency face challenges. This is primarily due to the highly complex microbial community structure within the activated sludge system, with intricate interactions between different microbial species and quantities that are difficult to control effectively using traditional empirical methods. Consequently, maintaining stable and efficient treatment performance is often challenging in practice, leading to significant fluctuations in treatment results and making it difficult to meet increasingly stringent environmental protection requirements.
[0004] To address this issue, researchers have been exploring more precise and effective ways to regulate the microbial community structure in activated sludge technology. However, the complexity and dynamism of the microbial community structure, coupled with the current lack of effective quantitative regulation methods, make solving this problem particularly difficult. Summary of the Invention
[0005] The purpose of this invention is to address the problem in the prior art of lacking effective quantitative control methods for the complexity and dynamism of microbial community structure, and to provide a two-stage activated sludge system optimization method and system based on microbial community prediction.
[0006] To achieve the above objectives, the present invention employs the following technical solution:
[0007] This invention proposes a two-stage activated sludge system optimization method based on microbial community prediction, comprising the following steps:
[0008] The first-stage microbial community succession prediction model was constructed to obtain data on the microbial community in activated sludge. Combined with the environmental parameters and influent data of the reactor, the microbial community was decomposed into diversity and abundance. The time factor was incorporated to track the dynamic changes of the microbial community over time. The microbial community state at a specific future moment was predicted through a time series enhancement strategy.
[0009] A second-stage pollutant removal efficiency prediction model for the activated sludge system was constructed. Using the output and control indicators of the first-stage model as inputs, the effluent indicators of the activated sludge were predicted, and the pollutant treatment performance of the activated sludge system was evaluated.
[0010] By combining intelligent algorithms and employing a non-dominated sorting genetic algorithm to integrate the prediction models of the first and second stages, a two-stage intelligent optimization model for the activated sludge system is constructed. The removal rates of chemical oxygen demand, total nitrogen, and total phosphorus are used as objective functions, and the control indicators are used as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system.
[0011] Preferably, the construction of the first-stage microbial community succession prediction model specifically includes the following steps:
[0012] Data preparation and preprocessing: Time series data containing the abundance of the top 15 microbial species were collected and organized. The data were cleaned to remove outliers and missing values and then standardized.
[0013] Multi-task neural network architecture design: A multi-task neural network is designed, in which each task corresponds to the abundance prediction of a microbial species. The network shares hidden layers to capture the potential correlation between the abundance of different species. Each task has an independent output layer for predicting the abundance of the corresponding species.
[0014] Feature selection and input: Select features related to microbial species abundance as input, including time, water inflow index, regulatory index, and the abundance of the top 15 species in the microbial community, and input the features into the input layer of the multi-task neural network.
[0015] Model training and optimization: The multi-task neural network is trained using the prepared dataset. Mean squared error is used as the loss function to measure the difference between the predicted results and the true values. The Adam optimization algorithm is used to update the weights of the network to minimize the loss function. The performance of the model is evaluated by the 5-fold cross-validation method, and hyperparameter tuning is performed.
[0016] Model application and feedback: The trained multi-task neural network is applied to the actual operation of the activated sludge system to make real-time predictions of microbial species abundance. Based on the prediction results, the control indicators are adjusted to optimize the succession process of the microbial community, and feedback data from the actual application is collected to improve and optimize the model.
[0017] Preferably, during the data preparation and preprocessing, historical data is collected and organized, including the composition and abundance of microbial communities at different time points, as well as the corresponding influent indicators and control parameters; the data is cleaned to remove outliers and missing values, and standardized to ensure data consistency and comparability.
[0018] The Min-Max method is used for data normalization, as shown in formula (1):
[0019]
[0020] Where, x min Let x be the minimum value of the variable. max x is the maximum value of the variable. new These are the normalized variables.
[0021] Preferably, the acquisition of data on the microbial community in activated sludge specifically involves:
[0022] Using a photo-sequential batch reactor of the same laboratory scale, the reactor was placed in a shaded position and surrounded by LED lights to provide constant intensity illumination. The illumination intensity was set to 5000±200 lux, and the illumination time was set to 12 hours of light / 12 hours of darkness.
[0023] The test water is introduced from the top of the reactor by a water pump. The liquid level is controlled by a liquid level relay. The drain outlet is located in the middle of the reactor and is controlled by a solenoid valve. The water volume is replaced by 50% in each cycle.
[0024] A sand core aeration head is installed at the bottom of the reactor for aeration, with an aeration rate of 2L / min.
[0025] The reactor was operated at room temperature with a cycle of 6 hours, including water inlet, aeration, sedimentation, drainage and idle. The water inlet and drainage stages each lasted 5 minutes, followed by 30 minutes of anaerobic idle after water inlet, 5 minutes of sedimentation, and 315 minutes of aeration.
[0026] Synthetic wastewater is used as the operating water source for the reactor, with sodium acetate, ammonium sulfate, and potassium dihydrogen phosphate as the carbon, nitrogen, and phosphorus sources, respectively.
[0027] The inoculated sludge was granular bacterial and algal sludge cultured in the previous stage;
[0028] The reactor was designed with R1-R3 controlling the organic loading rate, R4-R6 controlling the carbon-nitrogen ratio, and R7-R9 controlling the organic nitrogen content.
[0029] Influent, effluent, and biomass samples were collected. Influent parameters, including chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus, were measured. Community data of fungi and eukaryotic microorganisms were measured by performing 16S rRNA high-throughput sequencing and 18S rRNA high-throughput sequencing on biomass samples, respectively.
[0030] Preferably, the pollutant removal efficiency prediction model for constructing the second-stage activated sludge system specifically comprises:
[0031] Data preparation includes collecting pollutant removal efficiency data from the activated sludge system, cleaning the data to remove outliers and missing values, and standardizing the data.
[0032] Feature selection involves selecting features related to pollutant removal efficiency from the collected data, including microbial community indicators and regulatory indicators.
[0033] Model training involves training the model using the selected ensemble tree algorithm and the processed dataset, tuning the model's parameters, including the number and depth of trees, and optimizing the model's performance through cross-validation.
[0034] Preferably, the data preparation specifically includes:
[0035] The output data of the first-stage model and the actual control parameter data are integrated to form an initial dataset;
[0036] Data cleaning is performed to remove outliers and missing values to ensure data integrity and accuracy; data standardization or normalization is performed to improve the efficiency and accuracy of model training.
[0037] The output data includes effluent indicators of the activated sludge system, specifically the concentrations of pollutants COD-OUT, TN-OUT, and TP-OUT; the input data includes the number of bacterial microbial populations and their percentage of the total microbial population, the abundance of the top 15 microbial species and the percentage of algal microbial abundance among the top 15 species, and the control indicators specifically organic loading rate, nitrogen loading, and carbon-nitrogen ratio.
[0038] Preferably, the construction of the two-stage activated sludge system intelligent optimization model specifically includes:
[0039] Problem definition and model initialization: The goal is to find the combination of regulatory parameters that maximizes the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus under the current microbial community conditions, and to initialize the parameters of the non-dominated sorting genetic algorithm, including population size, generation number, crossover probability, and mutation probability. At the same time, an initial population representing candidate solutions for the regulatory parameters is randomly generated.
[0040] Fitness assessment involves substituting each individual in the initial population into the pollutant removal efficiency prediction model of the activated sludge system. The individual is defined as a combination of control index parameters. The corresponding chemical oxygen demand, total nitrogen, and total phosphorus removal rates are calculated, and the fitness value of each individual is calculated based on the removal rates.
[0041] Non-dominated ranking and selection involves ranking individuals in the population in a non-dominated manner, dividing individuals into different non-dominated levels based on their fitness values, and selecting parents based on the non-dominated levels and crowding information to generate the next generation of the population.
[0042] Genetic manipulation involves performing crossover and mutation operations on selected parent individuals to generate new offspring individuals. Crossover operation includes randomly selecting two parent individuals and exchanging some of their genes to generate two new offspring individuals. Mutation operation includes randomly selecting a gene in an offspring individual and modifying it to increase population diversity.
[0043] Population update and iteration: newly generated offspring individuals are merged with parent individuals to form a new population. Non-dominated sorting and crowding calculation are performed on the new population. Individuals with high fitness and uniform distribution are selected as the next generation population. The non-dominated sorting and selection and genetic operation steps are repeated until the preset number of generations is reached or the stopping condition is met.
[0044] The optimal solution is selected and output. The individual with the highest fitness and uniform distribution is selected from the final population as the optimal solution, and the combination of regulatory index parameters corresponding to the optimal solution is output as the optimal regulation scheme to meet the removal rates of chemical oxygen demand, total nitrogen and total phosphorus under the current microbial community state.
[0045] Verification and application: The obtained optimal combination of control index parameters is applied to the actual operation of the activated sludge treatment system to verify its effectiveness under actual conditions. Based on the verification results, the model is adjusted and optimized to improve prediction accuracy and practical application effect.
[0046] Preferably, the optimal solution is measured by an equilibrium value, which minimizes the difference between objective function values within the Pareto optimal solution set that satisfies the conditions. Specifically:
[0047]
[0048] Where: Eq is the equilibrium value of the three objective functions in the Pareto optimal solution; COD is the COD removal rate in the Pareto optimal solution; TP is the TP removal rate in the Pareto optimal solution; TN is the TN removal rate in the Pareto optimal solution; and k is the equilibrium value coefficient.
[0049] This invention proposes a two-stage activated sludge system optimization system based on microbial community prediction, comprising:
[0050] The first data processing module is used to construct the first-stage microbial community succession prediction model, obtain data on the microbial community in activated sludge, combine the environmental parameters and influent data of the reactor, decompose the microbial community into diversity and abundance, incorporate time factors to track the dynamic changes of the microbial community over time, and predict the state of the microbial community at a specific future moment through time series enhancement strategies.
[0051] The second data processing module is used to construct a pollutant removal efficiency prediction model for the second-stage activated sludge system. Using the output and control indicators of the first-stage model as input, it predicts the effluent indicators of the activated sludge and evaluates the pollutant treatment performance of the activated sludge system.
[0052] The model optimization module is used to integrate the first-stage and second-stage prediction models with intelligent algorithms and non-dominated sorting genetic algorithms to construct a two-stage intelligent optimization model for the activated sludge system. The model uses the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus as objective functions and the control indicators as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system.
[0053] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of a two-stage activated sludge system optimization method based on microbial community prediction.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] This invention presents a two-stage activated sludge system optimization method based on microbial community prediction. By constructing a microbial community succession prediction model and a pollutant removal efficiency prediction model, this invention enables precise control and optimization of the activated sludge (AS) system. This facilitates flexible adjustment of control parameters based on the microbial community state and pollutant removal requirements during actual operation, thereby improving the system's treatment efficiency and stability. To improve pollutant removal rates, the invention utilizes a non-dominated sorting genetic algorithm (NSGA) to solve for the optimal control parameters, finding a control scheme that maximizes the removal rates of chemical oxygen demand (COD), total nitrogen (TN), and total phosphorus (TP) under the current microbial community state. This not only helps improve pollutant removal efficiency but also reduces treatment costs and increases resource utilization. Furthermore, this invention combines machine learning algorithms and intelligent optimization technology to achieve intelligent and automated control of the AS system. This reduces human intervention and reliance, improving system operating efficiency and reliability. Simultaneously, the intelligent control method helps to promptly detect and respond to anomalies in system operation, ensuring stable system operation. Finally, the prediction model and control method of this invention exhibit good scalability and adaptability. As the operating conditions of the AS system change and the microbial community structure evolves, the predictive performance and control effects can be optimized by updating the dataset and retraining the model. This allows the invention to continuously adapt to different operating environments and treatment needs, maintaining its effectiveness and competitiveness in practical applications. Regarding energy conservation, emission reduction, and sustainable development, by optimizing the operating parameters and control strategies of the AS system, the invention helps reduce energy consumption and waste emissions, achieving energy conservation, emission reduction, and sustainable development goals. This aligns with current global trends in environmental protection and sustainable development, contributing to the innovation and advancement of water treatment technologies. Compared to existing technologies, the invention demonstrates significant beneficial effects in areas such as precise control and optimization of the AS system, improved pollutant removal rates, intelligent and automated control, scalability and adaptability, and energy conservation, emission reduction, and sustainable development, giving it broad application prospects and significant practical value in the field of water treatment.
[0056] This invention proposes a two-stage activated sludge system optimization system based on microbial community prediction. By dividing the system into a first data processing module, a second data processing module, and a model optimization module, the system adjusts control indicators to directionally regulate the succession of the microbial community, thereby achieving the maximum pollutant removal rate of the activated sludge system. The modular approach ensures that each module is independent, facilitating unified management of all modules. Attached Figure Description
[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a flowchart of the two-stage activated sludge system optimization method based on microbial community prediction of the present invention.
[0059] Figure 2 This is a structural diagram of the two-stage activated sludge system optimization control model of the present invention.
[0060] Figure 3 This is a structural diagram of the optimal multi-task neural network for the first-stage microbial community succession prediction model of the present invention.
[0061] Figure 4 This is a system diagram of the two-stage activated sludge system optimization based on microbial community prediction according to the present invention.
[0062] Figure 5 This is a schematic diagram of the structure of an electronic device according to the present invention. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0064] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0065] The present invention will now be described in further detail with reference to the accompanying drawings:
[0066] Example 1
[0067] This invention proposes a two-stage activated sludge system optimization method based on microbial community prediction, such as... Figure 1 The flowchart shown illustrates the construction process of the two-stage activated sludge system optimization and control model, which includes the following steps:
[0068] S1. Construct the first-stage microbial community succession prediction model to obtain data on the microbial community in activated sludge. Combine the environmental parameters and influent data of the reactor, decompose the microbial community into diversity and abundance, incorporate the time factor to track the dynamic changes of the microbial community over time, and predict the state of the microbial community at a specific future moment through a time series enhancement strategy.
[0069] The construction of the first-stage microbial community succession prediction model includes the following steps:
[0070] S1.1, Data preparation and preprocessing: Collect and organize time series data containing the abundance of the top 15 microbial species, clean the data to remove outliers and missing values, and perform standardization.
[0071] During data preparation and preprocessing, historical data were collected and organized, including the composition and abundance of microbial communities at different time points, as well as the corresponding influent indicators and control parameters. The data were cleaned to remove outliers and missing values, and standardized to ensure data consistency and comparability.
[0072] The Min-Max method is used for data normalization, as shown in formula (1):
[0073]
[0074] Where, x min Let x be the minimum value of the variable. max x is the maximum value of the variable. new These are the normalized variables.
[0075] S1.2, Multi-task neural network structure design: Design a multi-task neural network in which each task corresponds to the abundance prediction of a microbial species. The network shares hidden layers to capture the potential correlation between the abundance of different species. Each task has an independent output layer for predicting the abundance of the corresponding species.
[0076] S1.3 Feature selection and input: Select features related to microbial species abundance as input, including time, water inflow index, regulatory index, and the abundance of the top 15 species in the microbial community. Input the features into the input layer of the multi-task neural network.
[0077] S1.4, Model Training and Optimization: Train the multi-task neural network using the prepared dataset, use mean squared error as the loss function to measure the difference between the predicted results and the true values, use the Adam optimization algorithm to update the network weights to minimize the loss function, evaluate the model performance using the 5-fold cross-validation method, and perform hyperparameter tuning.
[0078] S1.5, Model Application and Feedback: The trained multi-task neural network is applied to the actual operation of the activated sludge system to make real-time predictions of microbial species abundance. Based on the prediction results, the control indicators are adjusted to optimize the succession process of the microbial community, and feedback data from the actual application is collected to improve and optimize the model.
[0079] When obtaining data on the microbial community in activated sludge, the specific steps are as follows:
[0080] A laboratory-scale sequential batch reactor was used, each with a height of 1100 mm, an inner diameter of 70 mm, a height-to-diameter ratio of 15.7, and an effective volume of 3 L. The reactor was placed in a shaded location, and LED lights were installed around it to provide constant intensity illumination. The illumination intensity was set to 5000 ± 200 lux, and the illumination time was set to 12 hours of light / 12 hours of darkness.
[0081] The test water is introduced from the top of the reactor by a water pump. The liquid level is controlled by a liquid level relay. The drain outlet is located in the middle of the reactor and is controlled by a solenoid valve. The water volume is replaced by 50% in each cycle.
[0082] A sand core aeration head is installed at the bottom of the reactor for aeration, with an aeration rate of 2L / min.
[0083] The reactor was operated at room temperature with a cycle of 6 hours, including water inlet, aeration, sedimentation, drainage and idle. The water inlet and drainage stages each lasted 5 minutes, followed by 30 minutes of anaerobic idle after water inlet, 5 minutes of sedimentation, and 315 minutes of aeration.
[0084] Synthetic wastewater is used as the operating water source for the reactor, with sodium acetate, ammonium sulfate, and potassium dihydrogen phosphate as the carbon, nitrogen, and phosphorus sources, respectively.
[0085] The inoculated sludge was granular bacterial and algal sludge cultured in the previous stage;
[0086] The reactor was designed with R1-R3 controlling the organic loading rate, R4-R6 controlling the carbon-nitrogen ratio, and R7-R9 controlling the organic nitrogen content.
[0087] Influent, effluent, and biomass samples were collected. Influent parameters, including chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus, were measured. Community data of fungi and eukaryotic microorganisms were measured by performing 16S rRNA high-throughput sequencing and 18S rRNA high-throughput sequencing on biomass samples, respectively.
[0088] High-throughput sequencing methods include: extracting total genomic DNA from sludge samples using the Ezup column-based soil DNA extraction kit, detecting the concentration and purity of the DNA, storing it at -20°C after it meets the standards, and then performing sequencing on the MiSeq platform. The primers for bacterial sequencing are 338F and 806R, and the primers for algal sequencing are 3NDF and V4-euk-R2R.
[0089] S2. Construct a pollutant removal efficiency prediction model for the second-stage activated sludge system. Using the output and control indicators of the first-stage model as inputs, predict the effluent indicators of the activated sludge and evaluate the pollutant treatment performance of the activated sludge system.
[0090] The second-stage pollutant removal efficiency prediction model for the activated sludge system is constructed as follows:
[0091] S2.1, Data preparation, including collecting pollutant removal efficiency data of the activated sludge system, cleaning the data to remove outliers and missing values, and standardizing the data.
[0092] The data preparation specifically includes:
[0093] The output data of the first-stage model and the actual control parameter data are integrated to form an initial dataset;
[0094] Data cleaning is performed to remove outliers and missing values to ensure data integrity and accuracy; data standardization or normalization is performed to improve the efficiency and accuracy of model training.
[0095] The output data includes effluent indicators of the algal granular sludge (ABGS) system, specifically the concentrations of pollutants COD-OUT, TN-OUT, and TP-OUT; the input data includes the number of bacterial microbial populations and their percentage of the total microbial population, the abundance of the top 15 microbial species and the percentage of algal microbial abundance among the top 15 species, and the control indicators specifically organic loading rate, nitrogen loading, and carbon-nitrogen ratio.
[0096] It should be noted that, in order to reflect the application scope of the present invention as much as possible, this example applies an activated sludge improvement technology with a more complex microbial community than the activated sludge system: the "ABGS system".
[0097] S2.2 Feature selection: Selecting features related to pollutant removal efficiency from the collected data, including microbial community indicators and regulatory indicators.
[0098] S2.3, Model Training: Train the model using the selected ensemble tree algorithm and the processed dataset, adjust the model parameters, including the number and depth of trees, and optimize the model's performance through cross-validation.
[0099] S3, combined with intelligent algorithms, integrates the first and second stage prediction models using a non-dominated sorting genetic algorithm to construct a complete two-stage intelligent optimization model for activated sludge systems. The model uses the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus as objective functions and the control indicators as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system.
[0100] A two-stage intelligent optimization model for activated sludge systems is constructed, specifically as follows:
[0101] S3.1 Problem Definition and Model Initialization: The goal is to find the combination of regulatory parameters that maximizes the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus under the current microbial community conditions, and to initialize the parameters of the non-dominated sorting genetic algorithm, including population size, generation number, crossover probability, and mutation probability. At the same time, an initial population representing candidate solutions for the regulatory parameters is randomly generated.
[0102] S3.2, Fitness Assessment: Substitute each individual in the initial population into the pollutant removal efficiency prediction model of the activated sludge system. The individual is the combination of control index parameters. Calculate the corresponding chemical oxygen demand, total nitrogen, and total phosphorus removal rates, and calculate the fitness value of each individual based on the removal rates.
[0103] S3.3, Non-dominated sorting and selection: Individuals in the population are sorted non-dominatedly, and individuals are divided into different non-dominated levels according to their fitness values. Parents are selected based on the non-dominated levels and crowding information to generate the next generation of the population.
[0104] S3.4, Genetic operations, performing crossover and mutation operations on selected parent individuals to generate new offspring individuals. The crossover operation includes randomly selecting two parent individuals and exchanging some of their genes to generate two new offspring individuals. The mutation operation includes randomly selecting a gene in an offspring individual and modifying it to increase population diversity.
[0105] S3.5, Population Update and Iteration: The newly generated offspring individuals are merged with the parent individuals to form a new population. Non-dominated sorting and crowding calculation are performed on the new population. Individuals with high fitness and uniform distribution are selected as the next generation population. The non-dominated sorting and selection and genetic operation steps are repeated until the preset number of generations is reached or the stopping condition is met.
[0106] S3.6, Optimal Solution Selection and Output: Select the individual with the highest fitness and uniform distribution from the final population as the optimal solution, and output the corresponding combination of regulatory index parameters as the optimal regulation scheme to meet the removal rates of chemical oxygen demand, total nitrogen and total phosphorus under the current microbial community state.
[0107] The optimal solution is measured by the equilibrium value. Within the set of Pareto optimal solutions that satisfy the conditions, the objective function values are minimized. Specifically:
[0108]
[0109] Where: Eq is the equilibrium value of the three objective functions in the Pareto optimal solution; COD is the COD removal rate in the Pareto optimal solution; TP is the TP removal rate in the Pareto optimal solution; TN is the TN removal rate in the Pareto optimal solution; and k is the equilibrium value coefficient.
[0110] S3.7, Verification and Application: The obtained optimal combination of control index parameters is applied to the actual operation of the activated sludge treatment system to verify its effectiveness under actual conditions. Based on the verification results, the model is adjusted and optimized to improve prediction accuracy and practical application effect.
[0111] Example
[0112] A model structure framework diagram of one embodiment of the present invention is shown below. Figure 2 As shown, the specific steps are as follows:
[0113] Step 1, Construction of the first-stage microbial community succession prediction model:
[0114] Data on the microbial community in activated sludge was acquired and combined with reactor environmental parameters and influent data. The model comprehensively decomposes the microbial community into two main aspects: diversity and abundance, and cleverly incorporates the time factor to accurately track the dynamic changes of the microbial community over time. By applying a time-series enhancement strategy, the model can accurately predict the state of the microbial community at a specific future moment based on initial conditions, time span, influent indicators, and control parameters.
[0115] The microbial community succession prediction model is built using support vector machine and multi-task neural network algorithms.
[0116] Microbial community diversity characteristics include: the number and proportion of fungal microbial populations, and microbial community abundance characteristics include: the abundance of the top 15 microbial species and the proportion of algal microbial abundance, among other key microbial indicators.
[0117] The output indicators of the first-stage model for predicting microbial community succession include: microbial community diversity and abundance characteristics. The input indicators include: past time data of the output indicators, the time interval from the past time point to the present, the influent indicators of the activated sludge system, and the indicators that can be controlled.
[0118] To demonstrate the effectiveness of this invention, the activated sludge process provided is an example of algal granular sludge (ABGS) with a more complex microbial community. First, microbial community data and other environmental indicator data related to the model established in this invention should be collected.
[0119] This includes obtaining data on the microbial community in activated sludge, combined with reactor environmental parameters and influent data, including:
[0120] For data acquisition, this example uses nine identical laboratory-scale photo-sequencing batch reactors (PSBRs).
[0121] Specifically, the photo-sequencing batch reactor (PSBR) has a height of 1100 mm, an inner diameter of 70 mm, a height-to-diameter ratio (H / D) of 15.7, and an effective volume of 3 L. During operation, the reactor is placed in a shaded location with LED lights providing constant intensity illumination around it. The test feed water enters from the top of the reactor via a pump, with the liquid level controlled by a level relay. The drain outlet is located in the middle of the reactor and controlled by a solenoid valve. Each cycle involves a 50% water replacement volume, with a replacement rate of 1.5 L. Aeration is achieved at the bottom of the reactor using sand core aerators, with the aeration rate controlled at 2 L / min by a glass rotor flow meter. The PSBR reactor operates at room temperature with a 6-hour cycle. The operating time for each stage is adjusted by a timer, including five stages: feed water, aeration, sedimentation, drainage, and idle. The feed water and drainage stages are each 5 minutes, followed by a 30-minute anaerobic idle after feed water, a 5-minute sedimentation time, and a 315-minute aeration time. LED lights were used as the light source and placed at the rear of the reactor. The light intensity was set to 5000 ± 200 lux near the inner wall of the reactor, and the illumination time was set to 12 hours of light and 12 hours of darkness. The reactor was operated using synthetic wastewater.
[0122] The carbon, nitrogen, and phosphorus sources were sodium acetate, ammonium sulfate, and potassium dihydrogen phosphate, respectively. The inoculum sludge was ABGS cultured previously.
[0123] To investigate the effects of different operating conditions on ABGS cultivation, this study designed reactors with R1-R3 controlling the organic loading rate (OLR), R4-R6 controlling the carbon-to-nitrogen ratio (C / N), and R7-R9 controlling the organic nitrogen content (ON). These were referred to as control parameters. For the initial inoculation of ABGS, the reactors with the three control parameters were taken from different batches of sludge previously cultivated by our research group. Influent, effluent, and biomass samples were collected daily to measure the following parameters at different times: influent parameters (chemical oxygen demand (COD-IN), ammonia nitrogen (NH3-IN), total nitrogen (TN-IN), and total phosphorus (TP-IN)) and the bacterial and algal microbial community of ABGS (16S rRNA high-throughput and 18S rRNA high-throughput). 16S rRNA high-throughput measured bacterial microorganisms, while 18S rRNA high-throughput measured eukaryotic microorganisms; the 18S rRNA high-throughput results also included algae. Other parameters were measured according to standard methods. To facilitate model construction,
[0124] Specifically, the high-throughput sequencing method was as follows: Total genomic DNA was extracted from sludge samples using the Ezup column-based soil DNA extraction kit from Sangon Biotech (Shanghai) Co., Ltd. The concentration and purity of the DNA were then measured using a Nandrop2000 micro spectrophotometer, and samples meeting the standards were stored at -20°C. Sequencing was performed by Pasenuo Biotechnology Co., Ltd. (Shanghai, China) on the MiSeq platform (Illumina, USA). For bacterial sequencing, the upstream primer was 338F (5'-ACTCCTACGGGAGGCAGCAG-3'), and the downstream primer was 806R (5'-GGACTACH VGGGTWTCTAAT-3'). For algal sequencing, the upstream primer was 3NDF (5'-GGCAAGTCTGGTGCCAG-3'), and the downstream primer was V4-euk-R2R (5'-ACGGTATCTRATCRTCTTCG-3').
[0125] The pollutant removal efficiency indicators for ABGS were set as the removal of carbon (C), nitrogen (N), and phosphorus (P), expressed as the removal rates of COD, TN, and TP (COD-OUT, TN-OUT, TP-OUT (%)), respectively. For ease of interpretation, these are referred to as effluent indicators. To facilitate model visualization, some indicators were redefined and described, and their detailed explanations are shown in Table 1. The range of all collected data is shown in Table 2.
[0126] Table 1 describes the characteristics present in the dataset used for modeling.
[0127]
[0128]
[0129] Table 2. Reactor index range
[0130]
[0131] Specifically, in step 1, based on initial conditions, time span, water inflow indicators, and control parameters, this invention provides a prediction method based on a multi-task neural network algorithm to accurately predict the state of the microbial community at a specific future moment, specifically including:
[0132] S11, Data Preparation and Preprocessing
[0133] Specifically, historical data is collected and organized, including the composition and abundance of microbial communities at different time points, as well as the corresponding influent indicators and control parameters.
[0134] Specifically, the data is cleaned to remove outliers and missing values, and standardized to ensure data consistency and comparability.
[0135] It should be noted that, since each influencing factor has a different dimension, each dimension is normalized. In this paper, the Min-Max method is used for data normalization, as shown in formula (1).
[0136] S12, Construction of Multi-Task Neural Network Model
[0137] Specifically, a multi-task neural network structure is designed, where each task corresponds to the abundance prediction of one or more microbial species, or the prediction of dynamic changes in different functional microbial communities.
[0138] Specifically, the network shares an underlying hidden layer to capture potential correlations between different tasks, while each task has an independent output layer to predict the future state of the corresponding microbial community.
[0139] Specifically, the output indicators include: the number of fungal microbial populations (16S) and its percentage of the total microbial population (PCT of 16S), the abundance of the top 15 microbial species (Ab), and the percentage of algal microbial abundance among the top 15 species (PCT of 18S). The input variables include the state and time of the output indicators at the previous time point, influent indicators (chemical oxygen demand (COD-IN), ammonia nitrogen (NH3-IN), total nitrogen (TN-IN), total phosphorus (TP-IN)) and regulation indicators (OLR, ON, C / N).
[0140] S13, Model Training and Optimization
[0141] Specifically, a multi-task neural network model is trained using a prepared dataset, and the network weights are updated using the backpropagation algorithm and the optimizer (Adam).
[0142] Specifically, 5-fold cross-validation and L2 regularization are used to prevent overfitting, while Bayesian optimization is applied to adjust hyperparameters such as network structure (number of layers, number of neurons) to optimize model performance.
[0143] The final optimized hyperparameters are: fully connected layers: 2, first layer size: 15, second layer size: 17, activation function: none, regularization strength (Lambda): 5.96*10^-8, and normalized data: yes. For the specific model structure of the multi-task neural network, see [link to model]. Figure 3 .
[0144] S15, Prediction and Verification
[0145] Specifically, the trained model is applied to new datasets or the actual ABGS system to predict the future state of the microbial community. The accuracy and reliability of the model's predictions are verified by comparing them with actual observational data.
[0146] S16, Results Analysis and Application
[0147] Specifically, the analysis and prediction results identify the dynamic trends of key microbial species or functional communities. Model performance is evaluated using the coefficient of determination (R²) and root mean square error (RMSE).
[0148] The performance of the microbial community succession prediction model in the first stage was: RMSR = 102.004, R2 = 0.951.
[0149] In summary, this invention, through a multi-task neural network algorithm combined with initial conditions, time span, influent indicators, and control parameters, achieves accurate prediction of microbial community succession in ABGS systems, providing strong support for the intelligent control and optimization of the system.
[0150] Step 2, the second-stage pollutant removal efficiency prediction model of the activated sludge system, includes:
[0151] This model uses the output and control indicators of the first-stage model as inputs to directly predict the effluent indicators of activated sludge, thereby comprehensively evaluating the pollutant treatment performance of the activated sludge system.
[0152] The pollutant removal efficiency prediction model of the activated sludge system is built using an ensemble tree algorithm.
[0153] This patent uses the removal rates of chemical oxygen demand (COD), total nitrogen (TN), and total phosphorus (TP) to characterize the pollutant treatment performance of activated sludge and as output indicators of the second-stage model.
[0154] In constructing an effluent indicator prediction model for an activated sludge system (specifically, an algal granular sludge ABGS system), this invention uses the output indicators (such as microbial community diversity and abundance changes) and control indicators (OLR, C / N, ON) of the first-stage model as input features. The model is trained using an ensemble tree (TE) algorithm to accurately predict effluent indicators of the ABGS system, such as the concentrations of COD, TN, and TP pollutants. This step specifically includes:
[0155] S21, Data Integration and Preprocessing
[0156] Specifically, the output data of the first-stage model and the actual control parameter data are integrated to form a training dataset.
[0157] Specifically, data cleaning involves removing outliers and missing values to ensure data integrity and accuracy. Data standardization or normalization is also performed to improve the efficiency and accuracy of model training.
[0158] Specifically, the output indicators include the concentrations of pollutants in the effluent from the ABGS system (COD-OUT, TN-OUT, TP-OUT). Input variables include: the abundance of bacterial microbial populations (16S) and their percentage of the total microbial population (PCT of 16S), the abundance of the top 15 microbial species (Ab), the percentage of algal microbial abundance among the top 15 species (PCT of 18S), and regulatory indicators (OLR, ON, C / N).
[0159] S22, Selection and Configuration of Ensemble Tree Algorithm
[0160] Choose an appropriate ensemble tree algorithm based on the complexity of the problem and the characteristics of the data, such as random forest, gradient boosting tree (GBDT), extreme gradient boosting (XGBoost), etc.
[0161] Specifically, configure algorithm parameters, such as the number of trees, tree depth, and learning rate, to optimize the model's predictive performance.
[0162] For example, the final optimized hyperparameters of the ABGS system effluent pollutant concentration prediction model (COD-OUT, TN-OUT, TP-OUT) are shown in Table 3.
[0163] Table 3 shows the Bayesian-optimized hyperparameters of the optimal machine learning algorithms for the seven models.
[0164]
[0165] S23, Model Training
[0166] Specifically, the ensemble tree model is trained using the preprocessed dataset, and the model parameters are updated iteratively to make the model gradually approximate the actual changes in water discharge indicators.
[0167] Specifically, during training, 5-fold cross-validation and regularization are used to prevent overfitting and improve the model's generalization ability.
[0168] S24, Model Validation and Optimization
[0169] Specifically, the predictive performance of the model is evaluated using an independent validation dataset, and metrics such as the coefficient of determination (R2) and root mean square error (RMSE) are calculated.
[0170] Specifically, the COD-OUT prediction model is: RMSE = 0.314, R 2 =0.980, the TN-OUT prediction model is: RMSE = 1.288, R 2 =0.944, while the TP-OUT prediction model has: RMSE = 0.200, R 2 =0.983.
[0171] S25, Model Application and Feedback
[0172] Specifically, the trained ensemble tree model is applied to the actual operation of the ABGS system to predict effluent indicators in real time.
[0173] In summary, this invention uses the output and control indicators of the first-stage model as inputs and employs the ensemble tree algorithm to train a model that predicts the effluent indicators of activated sludge, thereby achieving accurate prediction of the effluent quality of the ABGS system and providing strong support for the intelligent operation and performance optimization of the system.
[0174] Step 3, Construction of the intelligent optimization model for the two-stage activated sludge system:
[0175] Building upon steps 1 and 2, and leveraging the advantages of intelligent algorithms in finding optimal solutions, this invention employs a non-dominated sorting genetic algorithm (NSGA) to integrate the two-stage prediction model, constructing a complete two-stage microbial community optimization and control model. This model uses the removal rates of COD, TN, and TP as objective functions, with regulatory indicators as the solution space. By finely adjusting these indicators, it aims to directionally regulate the succession of the microbial community, striving to achieve the maximum pollutant removal rate of the activated sludge system.
[0176] For example, to demonstrate the final function of the present invention, four final pollutant removal requirements were selected: balanced and efficient removal of carbon, nitrogen and phosphorus, efficient removal of carbon and nitrogen, efficient removal of nitrogen and phosphorus, and efficient removal of carbon and phosphorus.
[0177] When constructing the optimal control model for a two-stage activated sludge system (specifically referring to the ABGS system of bacterial and algal granular sludge), the removal rates of COD, TN, and TP are used as objective functions, and the control indices (OLR, C / N, ON) are used as the solution space. A non-dominated sorting genetic algorithm (NSGA) is employed to integrate and optimize the two-stage prediction model, thereby establishing an optimal control model for a two-stage activated sludge system that comprehensively considers microbial community succession and pollutant removal efficiency. This step specifically includes:
[0178] S31, Definition of the objective function
[0179] Specifically, the objective function should be clearly defined, namely, maximizing the removal rates of COD, TN, and TP to reflect the processing efficiency of the ABGS system.
[0180] Specifically, the objective function is the pollutant removal performance prediction model of the activated sludge in the second stage of step 2.
[0181] S32, Solving the spatial boundary
[0182] Specifically, the range of values for the control indicators is determined, forming the solution space. The NSGA algorithm is chosen as the solution method because it can effectively handle multi-objective optimization problems and find a set of non-dominated solutions (i.e., Pareto optimal solutions). Algorithm parameters, such as population size, number of generations, crossover probability, and mutation probability, are configured to optimize the algorithm's search efficiency and performance.
[0183] For example, the solution space of this invention is the upper and lower bounds of the control index (OLR, C / N, ON), as shown in Table 2.
[0184] S33 Two-Stage Prediction Model Integration
[0185] Specifically, the first-stage microbial community succession prediction model and the second-stage effluent index prediction model are integrated to form a complete two-stage prediction model.
[0186] In the NSGA algorithm, the predicted model of the removal rate of the three pollutants in the second-stage activated sludge prediction model is used as the output as the basis for fitness assessment.
[0187] Given that the ABGS system must meet three objectives in terms of pollutant removal efficiency, the optimal solution obtained by the NSGA algorithm is not a single solution, but rather a set of Pareto optimal solutions (PS) that satisfy the conditions. To evaluate the quality of each solution within the PS, this study considers the stringent requirement in practical engineering applications that the removal rates of all three pollutants must meet the standards. Specifically, in the PS, the smaller the difference between the three objective function values, the better it reflects the characteristic of simultaneously achieving the optimal solution. Therefore, this study uses the degree of difference between the three objective variable values in the PS to measure the quality of the solution, and defines it as the equilibrium value (Eq), as detailed in formula (2).
[0188] Specifically, where: Eq is the equilibrium value of the three objective functions in the Pareto optimal solution; COD is the removal rate of COD in the Pareto optimal solution; TP is the removal rate of TP in the Pareto optimal solution; TN is the removal rate of TN in the Pareto optimal solution; and k is the equilibrium value coefficient, which is set to 4 in this study for ease of analysis of Eq.
[0189] S34, Genetic manipulation and fitness assessment
[0190] Specifically, genetic manipulations, such as crossover and mutation, are performed on individuals in the population (i.e., combinations of regulatory indicators) to generate new offspring individuals.
[0191] Specifically, a two-stage prediction model is used to assess the fitness of offspring individuals and calculate their corresponding COD, TN, and TP removal rates.
[0192] S35, Non-dominated sorting and selection
[0193] Specifically, individuals in the population are non-dominated and ranked according to their fitness values to form different non-dominated hierarchies.
[0194] Specifically, at each level, individuals with high fitness and uniform distribution are selected as candidate solutions for the next generation based on crowding information.
[0195] S36, Iterative Optimization and Pareto Optimal Solution Set Generation
[0196] Specifically, genetic operations, fitness assessments, and non-dominated sorting are repeated until the preset number of generations is reached or the stopping condition is met.
[0197] The hyperparameters of NSGA in this embodiment are set as follows: selection operation: tournament selection, optimal front-end individual coefficient: 0.25, population size: 1000, maximum number of generations: 1000, stopping generation: 1000, fitness function value deviation: 1e-100.
[0198] Specifically, by setting the parameters above, 250 Pareto optimal solution sets (PS) that meet the requirements can be selected.
[0199] Specifically, the output Pareto optimal solution set is a set of optimal control indicators within a given solution space, which balance the removal rates of COD, TN and TP to varying degrees.
[0200] S37, Application and Feedback of the Optimization Control Model
[0201] Specifically, the combination of control indices from the Pareto optimal solution set is applied to the actual operation of the ABGS system for performance verification and optimization.
[0202] The model of this invention simulated four different control strategies and their final control results for engineering requirements. The specific settings are as follows: For the EQ (Effective Fluidity) project, the objective is to achieve the most balanced removal rates for C, N, and P, with high-efficiency removal rate as a secondary objective. The lowest Eq values selected from 250 PS (Pollutants in the Quantitative Analysis) are: COD-OUT, TN-OUT, TP-OUT, and Eq are 87.16%, 88.11%, 89.97%, and 3.50, respectively. This method can achieve a removal rate of over 85% for each pollutant. For the NC (Negative Fluidity) project, the objective is to achieve the most efficient removal rates for N and P, with C removal rate as a secondary objective. The lowest equilibrium values selected from 250 PS are: COD-OUT, TN-OUT, TP-OUT, and Eq are 81.25%, 90.76%, 94.00%, and 16.23, respectively. This method can achieve a removal rate of over 90% for N and P. The NN project aims to maximize C and P removal efficiency, with N removal being a secondary objective. The lowest equilibrium values selected from 250 PSs are: COD-OUT, TN-OUT, TP-OUT, and Eq are 90.08%, 81.01%, 90.54%, and 13.17, respectively. This method can achieve C and P removal rates exceeding 90%. The NP project also aims to maximize C and N removal efficiency, with P removal being a secondary objective. The lowest Eq values selected from 250 PSs are: COD-OUT, TN-OUT, TP-OUT, and Eq are 90.00%, 90.84%, 74.09%, and 23.12, respectively. This method can achieve C and N removal rates exceeding 90%. The simulated influent treatment strategies in this study show that the NC effluent has a high COD content, approximately 95 mg / L; the NN effluent has a high TN content, approximately 14 mg / L; and the NP effluent has a high TP content, approximately 2.3 mg / L. The specific control results and strategies are shown in Table 4.
[0203] Table 4. Regulated output results of NSGA-3 under specific conditions
[0204]
[0205] Depending on different industrial effluent standards, engineers can select different microbial community regulation strategies to ensure that the pollutant removal efficiency of the ABGS system remains high and stable. This two-stage microbial community optimization model framework transforms the empirical operation of the ABGS system into a scientific operation, providing a reliable and intelligent guidance model for the operation of the ABGS system.
[0206] In summary, this invention establishes a two-stage optimal control model for activated sludge systems by integrating and optimizing the two-stage prediction model using a non-dominated sorting genetic algorithm (NSGA), which comprehensively considers microbial community succession and pollutant removal efficiency, providing strong support for the intelligent operation and performance optimization of ABGS systems.
[0207] Example 2
[0208] This invention proposes a two-stage activated sludge system optimization system based on microbial community prediction, such as... Figure 4 As shown, it includes:
[0209] The first data processing module is used to construct the first-stage microbial community succession prediction model, obtain data on the microbial community in activated sludge, combine the environmental parameters and influent data of the reactor, decompose the microbial community into diversity and abundance, incorporate time factors to track the dynamic changes of the microbial community over time, and predict the state of the microbial community at a specific future moment through time series enhancement strategies.
[0210] The second data processing module is used to construct a pollutant removal efficiency prediction model for the second-stage activated sludge system. Using the output and control indicators of the first-stage model as input, it predicts the effluent indicators of the activated sludge and evaluates the pollutant treatment performance of the activated sludge system.
[0211] The model optimization module is used to integrate the first-stage and second-stage prediction models with intelligent algorithms and non-dominated sorting genetic algorithms to construct a two-stage intelligent optimization model for the activated sludge system. The model uses the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus as objective functions and the control indicators as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system.
[0212] Example 3
[0213] Please see Figure 5As shown, the present invention also provides an electronic device 100 for a two-stage activated sludge system optimization method based on microbial community prediction; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0214] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the two-stage activated sludge system optimization method based on microbial community prediction described in Example 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0215] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.
[0216] The memory 101 in the electronic device 100 stores multiple instructions to implement a two-stage activated sludge system optimization method based on microbial community prediction, and the processor 102 can execute the multiple instructions to achieve the following:
[0217] The first-stage microbial community succession prediction model was constructed to obtain data on the microbial community in activated sludge. Combined with the environmental parameters and influent data of the reactor, the microbial community was decomposed into diversity and abundance. The time factor was incorporated to track the dynamic changes of the microbial community over time. The microbial community state at a specific future moment was predicted through a time series enhancement strategy.
[0218] A second-stage pollutant removal efficiency prediction model for the activated sludge system was constructed. Using the output and control indicators of the first-stage model as inputs, the effluent indicators of the activated sludge were predicted, and the pollutant treatment performance of the activated sludge system was evaluated.
[0219] By combining intelligent algorithms and employing a non-dominated sorting genetic algorithm to integrate the prediction models of the first and second stages, a two-stage intelligent optimization model for the activated sludge system is constructed. The removal rates of chemical oxygen demand, total nitrogen, and total phosphorus are used as objective functions, and the control indicators are used as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system.
[0220] Example 4
[0221] If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0222] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0223] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0224] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0225] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A two-stage activated sludge system optimization method based on microbial community prediction, characterized in that, Includes the following steps: The first-stage microbial community succession prediction model was constructed to obtain data on the microbial community in activated sludge. Combined with the environmental parameters and influent data of the reactor, the microbial community was decomposed into diversity and abundance. The time factor was incorporated to track the dynamic changes of the microbial community over time. The microbial community state at a specific future moment was predicted through a time series enhancement strategy. A second-stage pollutant removal efficiency prediction model for the activated sludge system was constructed. Using the output and control indicators of the first-stage model as inputs, the effluent indicators of the activated sludge were predicted, and the pollutant treatment performance of the activated sludge system was evaluated. By combining intelligent algorithms and employing a non-dominated sorting genetic algorithm to integrate the prediction models of the first and second stages, a two-stage intelligent optimization model for the activated sludge system is constructed. The removal rates of chemical oxygen demand, total nitrogen, and total phosphorus are used as objective functions, and the control indicators are used as the solution space. By adjusting the control indicators, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system. The construction of the two-stage intelligent optimization model for the activated sludge system specifically includes problem definition and model initialization, fitness assessment, non-dominated sorting and selection, genetic operations, population update and iteration, optimal solution selection and output, as well as verification and application. Problem definition and model initialization: The goal is to find the combination of regulatory index parameters that maximizes the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus under the current microbial community state, and to initialize the parameters of the non-dominated sorting genetic algorithm, including population size, number of generations, crossover probability, and mutation probability. At the same time, an initial population representing candidate solutions for the regulatory index parameters is randomly generated. Fitness assessment involves substituting each individual in the initial population into the pollutant removal efficiency prediction model of the activated sludge system. The individual is defined as a combination of control index parameters. The corresponding chemical oxygen demand, total nitrogen, and total phosphorus removal rates are calculated, and the fitness value of each individual is calculated based on the removal rates. The optimal solution is selected and output. The individual with the highest fitness and uniform distribution is selected from the final population as the optimal solution, and the combination of regulatory index parameters corresponding to the optimal solution is output as the optimal regulation scheme to meet the removal rates of chemical oxygen demand, total nitrogen and total phosphorus under the current microbial community state. Verification and application: The obtained optimal combination of control index parameters is applied to the actual operation of the activated sludge treatment system to verify its effectiveness under actual conditions. Based on the verification results, the model is adjusted and optimized to improve prediction accuracy and practical application effect.
2. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 1, characterized in that, The construction of the first-stage microbial community succession prediction model specifically includes the following steps: Data preparation and preprocessing: Time series data containing the abundance of the top 15 microbial species were collected and organized. The data were cleaned to remove outliers and missing values and then standardized. Multi-task neural network architecture design: A multi-task neural network is designed, in which each task corresponds to the abundance prediction of a microbial species. The network shares hidden layers to capture the potential correlation between the abundance of different species. Each task has an independent output layer for predicting the abundance of the corresponding species. Feature selection and input: Select features related to microbial species abundance as input, including time, water inflow index, regulatory index, and the abundance of the top 15 species in the microbial community, and input the features into the input layer of the multi-task neural network. Model training and optimization: The multi-task neural network is trained using the prepared dataset. Mean squared error is used as the loss function to measure the difference between the predicted results and the true values. The Adam optimization algorithm is used to update the weights of the network to minimize the loss function. The performance of the model is evaluated by the 5-fold cross-validation method, and hyperparameter tuning is performed. Model application and feedback: The trained multi-task neural network is applied to the actual operation of the activated sludge system to make real-time predictions of microbial species abundance. Based on the prediction results, the control indicators are adjusted to optimize the succession process of the microbial community, and feedback data from the actual application is collected to improve and optimize the model.
3. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 2, characterized in that, During the data preparation and preprocessing process, historical data is collected and organized, including the composition and abundance of microbial communities at different time points, as well as the corresponding influent indicators and control parameters; the data is cleaned to remove outliers and missing values, and standardized to ensure data consistency and comparability. The Min-Max method is used for data normalization, as shown in formula (1): in, The minimum value of the variable. For the maximum value of the variable, Let x be the normalized variable, and let x be all the measured data indicators.
4. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 1, characterized in that, Specifically, when obtaining data on the microbial community in activated sludge, the following steps are taken: Using a photo-sequential batch reactor of the same laboratory scale, the reactor was placed in a shaded position and surrounded by LED lights to provide constant intensity illumination. The illumination intensity was set to 5000±200 lux, and the illumination time was set to 12 hours of light / 12 hours of darkness. The test water is introduced from the top of the reactor by a water pump. The liquid level is controlled by a liquid level relay. The drain outlet is located in the middle of the reactor and is controlled by a solenoid valve. The water volume is replaced by 50% in each cycle. A sand core aeration head is installed at the bottom of the reactor for aeration, with an aeration rate of 2L / min. The reactor was operated at room temperature with a cycle of 6 hours, including water inlet, aeration, sedimentation, drainage and idle. The water inlet and drainage stages each lasted 5 minutes, followed by 30 minutes of anaerobic idle after water inlet, 5 minutes of sedimentation, and 315 minutes of aeration. Synthetic wastewater is used as the operating water source for the reactor, with sodium acetate, ammonium sulfate, and potassium dihydrogen phosphate as the carbon, nitrogen, and phosphorus sources, respectively. The inoculated sludge was granular bacterial and algal sludge cultured in the previous stage; The reactor was designed with R1-R3 controlling the organic loading rate, R4-R6 controlling the carbon-nitrogen ratio, and R7-R9 controlling the organic nitrogen content. Influent, effluent, and biomass samples were collected. Influent parameters, including chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus, were measured. Community data of fungi and eukaryotic microorganisms were measured by performing 16S rRNA high-throughput sequencing and 18S rRNA high-throughput sequencing on biomass samples, respectively.
5. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 1, characterized in that, The pollutant removal efficiency prediction model for the second-stage activated sludge system is specifically constructed as follows: Data preparation includes collecting pollutant removal efficiency data from the activated sludge system, cleaning the data to remove outliers and missing values, and standardizing the data. Feature selection involves selecting features related to pollutant removal efficiency from the collected data, including microbial community indicators and regulatory indicators. Model training involves training the model using the selected ensemble tree algorithm and the processed dataset, tuning the model's parameters, including the number and depth of trees, and optimizing the model's performance through cross-validation.
6. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 5, characterized in that, The data preparation specifically involves: The output data of the first-stage model and the actual control parameter data are integrated to form an initial dataset; Data cleaning is performed to remove outliers and missing values to ensure data integrity and accuracy; data standardization or normalization is performed to improve the efficiency and accuracy of model training. The output data includes effluent indicators of the activated sludge system, specifically the concentrations of pollutants COD-OUT, TN-OUT, and TP-OUT; the input data includes the number of bacterial microbial populations and their percentage of the total microbial population, the abundance of the top 15 microbial species and the percentage of algal microbial abundance among the top 15 species, as well as the control indicators specifically organic loading rate, nitrogen loading, and carbon-nitrogen ratio.
7. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 1, characterized in that, The construction of the two-stage intelligent optimization model for the activated sludge system is specifically as follows: Non-dominated ranking and selection involves ranking individuals in the population in a non-dominated manner, dividing individuals into different non-dominated levels based on their fitness values, and selecting parents based on the non-dominated levels and crowding information to generate the next generation of the population. Genetic manipulation involves performing crossover and mutation operations on selected parent individuals to generate new offspring individuals. Crossover operation includes randomly selecting two parent individuals and exchanging some of their genes to generate two new offspring individuals. Mutation operation includes randomly selecting a gene in an offspring individual and modifying it to increase population diversity. Population updates and iterations involve merging newly generated offspring individuals with parent individuals to form a new population. The new population undergoes non-dominated sorting and crowding calculations. Individuals with high fitness and uniform distribution are selected as the next generation population. The non-dominated sorting, selection, and genetic operations are repeated until the preset number of generations is reached or the stopping condition is met.
8. The two-stage activated sludge system optimization method based on microbial community prediction according to claim 7, characterized in that, The optimal solution is measured by an equilibrium value, which minimizes the difference between objective function values within the Pareto optimal solution set that satisfies the conditions. Specifically: in: Eq This represents the equilibrium value of the three objective functions in the Pareto optimal solution; COD In the Pareto optimal solution COD Removal rate; TP for COD In the Pareto optimal solution TP Removal rate; TN for COD In the Pareto optimal solution TN Removal rate; k This is the equilibrium value coefficient.
9. A two-stage activated sludge system optimization system based on microbial community prediction, characterized in that, include: The first data processing module is used to construct the first-stage microbial community succession prediction model, obtain data on the microbial community in activated sludge, combine the environmental parameters and influent data of the reactor, decompose the microbial community into diversity and abundance, incorporate time factors to track the dynamic changes of the microbial community over time, and predict the state of the microbial community at a specific future moment through time series enhancement strategies. The second data processing module is used to construct a pollutant removal efficiency prediction model for the second-stage activated sludge system. Using the output and control indicators of the first-stage model as input, it predicts the effluent indicators of the activated sludge and evaluates the pollutant treatment performance of the activated sludge system. The model optimization module is used to integrate the first-stage and second-stage prediction models with intelligent algorithms and non-dominated sorting genetic algorithms to construct a two-stage intelligent optimization model for activated sludge system. The objective function is the removal rate of chemical oxygen demand, total nitrogen and total phosphorus. The solution space is the regulation index. By adjusting the regulation index, the succession of the microbial community is directionally regulated to achieve the maximum removal rate of pollutants in the activated sludge system. The construction of the two-stage intelligent optimization model for the activated sludge system specifically includes problem definition and model initialization, fitness assessment, non-dominated sorting and selection, genetic operations, population update and iteration, optimal solution selection and output, as well as verification and application. Problem definition and model initialization: The goal is to find the combination of regulatory index parameters that maximizes the removal rates of chemical oxygen demand, total nitrogen, and total phosphorus under the current microbial community state, and to initialize the parameters of the non-dominated sorting genetic algorithm, including population size, number of generations, crossover probability, and mutation probability. At the same time, an initial population representing candidate solutions for the regulatory index parameters is randomly generated. Fitness assessment involves substituting each individual in the initial population into the pollutant removal efficiency prediction model of the activated sludge system. The individual is defined as a combination of control index parameters. The corresponding chemical oxygen demand, total nitrogen, and total phosphorus removal rates are calculated, and the fitness value of each individual is calculated based on the removal rates. The optimal solution is selected and output. The individual with the highest fitness and uniform distribution is selected from the final population as the optimal solution, and the combination of regulatory index parameters corresponding to the optimal solution is output as the optimal regulation scheme to meet the removal rates of chemical oxygen demand, total nitrogen and total phosphorus under the current microbial community state. Verification and application: The obtained optimal combination of control index parameters is applied to the actual operation of the activated sludge treatment system to verify its effectiveness under actual conditions. Based on the verification results, the model is adjusted and optimized to improve prediction accuracy and practical application effect.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the two-stage activated sludge system optimization method based on microbial community prediction as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Microflora-based sewage treatment aeration system control method
CN114573096A
Breeding object pathology deep learning prediction method based on microbial community abundance
CN117272176A