AI-based dynamic adaptive control system for aeration volume in aerobic tanks of wastewater treatment plants

CN122568950APending Publication Date: 2026-08-14WOVER (NANJING) ENVIRONMENTAL ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明提供污水处理好氧池曝气量动态自适应的AI调控系统,解决相关技术中在多分区好氧曝气共用鼓风与管网条件下,因分区耗氧信息获取滞后、曝气传氧效能随运行时间退化及区间气流耦合干扰而导致的溶解氧安全约束与供气能耗协同控制困难的技术问题

Benefits of technology

通过间歇主动探测与安全触发调度获取分区耗氧速率时序,并以长短时记忆网络融合进水负荷特征输出需氧量预测矩阵,再在氧质量平衡框架下将预测矩阵当前时步值与溶解氧实测变化量联立反算各分区实时效能系数;结合阶梯曝气测试建立的传氧效能标定曲线与气流量-风阀开度标定曲线,为反算与反解提供一致基准。执行阶段以效能系数对目标风量补偿后,再以历史数据辨识的区间气流耦合补偿矩阵反解各分区风阀开度并换算鼓风机频率指令,使控制链条在信息侧对齐生物需氧变化,在执行侧贴合多分区管网耦合与设备退化现实,从而提高溶解氧约束的可预期性与指令落地的可控性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122568950A_ABST
    Figure CN122568950A_ABST
Patent Text Reader

Abstract

This invention relates to the field of wastewater treatment and automatic control technology, and discloses an AI-based dynamic adaptive aeration control system for aerobic tanks in wastewater treatment. The system includes an OUR (Oxygen Demand) identification and equipment calibration module, an oxygen demand prediction module, an efficiency identification module, a constraint decision module, an execution command generation module, and an online iteration module. Driven by reinforcement learning strategies within a framework of intermittent active detection and long short-term memory network prediction, oxygen mass balance efficiency identification, and constrained Markov decision processes, the system completes the allocation of target airflow to multiple zones and generates execution commands for air valves and blowers. Priority experience playback and model fine-tuning maintain synchronous updates between the strategy and predictions. By coordinating the allocation and total amount adjustment under dissolved oxygen safety constraints in each zone, this invention improves the dissolved oxygen constraint satisfaction capability of aerobic tanks and the adaptability of aeration control to influent shocks and equipment aging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wastewater treatment and automatic control technology, and more specifically, to an AI control system for dynamically adaptive aeration of aerobic tanks in wastewater treatment. Background Technology

[0002] In urban wastewater treatment plants, the aerobic biological treatment stage typically employs a combination of dissolved oxygen setpoint feedback control and variable frequency blowers to maintain aeration intensity. This approach relies on online dissolved oxygen as the controlled variable and is easily implemented under stable operating conditions. However, when factors such as influent load fluctuations, seasonal water temperature changes, aeration disc blockage, and blower efficiency degradation are combined, the correlation between the control target and biological oxygen demand is weakened. Furthermore, when multiple zones share a common pipe network, the mutual influence of air supply is difficult to fully characterize using a single feedback loop. If the engineering design simplifies each zone into an independent regulation object, the aforementioned mismatch risks will be further amplified.

[0003] In existing technologies, some solutions introduce data-driven models to predict oxygen demand or optimize blower frequency, but they often rely mainly on short-term dissolved oxygen deviations, which are insufficient for obtaining online oxygen consumption rates for different zones. There is a lack of closed-loop correction between prediction and execution to assess the actual oxygen transfer efficiency of the equipment. When multiple zones share blowers and pipe networks, the airflow redistribution caused by valve regulation causes the actual air volume of each zone to deviate from the commanded value. Without coupling compensation, the allocation strategy will have a systematic deviation at the execution level. The engineering fit for online hard constraints and multi-actuator collaboration is still insufficient.

[0004] Therefore, when shock loads and equipment aging coexist, there is a dilemma that dissolved oxygen may drop below the safety limit for a short time or excessive aeration may occur in order to ensure safety. It is difficult to balance the stability of effluent quality and the control of aeration energy consumption. The process operation relies too much on human experience. A dynamic adaptive control scheme for aeration volume is needed that connects oxygen consumption perception, forward prediction, efficiency identification, constraint decision-making and coupled execution. Summary of the Invention

[0005] This invention provides an AI control system for dynamically adaptive aeration volume in aerobic tanks of wastewater treatment plants. It solves the technical problems in related technologies, such as the difficulty in coordinating dissolved oxygen safety constraints and air supply energy consumption control under the condition of shared blowers and pipe networks for aerobic aeration in multiple zones, due to the lag in obtaining oxygen consumption information of each zone, the degradation of aeration oxygen transfer efficiency over time, and the interference of airflow coupling between zones.

[0006] This invention provides an AI-based control system for dynamically adaptive aeration volume in aerobic tanks of wastewater treatment plants, comprising: The identification and equipment calibration module collects dissolved oxygen data for each zone, intermittently reduces aeration when dissolved oxygen is higher than the safety threshold, and linearly fits the dissolved oxygen decay time sequence to identify the oxygen consumption rate time sequence of each zone. At the same time, it establishes the oxygen transfer efficiency calibration curve and air flow-valve opening calibration curve for each zone through stepped aeration test. The oxygen demand prediction module outputs the future N-step oxygen demand prediction matrix for each zone through a long short-term memory network, based on the oxygen consumption rate time series and influent load characteristics. The efficiency identification module takes the predicted value of the prediction matrix at the current time step and the measured change in dissolved oxygen, and calculates the real-time efficiency coefficient of each partition using the oxygen mass balance equation. The constrained decision module, with the prediction matrix, the real-time efficiency coefficient of each zone and the measured total air volume of the blower as the state, outputs the target air volume of each zone under the constrained Markov decision framework through the Lagrange constrained soft actor-critic algorithm, which satisfies the lower limit constraint of dissolved oxygen safety in each zone. The execution instruction generation module divides the target air volume of each zone by the real-time efficiency coefficient to obtain the compensated target air flow rate. It then uses the interval airflow coupling compensation matrix identified by historical data to inversely solve the valve opening instructions for each zone and converts the blower frequency instructions into the total compensated air flow rate. The online iterative module updates the reinforcement learning strategy according to the priority of experience playback based on the regulation cycle, and fine-tunes the oxygen demand prediction model daily.

[0007] Preferably, before triggering detection, the OUR identification and equipment calibration module calculates the safety trigger threshold for each zone by summing the product of the dissolved oxygen safety lower limit, the oxygen consumption rate of the last identification, and the longest detection duration. When the dissolved oxygen in each zone is higher than the safety trigger threshold, the detection is initiated. The system performs detection in turn according to the detection scheduling queue of each partition, and at most one partition can be in the detection state at any one time.

[0008] Preferably, when the OUR identification and equipment calibration module performs detection, the opening degree of the target zone air valve is reduced to the preset minimum opening degree, the residual oxygen transfer rate during the detection period is calculated using the effective oxygen transfer benchmark model, and the residual oxygen transfer rate is subtracted from the oxygen consumption rate obtained by linear fitting of dissolved oxygen decay time sequence to obtain the corrected oxygen consumption rate identification value. When conducting a stepped aeration test, the system is considered to be in a steady state when the dissolved oxygen level stabilizes at each setting until the variation in the continuous sampling points is lower than the preset steady-state threshold.

[0009] Preferably, the input feature vector of the long short-term memory network of the oxygen demand prediction module consists of the historical time series of oxygen consumption rate of each zone, the time series of influent water quality and flow rate, and the feature encoding of water temperature and time period.

[0010] Preferably, the real-time efficiency coefficient calculation steps of the efficiency identification module are as follows: approximate the average dissolved oxygen of the time step by the average value of the measured dissolved oxygen at the start and end of the time step; query the oxygen transfer coefficient corresponding to the current valve opening by the oxygen transfer efficiency calibration curve, and obtain the nominal oxygen transfer increment by multiplying the difference between saturated dissolved oxygen and average dissolved oxygen by the time step length. The actual oxygen transfer increment is obtained by multiplying the measured dissolved oxygen change at each time step by the predicted value of the current time step and the time step length; the real-time efficiency coefficient is the ratio of the actual oxygen transfer increment to the nominal oxygen transfer increment; when the nominal oxygen transfer increment is lower than the preset minimum calculation threshold, the efficiency coefficient estimate of the previous time step is maintained.

[0011] Preferably, the efficiency identification module adopts a dual-mode strategy for updating the efficiency coefficient: when the water intake condition is stable, the calculation result of the current step is included in the weighted moving average update with the first update weight; When the change in influent water quality relative to the moving average exceeds the preset change judgment threshold within a preset time window, the system switches to the second update weight for gradual updates until the operating conditions return to stability; the first update weight is greater than the second update weight.

[0012] Preferably, the reward function of the constraint decision module is composed of the main reward item of energy consumption and the auxiliary reward item of dissolved oxygen deviation, weighted by a preset weight coefficient: the main reward item is the negative value of the actual energy consumption of the blower in the current step after normalization based on the rated power; The auxiliary reward item is the negative value of the deviation of the measured dissolved oxygen in each region from the ideal dissolved oxygen operating point, normalized based on the difference between saturated dissolved oxygen and the lower limit of dissolved oxygen; the ideal dissolved oxygen operating point in each region is calculated and determined based on the current saturated dissolved oxygen, the current time step predicted value of the oxygen demand prediction matrix, and the real-time efficiency coefficient.

[0013] Preferably, the constraint decision module establishes independent Lagrange multipliers for each aeration zone of the aerobic tank: when the dissolved oxygen in a certain zone is lower than the safe lower limit of dissolved oxygen, the corresponding zone multiplier is automatically increased to enhance the penalty weight for constraint violation, and when the constraint is satisfied, the corresponding multiplier is gradually decreased; Each partition multiplier is adjusted independently during each policy update using the dual gradient descent method, and the partition multipliers do not interfere with each other.

[0014] Preferably, in the instruction generation module, the interval airflow coupling compensation matrix is ​​identified by multivariate regression of historical operating data, and each element of the matrix is ​​fitted with the change in the opening of the air valve in each zone as the independent variable and the change in the measured airflow in each zone as the dependent variable. Establish corresponding coupling matrices for the first and second operating frequency ranges of the blower, and select the corresponding version according to the current blower frequency during real-time compensation; use the deviation between the compensated target air flow and the measured air flow as the residual for each control cycle, and perform online correction of the matrix increment by recursive least squares method.

[0015] Preferably, the online iterative module uses a fast cycle with a control period: it calculates immediate rewards based on the actual energy consumption and dissolved oxygen measured values ​​of each zone, generates constraint cost signals based on whether the dissolved oxygen in each zone is below the dissolved oxygen safety lower limit, samples from the experience buffer pool with priority based on the absolute value of the time-series difference error to perform gradient updates on the policy network, and independently adjusts the corresponding Lagrange multipliers according to the constraint cost signals of each zone; and performs fine-grained fine-tuning of the oxygen demand prediction model on a daily basis.

[0016] The beneficial effects of this invention are as follows: The oxygen consumption rate time series of each zone is obtained through intermittent active detection and safety-triggered scheduling. An oxygen demand prediction matrix is ​​output by integrating influent load characteristics using a long-short time memory network. Then, under the oxygen mass balance framework, the current time step value of the prediction matrix is ​​combined with the measured change in dissolved oxygen to calculate the real-time efficiency coefficient of each zone. Combined with the oxygen transfer efficiency calibration curve and airflow-valve opening calibration curve established by the stepped aeration test, a consistent benchmark is provided for the back-calculation and back-solution. During the execution phase, the target airflow is compensated with the efficiency coefficient, and then the interval airflow coupling compensation matrix identified from historical data is used to back-solution the valve opening of each zone and convert it into blower frequency commands. This ensures that the control chain aligns with biological oxygen demand changes on the information side and conforms to the reality of multi-zone pipe network coupling and equipment degradation on the execution side, thereby improving the predictability of dissolved oxygen constraints and the controllability of command implementation. Within the constrained Markov decision process framework, a Lagrange-constrained soft actor-critic algorithm is employed. Constraint costs are constructed using the dissolved oxygen safety lower limit for each zone, and independent Lagrange multipliers are assigned to each zone, achieving a balance between zoned constraint response and energy consumption targets. The reward signal integrates the deviations between blower energy consumption and dissolved oxygen relative to the ideal operating point, mitigating the learning difficulties caused by the lag in online feedback of effluent water quality. The strategy output, under the reference constraint of the measured total blower airflow, forms the target airflow for each zone, providing a decomposable decision interface for the execution chain. On the online iterative side, a fast-loop control cycle is used to complete priority experience-driven strategy updates, and the oxygen demand prediction model is fine-tuned on a daily scale. This ensures that allocation optimization and total blower volume adjustment maintain cyclical consistency despite differences in response speed. The strategy and model can be continuously corrected during operating condition drift, reducing the uncertainty and maintenance burden of frequent manual parameter tuning. Attached Figure Description

[0017] Figure 1 This is a block diagram of the AI ​​control system for dynamic adaptive aeration volume in the aerobic tank of wastewater treatment according to the present invention. Figure 2 This invention relates to a process flow of an AI-based system for dynamically adaptive aeration volume control in aerobic tanks of wastewater treatment. Figure 1 ; Figure 3 This invention relates to a process flow of an AI-based system for dynamically adaptive aeration volume control in aerobic tanks of wastewater treatment. Figure 2 . Detailed Implementation

[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0019] In one embodiment of the present invention, focusing on the aerobic biochemical treatment section in the field of urban wastewater treatment, a dynamic adaptive control method for aeration volume is proposed. This method addresses the systemic technical challenges commonly found in multi-zone aeration systems of aerobic tanks, such as the inability to sense biological oxygen consumption rates online, the continuous degradation of aeration equipment performance over time, and mutual interference between airflow regulation in multiple zones. The method uses intermittent active detection and identification of zone oxygen consumption rates as the sensing core and deep reinforcement learning with a constrained Markov decision process framework as the decision-making core.

[0020] Taking a wastewater treatment plant in a northern city as an example, the plant employs an anaerobic-anoxic-aerobic (AAO) biological treatment process. The aerobic zone has a corridor-style partitioned structure, with three independent aeration zones (hereinafter referred to as Zone 1, Zone 2, and Zone 3) set up along the wastewater flow direction. Each zone is equipped with electrically adjustable air valves and online dissolved oxygen (DO) sensors. The three zones share one set of variable frequency centrifugal blowers and one common air supply main. The blower unit has been operating continuously for more than 5 years, and some micropores of the aeration discs have become clogged. The actual oxygen transfer efficiency of each zone deviates from the design parameters to varying degrees, and the degree of deviation varies between zones. The influent water quality of the plant fluctuates frequently due to the influence of industrial wastewater, with a temperature difference of nearly 20 degrees Celsius between winter and summer. The oxygen consumption metabolic rate of the activated sludge in the aerobic zone exhibits obvious seasonal variations combined with diurnal cycle fluctuations. Against this backdrop, existing control methods, represented by traditional DO setpoint PID control, exhibit significant response lag under influent impact conditions and continuously expanding command-execution deviation under aging equipment conditions. These problems make it difficult to achieve ideal levels of effluent water quality stability and energy efficiency in aeration systems.

[0021] The control method of this invention consists of six modules, forming a complete closed loop of "sensing-prediction-identification-decision-execution-iteration". The identification and equipment calibration module completes the system acquisition of multi-source sensor data and identifies the real-time oxygen consumption rate time series of each zone online through an intermittent active detection mechanism. Simultaneously, it calibrates two sets of benchmark parameters through stepped aeration tests: the KLa-valve opening oxygen transfer efficiency calibration curve and the air flow-valve opening calibration curve. The oxygen demand prediction module uses the oxygen consumption rate time series of each zone as the core prediction variable, integrates influent load characteristics, and processes the data through a multivariate long short-term memory (LSTM) network. The system uses a memory module to predict the oxygen demand vector for each zone during future control cycles. The efficiency identification module uses the current time step OUR prediction value in the first column of the oxygen demand prediction matrix as a known quantity. Combined with the current aeration execution feedback and oxygen transfer efficiency baseline curve, it calculates the real-time efficiency coefficients and effective oxygen transfer models for each zone's equipment online using the zone's oxygen mass balance equation. The constraint decision module uses the future oxygen demand trends for each zone described in all N columns of the oxygen demand prediction matrix and the real-time efficiency coefficients of the equipment identified by the efficiency identification module as core state quantities. Within the framework of a Constrained Markov Decision Process (CMDP), it uses a Lagrangian constraint soft actor-commentator (LCB) algorithm that has undergone offline pre-training and online iteration. The SAC algorithm uses the current measured total air volume of the blower as a reference constraint to output the optimal aeration volume allocation ratio for each zone. The execution instruction generation module, based on the target air volume of each zone output by the constraint decision module, first performs efficiency compensation using the real-time efficiency coefficient of each zone to obtain the compensated target airflow. Then, it solves the valve opening execution instruction for each zone through the interval airflow coupling compensation matrix. At the same time, it determines the blower operating frequency instruction based on the sum of the compensated airflow of each zone and sends it to each actuator via PLC (Programmable Logic Controller). The online iteration module completes the online iteration of reward evaluation, experience storage and reinforcement learning strategy in a fast cycle of 5 minutes, and completes the adaptive update of the LSTM prediction model in a slow cycle of 24 hours, so that the whole system continuously self-optimizes during operation. The constraint decision module focuses on solving the optimal allocation ratio among the zones under the current total air volume reference constraint. Based on this, the execution instruction generation module determines the blower frequency according to the actual compensated total demand. The two form a two-level control structure where "allocation optimization" and "total volume regulation" are decoupled, which is suitable for the engineering reality that the blower's variable frequency speed regulation response is relatively slow, while the damper's adjustment response is relatively fast. At least one embodiment of this invention discloses an AI control system for dynamically adaptive aeration volume in a wastewater treatment aerobic tank, such as... Figures 1 to 3 As shown, it includes: The identification and equipment calibration module collects dissolved oxygen data for each zone, intermittently reduces aeration when dissolved oxygen is higher than the safety threshold, and linearly fits the dissolved oxygen decay time sequence to identify the oxygen consumption rate time sequence of each zone. At the same time, it establishes the oxygen transfer efficiency calibration curve and air flow-valve opening calibration curve for each zone through stepped aeration test. This module has two core tasks: first, to establish a multi-source real-time sensor data acquisition system required for system operation; and second, to identify the real-time oxygen uptake rate (OUR) of activated sludge in each zone online through an intermittent active detection mechanism. The basic principle of intermittent detection is that after a brief reduction in aeration in a zone, the rate of decrease in dissolved oxygen (DO) in that zone is approximately equal to the current biological oxygen uptake rate of the activated sludge. Applying this principle to a continuously operating industrial system presents a fundamental engineering contradiction: detection requires aeration reduction, which directly threatens the safety of aerobic biological treatment. This module addresses this contradiction by designing a safety threshold trigger logic, a minimum aeration rate retention anti-clogging mechanism, dynamic detection duration control, and a substitution estimation mechanism for low DO levels. This makes online OUR identification possible without affecting treatment efficiency, which is the core design element of engineering this sensing principle into the online control system.

[0022] At the data acquisition level, the system synchronously collects raw data from the following four types of sensors: online DO sensors (resolution not less than 0.01 mg / L) in each zone of the aerobic tank to obtain the real-time dissolved oxygen concentration of each zone; feedback sensors on the opening of the air valves and the air flow sensors of each zone's aeration system to obtain the current actual air supply status of each zone; online water quality instruments installed at the inlet of the aerobic tank to collect operating characteristic quantities such as influent COD concentration, ammonia nitrogen concentration, flow rate, and water temperature; and variable frequency drive operating data of the aeration blower units, including the current operating frequency of the blower, outlet air pressure, and total air volume. The unified acquisition cycle for the above data is set to once every 10 seconds, and after acquisition, it is uploaded to the upper computer control system for real-time processing via industrial Ethernet. The water quality data at the outlet of the aerobic tank is provided by an online water quality analyzer. Due to the 20 to 40 minute analysis delay in effluent water quality detection, its data is mainly used in the strategy evaluation stage of the online iterative module and does not directly participate in the real-time control calculation from the identification and equipment calibration module to the execution instruction generation module.

[0023] For detection scheduling, the system maintains a detection scheduling queue for each partition, performing detection on each partition in turn to ensure that at most one partition is in detection mode at any given time, while the remaining partitions maintain normal aeration. Before triggering detection in a partition, the system first calculates the safety threshold for this detection: the safety threshold equals the lower limit of DO safety plus the product of the estimated OUR value of the partition obtained from the previous identification and the longest detection duration (1.5 minutes), ensuring that even under the longest detection duration, the partition's DO will not fall below the safety lower limit (the default value of the lower limit of DO safety is 2.0 mg / L, and for conditions that need to simultaneously ensure nitrification, this value is not lower than 1.5 mg / L). When the system first runs, each partition has no historical identification records, so the initial estimated OUR value (default value is 0.5 mg / L·min) is substituted into the safety threshold calculation. After the first detection, the actual identification value is replaced, and thereafter it is updated on a rolling basis. If the current DO is higher than the safety threshold calculated this time, the detection condition is confirmed to be met, and detection is started; otherwise, the detection is postponed.

[0024] Once safety conditions are met, the system immediately reduces the opening of the electric air valve in the target zone to 15% to 20% of its current normal operating opening. This minimum opening is maintained instead of a complete shutdown because if the aeration disc operates under extremely low airflow for an extended period, the pressure differential in the membrane pores drops sharply, easily exacerbating sludge accumulation and blockage on the micropore surface. Maintaining a small amount of airflow is a proactive protection for the equipment's lifespan. The reduction duration is dynamically determined based on the margin between the DO (displacement) level before reduction and the safety lower limit. When the margin is sufficient, it can last up to 90 seconds; when the margin is tight, it is shortened to 60 seconds. Throughout the reduction period, the DO sensor continuously records the DO value at 10-second sampling intervals, forming a set of DO decay time-series data consisting of 6 to 9 sample points. A first-order linear fit is performed on this data to obtain the DO decay slope over time; its absolute value is the preliminary estimate of the OUR (Original Rate of Decay). Since the safety threshold design ensures that the initial DO concentration for detection is not lower than 2.0 mg / L, and the change in the oxygen consumption rate of activated sludge with DO concentration within this DO range does not exceed 1% (the Monod half-saturation constant for activated sludge oxygen consumption is on the order of 0.1 to 0.2 mg / L, which is far lower than the DO level within the detection range), the linear fitting assumption holds under engineering precision, and there is no need to introduce a nonlinear correction term.

[0025] Because the minimum aeration opening was maintained at 15% to 20% during the detection period, a small amount of residual oxygen transfer remained in the zone, which needed to be subtracted from the fitted slope. The system calculated the estimated residual oxygen transfer rate based on the actual valve opening and blower operating parameters during the detection period, combined with the effective oxygen transfer baseline model updated by the efficiency identification module at the end of the previous control cycle. After subtraction, a more accurate OUR identification value was obtained. The use of the baseline model updated in the previous cycle is explicitly stated here to ensure the unidirectional transmission of parameters between modules: at the end of each 5-minute control cycle, the efficiency identification module uniformly refreshes the baseline model parameters, and the identification and equipment calibration modules of the next control cycle directly use the refreshed version, without any cyclic calls within the same time step. After detection, the target zone's valves immediately returned to their normal opening before detection, and the zone's DO returned to normal levels within 1 to 2 minutes.

[0026] When a sudden change in influent load causes the dissolved oxygen (DO) of a certain zone to remain below the detection safety threshold, the detection is postponed. During this period, the system activates a backup estimation mechanism: it reads the influent COD concentration and flow rate data from online water quality instruments at the influent end, and combines this with the system's internally stored historical OUR-influent load mapping model (established from long-term historical operational data statistics) to perform a substitute estimation of the current zone's OUR, marking it as low-confidence data for subsequent modules. Once the DO recovers to above the detection safety threshold, the system prioritizes performing a compensation detection on that zone, replacing the substitute estimated value with the actual identified value to restore the integrity of the OUR time series. After each detection, the identified OUR value is recorded with a timestamp and stored independently as a sliding time series window for each zone, forming the OUR_1, OUR_2, and OUR_3 time series sequences, which serve as the core input to the oxygen demand prediction module's prediction model.

[0027] Regarding the identification of baseline performance parameters, during the initial deployment of the system and after each subsequent maintenance of large equipment, a systematic stepped aeration test is conducted on each zone during a period of relatively stable influent conditions. The opening degree of the air valves in each zone is set sequentially to typical opening levels of 30%, 50%, 70%, and 90%. Each level is stabilized until the reading variation of DO at 3 to 5 consecutive sampling points (i.e., 30 to 50 seconds) in that zone is less than 0.05 mg / L, which is considered a steady state. The steady-state DO rise rate and the actual air flow rate of the zone are recorded under this state. If the steady-state judgment condition is not met within 10 minutes at a certain level, the average value of the DO readings in the next 5 minutes at that level is taken as the steady-state estimate. The stepped aeration test simultaneously outputs two sets of calibration data: First, it uses the steady-state DO rise rate at each level to back-calculate the corresponding oxygen transfer coefficient (KLa) value through the zone oxygen mass balance equation, constructing a KLa-valve opening oxygen transfer efficiency calibration curve for use by the efficiency identification module. Second, it records the measured zone airflow at each level, constructing an airflow-valve opening calibration curve for use by the execution command generation module when converting the compensated target airflow into valve opening. Both calibration curves are stored together with operating condition markers (time, water temperature, etc.), jointly forming the aeration equipment efficiency profile for that zone.

[0028] For scenarios involving initial deployment and lack of historical operational data for the plant, a three-stage progressive startup scheme is adopted. The first stage lasts approximately two weeks, during which the system only runs the intermittent detection and identification and baseline performance calibration of the identification and equipment calibration modules. Aerobic tank aeration is maintained using the existing DO-PID method, while OUR identification time-series data is accumulated, baseline performance parameter calibration is completed, and historical data required for coupling matrix identification is accumulated through conscious zone-based step-by-step adjustment operations. In the second stage (weeks 2-8), the accumulated OUR data is combined with commonly used activated sludge kinetic parameters in the field for initial training of the LSTM. A proportional rule strategy (the target aeration rate for each zone is proportional to the oxygen demand predicted by the LSTM, and compensated according to the performance coefficient) replaces the constraint decision module, achieving forward-looking zone-based differentiated aeration control. Simultaneously, the reinforcement learning strategy is pre-trained offline in a simulation environment built using historical data. In the third stage (after week 8), after the simulation pre-training is completed and sufficient online data has been accumulated, the reinforcement learning strategy takes over the decision-making function of the constraint decision module, and the system enters the complete six-module operation mode. Thereafter, the LSTM of the oxygen demand prediction module and the reinforcement learning strategy of the constraint decision module are continuously optimized by the online iteration mechanism of the online iteration module.

[0029] The final outputs of the identification and equipment calibration module are: the OUR time series for each zone (OUR_1, OUR_2, OUR_3, sliding window data), the KLa-valve opening oxygen transfer efficiency calibration curve (efficiency profile) for each zone, and the air flow-valve opening calibration curve (air flow profile) for each zone. The OUR time series serves as the core input to the oxygen demand prediction module's prediction model; the oxygen transfer efficiency calibration curve provides a benchmark for the efficiency coefficient identification of the efficiency identification module; and the hydraulic calibration curve provides a basis for the valve opening conversion of the execution command generation module.

[0030] The oxygen demand prediction module outputs the future N-step oxygen demand prediction matrix for each zone through a long short-term memory network, based on the oxygen consumption rate time series and influent load characteristics. The OUR time series identified by the identification and equipment calibration module describes the historical dynamics of biological oxygen demand in each zone. However, due to the frequency limitation of intermittent detection (typically, a detection cycle is completed for each zone every 30 to 60 minutes), there are time intervals between the detection and identification results, making it impossible to directly use as the basis for real-time oxygen demand in each control cycle (a decision is made every 5 minutes). More importantly, relying solely on historical OUR data for the current moment's response is essentially a delayed perception. When a large amount of high-concentration organic wastewater suddenly floods the influent, the OUR will inevitably rise rapidly in the following few minutes to more than ten minutes. If the control system has not yet completed the next detection and identification, the sudden increase in oxygen demand will still be invisible to the control system. Therefore, the task of this module is to: use the historical OUR time series of each zone as the core, integrate the load characteristics that can be obtained in real time from the influent end, and actively predict the changes in oxygen demand in each zone in future control cycles through a multivariate time series prediction model, so as to make the control decision more forward-looking.

[0031] The predictive model employs a multivariate long short-term memory (LSTM) network architecture. LSTM selectively retains and updates historical information through three gating mechanisms: a forget gate, an input gate, and an output gate. The forget gate determines which historical information to discard from the memory unit based on the current input and the hidden state from the previous time step. The input gate controls the proportion of new information written into the memory unit at the current time step. The output gate determines which information in the memory unit is output as the hidden state at the current time step. This mechanism enables LSTM to capture patterns across different time scales in the OUR time series, including intraday periodic fluctuations (OUR peaks caused by morning and evening peak inflows) and longer-term seasonal drifts (slow changes in microbial metabolic rates caused by water temperature variations). Both are dynamic patterns that require proactive responses in aeration rate regulation.

[0032] The input feature vector to the prediction model consists of two types of data with different properties, which complement rather than redundant each other in the prediction mechanism. The historical time series window of the OUR (Oxygen Demand) for each zone provides a "metabolic state baseline" for the biological system, representing the level of basic oxygen demand capacity of the activated sludge in the aerobic zone under the current water temperature, sludge age, and microbial community state. For example, comparing winter and summer, the same influent COD shock will cause drastically different OUR responses in summer (water temperature 28℃) and winter (water temperature 8℃). This difference in biological state is precisely reflected in the recent historical OUR time series, which LSTM perceives by learning the background trend of the OUR time series. Influent COD concentration, influent flow rate, and their historical time series provide "load disturbance feedforward," enabling the prediction model to predict the direction of oxygen demand change in advance based on the real-time signal from the influent within the control interval before the OUR detection is updated. Both types of input are indispensable: relying solely on influent COD cannot distinguish the differences in response between different biological states under the same load; relying solely on historical OUR cannot perceive the immediate load shock caused by sudden changes in influent. In addition, historical and current water temperature values ​​are incorporated as key environmental parameters for biological metabolic rates. The time-segment characteristics of the current moment are added using sine-cosine periodic encoding (mapping the current moment of a 24-hour period into sine and cosine values ​​respectively), enabling the model to perceive intraday load cycles such as morning and evening peaks. The length of the OUR historical time-series window for each zone is determined based on the autocorrelation analysis results of historical data from the stations, with a typical setting covering records from the past 3 to 6 hours.

[0033] Before entering the LSTM, the various numerical values ​​in the input feature vector are normalized using the mean and standard deviation of each feature obtained from historical operating data to eliminate the interference of different units on model training. For time gaps in the OUR time series caused by limited detection frequency, linear interpolation driven by the influent COD trend and water temperature information during that period is used to fill the gaps and maintain the continuity of the input time series. The output layer of the LSTM network is designed as a multi-step prediction structure. The model outputs the predicted oxygen demand values ​​for each zone for the next N control time steps (with a period of 5 minutes per step, N is usually set to 6, i.e., predicting the next 30 minutes) at once. The output format is a 3-row N-column prediction matrix, with each row corresponding to a zone and each column corresponding to a future time step. The output values ​​are denormalized to restore the physical units (unit: mg / L·min).

[0034] The training of the LSTM network is divided into two stages: offline pre-training and online fine-tuning. In the offline pre-training stage, a training sample set is constructed using the plant's historical operational data (including historical OUR identification records, influent water quality records, water temperature records, etc., typically covering at least one year to encompass the complete seasonal variation cycle; for initial deployments with insufficient historical data, data augmentation can be performed using commonly used activated sludge kinetic parameters in this field, followed by pre-training with publicly available data from similar process plants and then fine-tuning with the plant's own data). The multi-step prediction weighted mean square error is used as the loss function (with higher weights given to prediction errors in more recent future steps), and the Adam optimizer is used to iteratively train the network weights until the prediction accuracy on the validation set meets the preset requirements. The specific execution method of the online fine-tuning stage is detailed in the online iteration module.

[0035] The output of the oxygen demand (OD) prediction module is an OD prediction matrix (3 rows and N columns) for each zone over the next N control time steps. Within each control cycle, the first column of this matrix (corresponding to the predicted OUR value for the current time step) serves as the OD input for the performance identification module's oxygen mass balance back-calculation, used to calculate the performance coefficient for the current time step. All N columns of this matrix (reflecting the OD change trend over the next N steps) serve as the state vector input for the constraint decision module, supporting the forward-looking allocation decisions of the policy network. The baseline performance parameters output by the identification and equipment calibration module are also input to the performance identification module along with the prediction matrix.

[0036] The efficiency identification module takes the predicted value of the prediction matrix at the current time step and the measured change in dissolved oxygen, and calculates the real-time efficiency coefficient of each partition using the oxygen mass balance equation. The oxygen transfer efficiency of aeration equipment is a key parameter connecting the "control command" and the "actual oxygen supply effect." As the aeration disc diaphragms age and their micropores gradually become clogged in the underwater environment over time, and as the efficiency of the blower impeller slowly declines with operating time, the actual oxygen transfer under the same control command will consistently be lower than the nominal value. Furthermore, the aging degree of equipment in different zones varies, leading to different deviations in oxygen transfer efficiency between zones. If the control system does not detect and compensate for this deviation online, the "optimal aeration allocation scheme" output by the constraint decision module will produce systematic errors during actual execution due to efficiency deviations, failing to achieve the expected control effect.

[0037] The technical approach of this module is based on the online back-calculation of the zoned oxygen mass balance equation. For any aeration zone, the change in dissolved oxygen within a control time step is equal to the amount of dissolved oxygen transferred into the water by aeration within that time step minus the amount of dissolved oxygen consumed by the activated sludge: the amount transferred in is determined by the oxygen transfer coefficient (KLa) of the aeration equipment, the difference between the measured DO and saturated DO in the water, and the time step length; the amount consumed is determined by the current predicted OUR value of the zone (taken from the current time step predicted value of the corresponding zone in the first column of the prediction matrix of the oxygen demand prediction module) and the time step length; the saturated dissolved oxygen concentration is a function of water temperature and is accurately calculated from the water temperature measurement value using the temperature-saturated DO standard formula, without the need for additional measurement.

[0038] The efficiency coefficient calculation process is completed in four steps. First, the initial and final measured DO values ​​for this control time step in the zone are read. The average of these two values ​​is used to approximate the average DO within this time step, which is then used to calculate the subsequent oxygen transfer driving force. Second, the nominal KLa value corresponding to the current valve opening is retrieved from the oxygen transfer efficiency calibration curve of the identification and equipment calibration module. This value is then multiplied by the product of the KLa value and the difference between the saturated DO and the average DO within this time step, and then multiplied by the time step length to obtain the nominal theoretical oxygen transfer increment. Third, the actual dissolved oxygen increment transferred into the water is obtained by adding the measured DO change (final value minus initial value) of this time step to the product of the current time step OUR predicted value from the oxygen demand prediction module and the time step length. Fourth, the efficiency coefficient η is equal to the actual oxygen transfer increment obtained in the third step divided by the nominal theoretical oxygen transfer increment obtained in the second step. An η close to 1 indicates that the oxygen transfer efficiency is basically consistent with the nominal value, while a value significantly lower than 1 indicates a degradation in oxygen transfer efficiency. When the nominal theoretical oxygen transfer increment is less than the minimum calculation threshold (default value is 0.05 mg / L), it indicates that the current DO is close to saturation and the oxygen transfer driving force is extremely small. The calculation result of the efficiency coefficient in this time step is set to invalid, and the estimated value of the previous step is retained to avoid numerical instability caused by a small denominator. The parameter update of this module adopts the timing arrangement of "updated in the previous control cycle, used in the current control cycle". That is, at the end of each 5-minute control cycle, the efficiency coefficient and effective oxygen transfer model are refreshed uniformly for the identification and equipment calibration module's residual oxygen transfer correction and constraint decision module to use directly in the next cycle, ensuring the unidirectional parameter transfer between modules.

[0039] To improve the stability of efficiency coefficient estimation, the system employs a dual-mode strategy for updating the efficiency coefficient. Under normal operating conditions with stable influent conditions, the current calculation result is included in the weighted moving average update with a higher weight to quickly track slow changes in equipment status. When the influent water quality sensor detects that the increase in influent COD relative to the 30-minute moving average exceeds the first abrupt change threshold (default value is 25%) within 10 minutes, the system automatically switches to a low-weight progressive update mode. During abrupt events, this mode reduces the correction magnitude of the current efficiency estimation result to the historical average, using multi-step accumulated data over a longer time window to smooth the OUR prediction error caused by the abrupt change, until the influent conditions return to stability and the system switches back to normal update mode. The fundamental reason for this switch is that the OUR prediction error of the oxygen demand prediction module will significantly increase during abrupt changes in influent conditions. If the efficiency coefficient is still updated with a high weight, the prediction error will be misinterpreted as a change in equipment efficiency, leading to false fluctuations in the efficiency coefficient.

[0040] Each partition independently maintains its historical time-series record of efficiency coefficients. When the moving average of a partition's efficiency coefficient continuously falls below the second efficiency degradation threshold (the default value is 75% of the initial calibrated efficiency coefficient, i.e., oxygen transfer efficiency deteriorates by more than 25% relative to the initial state), the system sends an equipment performance degradation warning to the host management platform, prompting maintenance personnel to arrange for aeration disc cleaning or replacement maintenance. During the waiting period for maintenance, the control system uses the current real-time efficiency coefficient as the standard and automatically allocates more aeration to that partition during the decision-making phase of the constraint decision module to compensate for the insufficient oxygen transfer efficiency, ensuring that the actual dissolved oxygen supply does not continue to decline due to equipment aging. This module ultimately outputs the real-time efficiency coefficient vector of each partition and the updated effective oxygen transfer model (reflecting the corrected relationship that the effective value of KLa is equal to the nominal value of KLa multiplied by η, reflecting the current equipment efficiency state). These two, together with the oxygen demand prediction matrix of the oxygen demand prediction module, serve as the input state variables for the constraint decision module's optimization decision.

[0041] The constrained decision module, with the prediction matrix, the real-time efficiency coefficient of each zone and the measured total air volume of the blower as the state, outputs the target air volume of each zone under the constrained Markov decision framework through the Lagrange constrained soft actor-critic algorithm, which satisfies the lower limit constraint of dissolved oxygen safety in each zone. This module is the decision-making core of the entire control system, responsible for transforming the sensing and predictive information provided by the first three modules into optimal control actions. The decision objective can be clearly stated as: minimizing the total energy consumption of the aeration system in the current and future control time domains, while ensuring that the dissolved oxygen in each zone of the aerobic tank is maintained within the safe range for biological treatment. This is a continuous action space optimization problem with hard constraints, and standard unconstrained reinforcement learning algorithms cannot be directly applied. This module uses the Constrained Markov Decision Process (CMDP) framework to rigorously establish the mathematical structure of this optimization problem and solves it using the Lagrangian Constrained Soft Actor-Critic (Lagrangian SAC) algorithm, which is suitable for continuous action spaces.

[0042] In establishing the CMDP problem, the state vector contains the following types of information: the current measured DO value of each zone (3 scalars), all N columns of the prediction matrix of the oxygen demand prediction module (the oxygen demand prediction sequence for each zone in the next N steps, enabling the strategy to perceive the trend of oxygen demand changes over a period of time and support forward-looking allocation decisions), the real-time efficiency coefficient of each zone's equipment (from the efficiency identification module, 3 scalars, enabling the strategy to perceive the current oxygen transfer capacity of the equipment), the current measured total air volume of the blower (1 scalar, taken from the real-time reported total outlet flow measurement value of the frequency converter driver, representing the actual total amount of resources available for allocation in this control step), the current water temperature, and the time period feature code (sine and cosine pairs). All the above state components are normalized and then concatenated into a state vector, which serves as the input to the SAC policy network and value network.

[0043] The action space is defined as the aeration airflow allocation ratio for each zone. Specifically, at each control time step, the strategy outputs the proportion of the airflow each of the three zones should receive relative to the current measured total airflow of the blower. The sum of these three proportions is constrained to be 1, and each proportion is between 0 and a preset maximum upper limit. The output layer of the strategy network naturally satisfies the constraint that the sum of the allocation ratios is 1 through the Softmax function. Multiplying the allocation ratio by the current measured total airflow of the blower yields the target absolute airflow for each zone. This target airflow is then converted into valve commands in the execution command generation module after performance compensation and coupling compensation. Simultaneously, the execution command generation module independently determines the blower frequency command based on the sum of the compensated airflow rates for each zone, completing the overall adjustment. It should be noted here that the current measured total air volume of the blower used by the constraint decision module is the reference input of the strategy, not a rigid constraint upper limit for the execution command generation module. After the execution command generation module performs efficiency compensation on the target air volume of each zone, the sum of the actual required air volume of each zone may exceed the total air volume referenced by the constraint decision module when making the decision. At this time, the execution command generation module adjusts the blower frequency according to the actual compensated total demand. The measured total air volume read by the constraint decision module in the next control cycle will be automatically updated to the new value adjusted by the execution command generation module. In the online iteration of the online iteration module, the strategy network gradually learns to reserve efficiency compensation margin in the allocation decision, so that the two-level control tends to be coordinated and consistent through periodic iteration.

[0044] The reward function consists of two parts. The main reward term is based on the negative value of the actual operating energy consumption of the blower during the current control period, normalized to the rated power of the blower, and then taken as a negative value (the normalized range is approximately...). The direct-drive strategy optimizes towards reducing ineffective aeration. The intermediate auxiliary reward is designed based on the deviation of the measured DO from the ideal DO operating point in each partition: the sum of the absolute values ​​of the DO deviations in each partition is normalized using the difference between the saturated DO and the DO safety lower limit as a benchmark, and then negative (with a range of approximately...). The total reward is the product of the main reward item, the weighting coefficient λ, and the auxiliary reward item. The default value of λ is 0.3. Energy consumption optimization is the primary driving force, with DO deviation as a secondary guide. It can be adjusted according to the priority between effluent water quality and energy saving in the actual scenario. Ideal DO operating point. The steady-state level of dissolved oxygen (DO) in a given zone, where the oxygen transfer from aeration exactly balances the oxygen consumption at the current biological oxygen demand rate, is determined by the following formula:

[0045] in Current water temperature The saturated dissolved oxygen concentration (mg / L) at the specified temperature is calculated using the standard equation for temperature-saturated DO. The first column of the prediction matrix for the oxygen demand prediction module. The predicted oxygen demand (mg / L·min) at the current time step for each zone; For identification and equipment calibration module calibration of the first Regional baseline oxygen transfer coefficient (1 / min); The first identified by the performance identification module The real-time performance coefficient (dimensionless) of the partitioned equipment. The core reason for introducing this intermediate reward term is that there is a detection delay of tens of minutes in the effluent water quality test results (effluent COD, ammonia nitrogen). If the strategy relies solely on the delayed effluent water quality as the reward signal, the convergence efficiency is extremely low; however, the performance coefficient can be calculated based on the current time step. The deviation provides high-frequency, low-latency learning feedback and has a clear physical causal relationship with the final effluent quality (when the dissolved oxygen (DO) in each zone is always close to the target effluent quality). At this time, the biochemical degradation process is most efficient, and the effluent quality is optimal, fundamentally solving the reward delay problem in reinforcement learning in wastewater treatment scenarios. When the predicted OUR value for a certain partition is extremely high while the efficiency coefficient is extremely low, the calculation result of the above formula may approach or even fall below the safe lower limit of DO. In this case, the safe lower limit of DO is used instead. Used for intermediate reward calculations to ensure that the reward signal does not contradict the direction of the hard constraint.

[0046] The definition and handling of hard constraints are the core difference between CMDP and standard reinforcement learning. This module defines the requirement that "the DO of each zone must not fall below the safe lower limit threshold after each control step" as a hard constraint that must be met. The LagrangianSAC algorithm establishes independent Lagrange multipliers for each of the three aeration zones. ): When the When a partition DO violates the lower safety constraint, the corresponding Automatically increasing the weight of the penalty term for constraint violation is added to the overall objective function with greater weight overlap, forcing the policy to more aggressively ensure the DO safety of that partition in the next update; when the constraint is satisfied, The Lagrange multipliers are gradually reduced, allowing the strategy to focus more on energy consumption optimization. Each partition's Lagrange multiplier is updated independently without interference, enabling the strategy to respond differently to the constraint satisfaction status of each partition and avoiding the coupling oscillation problem that occurs when a single multiplier simultaneously regulates the constraints of three partitions. The Lagrange multipliers are adjusted synchronously with each policy update using dual gradient descent, with the learning rate set to a value smaller than the policy network's learning rate to ensure smooth multiplier updates.

[0047] The SAC policy network consists of two parts: an actor network and a dual critic network. The actor network takes the current state vector as input and outputs the mean and logarithmic standard deviation of the aeration distribution ratio for each zone through multiple fully connected layers, forming a multivariate Gaussian distribution. During training, specific actions are sampled from this distribution, with exploration noise embedded in the sampling process. The dual critic networks each take state-action pairs as input and output value estimates. During training, the minimum of the two Q values ​​is used to update the actor network, effectively reducing overestimation bias. Before formal deployment, the policy undergoes an offline pre-training phase. Pre-training is conducted in a simulation environment constructed using historical operating data. This simulation environment approximates the dynamic response of the aerobic tank using historical water quality data and a process model, and incorporates equipment performance degradation characteristics into the simulation environment. This ensures that the pre-trained policy has a certain adaptability to equipment aging scenarios before deployment. Pre-training continues until the policy can stably satisfy constraints in the simulation environment and the energy consumption optimization effect reaches the convergence condition. During the online operation phase after deployment, the strategy network directly outputs the aeration allocation ratio of each zone based on the current state vector at each control time step. The action output is deterministic inference (taking the mean of the actor network output distribution) without random sampling to ensure the stability of the decision during deployment. The final output of the constraint decision module is the target absolute air volume of each zone at the current control time step. This result serves as the direct input for the performance compensation and coupling compensation calculations of the execution instruction generation module.

[0048] The execution instruction generation module divides the target air volume of each zone by the real-time efficiency coefficient to obtain the compensated target air flow rate. It then uses the interval airflow coupling compensation matrix identified by historical data to inversely solve the valve opening instructions for each zone and converts the blower frequency instructions into the total compensated air flow rate. The constraint decision module outputs the ideal allocation values ​​of the target air volume for each zone. However, in an actual aerobic tank aeration system, to accurately convert this set of ideal target air volumes into executable physical control commands, two key engineering issues need to be addressed sequentially: First, there is the issue of equipment efficiency compensation. Since the aging levels of equipment in each zone vary, the actual oxygen transfer generated by the same target air volume differs across zones, requiring compensation for the target air volume based on the efficiency coefficient of each zone. Second, there is the issue of inter-zone airflow coupling interference. Multiple zones share a main air supply network, and adjusting the opening of a damper in one zone will affect the actual airflow in other zones through changes in network pressure. This mutual interference must be eliminated using a coupling compensation matrix. These two issues are addressed in the order of efficiency compensation first, followed by coupling compensation. The result of efficiency compensation is the actual operational target for all subsequent calculations.

[0049] Regarding equipment performance compensation, the performance identification module identifies the real-time performance coefficients of each zone ( This reflects the degree of decline in the current oxygen transfer capacity of each zone's equipment relative to its nominal value. The target airflow for each zone, output by the constraint decision module, is divided by the corresponding zone's efficiency coefficient to obtain the "compensated target airflow" for each zone. The physical meaning of this is: if the oxygen transfer efficiency of a zone's equipment has decreased to 80% of its nominal value (i.e., η=0.8), then 1.25 times the target airflow needs to be actually delivered to that zone to provide the water with the target dissolved oxygen level, compensating for the oxygen transfer loss caused by equipment aging.

[0050] Regarding inter-zone airflow coupling compensation, in a structure where multiple zones share a single main air supply network, there is a hydrodynamic coupling between the motorized dampers of each zone and the main network. When the damper opening of a zone increases, the airflow resistance in that zone decreases, and the airflow into that zone increases accordingly. However, this also slightly decreases the gas pressure in the main network, resulting in a reduction in the actual airflow received by other zones when the damper opening remains unchanged, and vice versa. If the airflow control of each zone is handled independently, ignoring inter-zone coupling, and the damper opening of each zone is directly set according to the compensated target airflow, the actual airflow of each zone will ultimately deviate from the target value in a correlated manner, which is difficult to eliminate through simple single-zone correction.

[0051] This module uses a coupling compensation matrix C to systematically compensate for the interval coupling effect. The coupling compensation matrix C is a 3x3 matrix. The element in the i-th row and j-th column describes the influence coefficient of the change in the opening of the damper in the j-th zone on the actual airflow in the i-th zone: diagonal elements are positive (reflecting the direct influence of the damper opening on the airflow in this zone), and off-diagonal elements are negative (reflecting the disturbance influence of damper adjustments in other zones on the airflow in this zone). The identification of matrix C is achieved through multivariate regression analysis of historical operating data: Operating segments with relatively stable blower frequencies are selected from the historical data, with the change in the opening of the dampers in each zone as the independent variable and the change in the measured airflow in each zone as the dependent variable. The least squares method is used to fit the estimated values ​​of each element of matrix C. During identification, two operating conditions, high frequency and low frequency, are distinguished for the blower, and corresponding coupling matrix versions are established for each. The corresponding version is selected based on the current blower frequency during real-time compensation calculations. The airflow-valve opening calibration curve established by the step aeration test in the identification and equipment calibration module provides a basic correspondence between airflow and opening for the initial identification of the diagonal elements of matrix C when each zone operates independently. At the same time, it provides a reference for verifying the reasonable range of valve opening when solving the matrix equation.

[0052] In actual compensation calculations, the "compensated target airflow" vector for each zone is used as the solution objective. A system of linear equations is established using the current coupling matrix C. Inverse solving enables the allocation of the damper opening settings for each zone to the compensated target airflow under the condition of coupling effect. For cases where the matrix equations have multiple solutions, the system selects a solution based on the secondary optimization principle of "minimizing the total damper adjustment," reducing unnecessary frequent and large-scale damper adjustments and extending the lifespan of the actuators. After each control cycle is completed, the deviation between the compensated target airflow and the measured actual airflow for each zone in that step is used as the identification residual. The matrix C is then incrementally corrected online using a recursive least squares algorithm, allowing the matrix to slowly and adaptively update with changes in system operating conditions.

[0053] Regarding the generation of blower frequency commands, the sum of the "compensated target airflow" for each zone determines the total airflow required for the current control step. The system calculates the minimum frequency setpoint that can provide this total airflow using the blower unit's performance curve (frequency-flow relationship, established by calibrating the blower operating data recorded during the stepped aeration test by the identification and equipment calibration module). A rate limit is simultaneously applied to the frequency change (the change in each adjustment does not exceed the preset step size limit) to prevent large frequency jumps from causing mechanical shock to the blower. When the total demand airflow after compensation exceeds the current measured total airflow referenced by the constraint decision module, the command generation module calculates the frequency command based on the actual total demand after compensation and actively requests the blower to provide more total airflow. The measured total airflow read by the constraint decision module in the next control cycle will be automatically updated to the new actual value. The strategy network gradually learns to internalize the efficiency compensation demand into the allocation decision through online iteration of the online iteration module, and the two-level control tends to be coordinated through periodic iteration.

[0054] The control commands output by this module include: the precise target opening degree of each zone's electric air valve (expressed as a percentage, accurate to one decimal place) and the set value of the blower unit's operating frequency (in Hertz, accurate to one decimal place). These commands are sent from the host computer to the field PLC via the industrial Ethernet protocol. After the PLC confirms the receipt of the commands, it immediately drives the corresponding frequency converter and air valve actuator to perform the adjustment. The time delay of the entire sending-execution link is controlled within 1 second.

[0055] The online iterative module updates the reinforcement learning strategy with priority experience playback according to the control cycle, and fine-tunes the oxygen demand prediction model daily. The first five modules collectively complete a single-cycle control loop from perception to execution. However, with long-term system operation, the influent water quality characteristics will slowly drift due to factors such as seasonal changes, altered drainage patterns, and changes in external industrial sources; the efficiency of aeration equipment will continuously decline; and the microbial community structure of the activated sludge in the aerobic zone will evolve over time, leading to changes in the OUR characteristics under the same influent conditions. If the model parameters of the control system remain fixed at the initial training state, its ability to adapt to changes in the actual operating environment will gradually decrease over time. The task of this module is to drive all data-driven models within the system to continuously learn the latest operating experience through an online learning mechanism, maintaining long-term stability of control performance.

[0056] The execution of this module is divided into two independent time granularity levels: a fast loop with each 5-minute control cycle and a slow loop with each 24-hour cycle. The fast loop is triggered at the end of each control cycle, completing three operations: reward evaluation, experience storage, and online update of the policy network. The system reads the stable response values ​​of each partition's DO within 1 to 2 minutes after the execution of the command issued by the execution command generation module, as well as the actual operating power record of the blower in this time step, and calculates the immediate reward: the main reward is the negative value of the actual blower energy consumption in this time step (normalized based on rated power), and the intermediate auxiliary reward is the negative value of the sum of the absolute values ​​of the deviations of the actual values ​​of each partition's DO from the ideal DO operating point defined by the constraint decision module (normalized based on the difference between the saturated DO and the DO safety lower limit in each partition), and the two are weighted and summed by λ; the constraint cost signal is a 3-dimensional Boolean vector ( ), Indicates the first Partition DO is below the safe lower limit. This indicates that the constraint is met. When the detection results (COD and ammonia nitrogen) of the effluent water quality sensor are readable (usually refreshed every 30 to 40 minutes), the degree of deviation of the effluent water quality from the target is converted into a delayed reward signal, which is retroactively superimposed on the corresponding historical time step reward in a way with a discount coefficient, so that the strategy can obtain correction signals from long-term effluent water quality feedback.

[0057] The system maintains a fixed-capacity priority experience replay buffer, with each storage unit recording a complete 5-tuple: current state vector, executed action vector, immediate reward value, next step state vector, and constraint cost signal (i.e., ...). (3D Boolean vector). When the buffer pool capacity reaches its limit, the oldest sample is eliminated using a first-in, first-out (FIFO) strategy to maintain the timeliness of the data in the buffer pool. The sampling priority of each experience in the buffer pool is determined based on the absolute value of the temporal difference (TD) error: the larger the absolute value of the TD error, the more "unexpected" information the sample contains for the current policy, and the higher the priority; recent time step experience samples are additionally weighted by a timeliness weighting coefficient, making the policy pay slightly more attention to the latest running situation than to older historical samples. Whenever a sufficient number of new samples accumulate in the buffer pool, the system samples a small batch of experience from the buffer pool according to priority, and performs a small-step gradient update on the SAC policy network and the constraint cost estimation network; the Lagrange multipliers of each of the three partitions ( Independent dual updates are performed based on the mean of the constraint cost signal of the corresponding partition in this batch: if the proportion of constraint violation samples is high, the multiplier is increased; if the constraint is sufficiently satisfied, the multiplier is decreased. Each partition is adjusted independently to maintain a dynamic balance between constraint satisfaction and energy consumption optimization.

[0058] The slow loop executes independently every 24 hours to adaptively update the LSTM prediction model. The system aggregates all OUR data and corresponding operational feature data obtained from daily detections and merges them with the existing training dataset. This allows the LSTM network to undergo a round of fine-grained training with a very small learning rate, enabling the network to gradually perceive the slow evolution of operational patterns. During fine-tuning, the system continues to use the previously trained LSTM version for prediction. After fine-tuning, it switches to the new version in the next adjustment cycle, ensuring uninterrupted prediction service during the switch. When the system detects that the root mean square error (RMS) of the OUR prediction (assessed by the deviation between the actual detection and predicted values) of the LSTM exceeds the third prediction error threshold (the default value is twice the historical average prediction error at the initial deployment stage) over several recent adjustment cycles, a priority fine-tuning process is triggered. Without waiting for the 24-hour cycle, a targeted update is immediately performed using the latest data to quickly correct the prediction accuracy degradation issue.

[0059] The parameter updates for the equipment performance identification model are already being performed online in a cycle-by-cycle manner within the performance identification module, and will not be triggered additionally in this module. At this point, after the fast and slow cycles have completed, all data-driven models within the system have been updated to the latest state. When the next adjustment time step arrives, the complete process from the identification and equipment calibration module to the online iteration module will be executed again in a loop.

[0060] In one embodiment of the present invention, the aerobic tank of a wastewater treatment plant in a northern city is used as a specific application scenario. Example data is calculated based on the plant's process parameters and actual operating patterns to reflect the system response characteristics under typical operating conditions. The aerobic tank of this plant has three aeration zones, and the aeration discs have been in operation for approximately 5 years. The peak influent flow period is from 7:00 AM to 9:00 AM daily, during which the influent COD significantly increases compared to the nighttime baseline. Consequently, the OUR (or osmotic pressure) of the activated sludge in the aerobic zone rises rapidly when high-load influent arrives. Taking the morning peak period (07:10 AM to 08:10 AM) of a weekday as an example, the system control process is as follows: At 07:10, the identification and equipment calibration module completed the second round of detection and identification for the day (the previous round was completed at 06:30). The DO in each zone was above the detection safety threshold. The three zones completed intermittent active detection in sequence, identifying the current oxygen consumption rate of each zone. Compared with the identification value at 06:30, the OUR of zone 1 (the front section of the aerobic tank, which first receives high COD influent) has increased noticeably, while the OUR of zones 2 and 3 is still close to the baseline level.

[0061] The LSTM of the oxygen demand prediction module takes the currently updated OUR time series (including the 07:10 identification value) and the current influent COD sensor reading (a COD upward trend has been detected) as inputs to predict that the oxygen demand of each zone will continue to increase in the next 30 minutes (6 control steps). It predicts that the increase in zone 1 will be significantly greater than that in zones 2 and 3. Zone 3 is expected to see an increase in oxygen demand in about 25 minutes (after the high COD substances in the influent are treated by zones 1 and 2, some residual organic matter continues to be consumed in zone 3).

[0062] The efficiency identification module performs online back-calculation of efficiency based on the current OUR identification value and the latest aeration execution feedback. The efficiency coefficients of Zone 1 and Zone 2 are both below 1 (representing typical differences in aging levels), while the efficiency coefficient of Zone 3 is relatively close to 1. The efficiency coefficients of all three have been incorporated into the state vector of the constraint decision module. Based on the current state (including DO, oxygen demand prediction vector, efficiency coefficient, etc. for each zone), the constraint decision module outputs the optimal aeration allocation scheme for each zone in this control step: because a high-load influent impact is predicted for Zone 1, the strategy pre-allocates more aeration to Zone 1, while increasing the total air volume to cope with the overall increase in oxygen demand, and maintaining appropriate air supply to Zones 2 and 3 to keep DO within a safe range. The execution instruction generation module first compensates for the target air volume of each zone according to the efficiency coefficient, and then applies coupled compensation correction to the larger incremental air volume instruction of Zone 1 (the impact of the pressure drop in Zones 2 and 3 caused by the large opening of the valve in Zone 1 is compensated through matrix correction). Finally, it outputs the valve opening instruction and the corresponding blower frequency instruction for each zone, which are then issued and executed by the PLC.

[0063] As shown in Tables 1 and 2, examples of sensor data acquisition and corresponding control outputs of the system during this period are presented.

[0064] Table 1. Examples of sensor data collection in each zone of the aerobic tank (07:10-08:10)

[0065] Table 2, Example of system control decision output (07:10—08:10)

[0066] Note: The OUR value at 07:10 is the measured result of this round of intermittent detection and identification; the OUR value from 07:22 to 07:58 is the LSTM prediction value, and the next round of complete detection and identification will be carried out between approximately 07:40 and 08:10.

[0067] As shown in Table 2, the CMDP decision strategy of the constraint decision module detected the continuously rising trend of influent COD at 07:22 (relying on the LSTM prediction look-ahead of the oxygen demand prediction module), and increased the aeration allocation of zone 1 in advance. The blower frequency rose to its peak (49.8 Hz) at 07:34. At 07:34, the DO of each zone remained within the safe range (as shown in Table 1, the DO of zone 1 at 07:34 was 2.1 mg / L, which did not fall below the safe lower limit). After 07:46, as the influent COD decreased, the blower frequency and the aeration of zone 1 were synchronously reduced, avoiding energy waste caused by inertial over-aeration. The online iteration module incorporated the complete experience of this influent shock event into the buffer pool, further strengthening the strategy's forward-looking response capability to the morning peak influent shock pattern.

[0068] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. An AI-based dynamic adaptive control system for aeration volume in an aerobic tank of wastewater treatment plant, characterized in that: include: The identification and equipment calibration module collects dissolved oxygen data for each zone, intermittently reduces aeration when dissolved oxygen is higher than the safety threshold, and linearly fits the dissolved oxygen decay time sequence to identify the oxygen consumption rate time sequence of each zone. At the same time, it establishes the oxygen transfer efficiency calibration curve and air flow-valve opening calibration curve for each zone through stepped aeration test. The oxygen demand prediction module outputs the future N-step oxygen demand prediction matrix for each zone through a long short-term memory network, based on the oxygen consumption rate time series and influent load characteristics. The efficiency identification module takes the predicted value of the prediction matrix at the current time step and the measured change in dissolved oxygen, and calculates the real-time efficiency coefficient of each partition using the oxygen mass balance equation. The constrained decision module, with the prediction matrix, the real-time efficiency coefficient of each zone and the measured total air volume of the blower as the state, outputs the target air volume of each zone under the constrained Markov decision framework through the Lagrange constrained soft actor-critic algorithm, which satisfies the lower limit constraint of dissolved oxygen safety in each zone. The execution instruction generation module divides the target air volume of each zone by the real-time efficiency coefficient to obtain the compensated target air flow rate. It then uses the interval airflow coupling compensation matrix identified by historical data to inversely solve the valve opening instructions for each zone and converts the blower frequency instructions into the total compensated air flow rate. The online iterative module updates the reinforcement learning strategy according to the priority of experience playback based on the regulation cycle, and fine-tunes the oxygen demand prediction model daily.

2. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, Before triggering detection, the OUR identification and equipment calibration module calculates the safety trigger threshold for each zone by summing the product of the dissolved oxygen safety lower limit, the oxygen consumption rate of the last identification, and the longest detection duration. When the dissolved oxygen in each zone is higher than the safety trigger threshold, the detection is initiated. The system performs detection in turn according to the detection scheduling queue of each partition, and at most one partition can be in the detection state at any one time.

3. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, When the OUR identification and equipment calibration module performs detection, it reduces the opening of the target zone air valve to the preset minimum opening, calculates the residual oxygen transfer rate during the detection period using the effective oxygen transfer benchmark model, and subtracts the residual oxygen transfer rate from the oxygen consumption rate obtained by linear fitting of dissolved oxygen decay time sequence to obtain the corrected oxygen consumption rate identification value. When conducting a stepped aeration test, the system is considered to be in a steady state when the dissolved oxygen level stabilizes at each setting until the variation in the continuous sampling points is lower than the preset steady-state threshold.

4. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The input feature vector of the long short-term memory network of the oxygen demand prediction module consists of the historical time series of oxygen consumption rate of each zone, the time series of influent water quality and flow rate, and the feature encoding of water temperature and time period.

5. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The real-time efficiency coefficient calculation steps of the efficiency identification module are as follows: approximate the average dissolved oxygen of the time step by the average value of the measured dissolved oxygen at the start and end of the time step; query the oxygen transfer coefficient corresponding to the current valve opening by the oxygen transfer efficiency calibration curve, and obtain the nominal oxygen transfer increment by multiplying the difference between saturated dissolved oxygen and average dissolved oxygen by the time step length. The actual oxygen transfer increment is obtained by adding the measured change in dissolved oxygen at each time step to the predicted value of the current time step of the oxygen demand prediction matrix and the time step length. The real-time efficiency coefficient is the ratio of the actual oxygen transfer increment to the nominal oxygen transfer increment. When the nominal oxygen transfer increment is lower than the preset minimum calculation threshold, the previous time step efficiency coefficient estimate is maintained.

6. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The efficiency identification module adopts a dual-mode strategy for updating efficiency coefficients: when the water intake condition is stable, the calculation result of the current step is included in the weighted moving average update with the first update weight; When the change in influent water quality relative to the moving average exceeds the preset change judgment threshold within a preset time window, the system switches to the second update weight for gradual updates until the operating conditions return to stability; the first update weight is greater than the second update weight.

7. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The reward function of the constraint decision module is composed of the main energy consumption reward item and the auxiliary dissolved oxygen deviation reward item, weighted by preset weight coefficients: the main reward item is the negative value of the actual energy consumption of the blower in the current step, normalized based on the rated power; The auxiliary reward item is the negative value of the deviation of the measured dissolved oxygen in each region from the ideal dissolved oxygen operating point, normalized based on the difference between saturated dissolved oxygen and the lower limit of dissolved oxygen; the ideal dissolved oxygen operating point in each region is calculated and determined based on the current saturated dissolved oxygen, the current time step predicted value of the oxygen demand prediction matrix, and the real-time efficiency coefficient.

8. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The constraint decision module establishes independent Lagrange multipliers for each aeration zone of the aerobic tank: when the dissolved oxygen in a certain zone is lower than the safe lower limit of dissolved oxygen, the corresponding zone multiplier is automatically increased to enhance the penalty weight for constraint violation, and the corresponding multiplier is gradually decreased when the constraint is satisfied. Each partition multiplier is adjusted independently during each policy update using the dual gradient descent method, and the partition multipliers do not interfere with each other.

9. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, In the instruction generation module, the interval airflow coupling compensation matrix is ​​identified by multivariate regression of historical operating data, and the elements of the matrix are fitted with the change in the opening of the damper in each zone as the independent variable and the change in the measured airflow in each zone as the dependent variable. Establish corresponding coupling matrices for the first and second operating frequency ranges of the blower, and select the corresponding version according to the current blower frequency during real-time compensation; use the deviation between the compensated target air flow and the measured air flow as the residual for each control cycle, and perform online correction of the matrix increment by recursive least squares method.

10. The AI ​​control system for dynamically adaptive aeration volume in the aerobic tank of wastewater treatment according to claim 1, characterized in that, The online iterative module operates on a fast cycle based on the regulation period: it calculates immediate rewards based on the actual energy consumption and measured dissolved oxygen values ​​of each zone, generates constraint cost signals based on whether the dissolved oxygen in each zone is below the dissolved oxygen safety limit, samples from the experience buffer pool with priority based on the absolute value of the time-series difference error to perform gradient updates on the policy network, and independently adjusts the corresponding Lagrange multipliers according to the constraint cost signals of each zone; and performs fine-grained fine-tuning of the oxygen demand prediction model on a daily basis.