Intelligent regulation and control system and method for sewage treatment process based on multi-time scale reinforcement learning

The intelligent control system for wastewater treatment processes, which utilizes multi-timescale reinforcement learning, solves the problems of insufficient control accuracy and adaptability of wastewater treatment systems under complex operating conditions. It achieves efficient, stable, and energy-saving operation of wastewater treatment systems and is applicable to various processes such as SBR, A²/O, and MBR.

CN121635170APending Publication Date: 2026-03-10ZHEJIANG YUTENG BAINUO ENVIRONMENTAL PROTECTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing wastewater treatment systems lack sufficient control precision when facing complex operating conditions, making it difficult to simultaneously ensure stable effluent quality and energy conservation. Furthermore, existing technical solutions have weak adaptability and cannot meet the control requirements across multiple time scales.

Method used

An intelligent control system for wastewater treatment processes based on multi-timescale reinforcement learning is adopted, including modules for data acquisition, data preprocessing, short-timescale control, long-timescale configuration, process adaptation, and multi-timescale fusion control. Combined with reinforcement learning algorithms such as deep Q-networks and policy gradient methods, it can achieve multi-dimensional optimization and dynamic adjustment of wastewater treatment processes.

Benefits of technology

It achieves rapid response and precise control to fluctuations in water quality and quantity, significantly reduces energy and chemical consumption, improves effluent compliance rate, and has high efficiency, adaptability, and intelligence, making it suitable for various wastewater treatment processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121635170A_ABST
    Figure CN121635170A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of environmental engineering and intelligent control, in particular to a sewage treatment process intelligent regulation and control system and method based on multi-time scale reinforcement learning, and the method comprises the steps: collecting real-time operation parameters in a sewage treatment process, and obtaining a cleaning data set through anomaly detection and filtering processing; by taking a minute-hour level as a regulation and control cycle, optimizing the short-term operation parameters in real time to obtain short-term optimization parameters; based on the cleaning data set, the historical operation database and a preset evaluation index, optimizing long-term configuration parameters by taking a day-week level as a period, and considering a working condition change rule and a long-term operation result; dynamically adjusting the parameter dimension of the short / long-term regulation and control module according to the process type, combining the real-time working condition adaptive weight coefficient, and fusing the two types of optimization parameters to generate a comprehensive regulation and control instruction; and driving an execution mechanism in the sewage treatment process to act through the comprehensive regulation and control instruction. The system can adapt to various sewage treatment processes in a self-adaptive manner, and gives consideration to real-time parameter optimization response and long-term operation configuration stability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of environmental engineering and intelligent control technology, in particular to a sewage treatment process intelligent regulation and control system and method based on multi-time scale reinforcement learning. BACKGROUND

[0002] In the sewage treatment industry, stable and efficient operation of sewage treatment plants is the core link to achieve sewage discharge standards and ensure water environment safety. During daily operation, real-time response to dynamic changes in water quality, water quantity and external environmental conditions is required, and continuous and high-frequency regulation and control of key process parameters such as aeration time, dissolved oxygen concentration, sludge age, and reagent dosage is implemented. The accuracy and timeliness of this regulation and control process directly determine the water quality standard compliance rate and the operation energy consumption level of the sewage treatment system.

[0003] Currently, the regulation and control of process parameters of most sewage treatment systems still relies on the experience of operation and maintenance personnel, lacking systematic regulation and control logic, standardized operation procedures, and self-adaptive dynamic adjustment capabilities. This traditional regulation and control mode has significant limitations in practical application: when facing complex conditions such as large fluctuations in influent water quality and frequent changes in pollutant load, parameter adjustment may lag, control accuracy may be insufficient, and operation efficiency may be low, resulting in difficulties in simultaneously achieving both water quality stability and standard compliance and energy saving and consumption reduction in operation, which may lead to water quality exceeding standards due to untimely regulation and control, or excessive consumption of energy and reagents due to unreasonable parameter matching.

[0004] To break through the bottleneck of traditional experience-based regulation and control and improve the intelligent operation level of sewage treatment systems, some research fields have attempted to introduce model predictive control (MPC), expert rule library and other technical methods to optimize the regulation and control strategy of process parameters. However, existing technical solutions still have obvious shortcomings: on the one hand, methods such as model predictive control (MPC) and expert rule library have weak adaptability to complex conditions and are difficult to cope with sudden and large-scale changes in water quality and quantity; on the other hand, their optimization period is mostly a single fixed scale, which cannot match the differentiated regulation and control needs of different process links (such as short-term water quality fluctuation response and long-term system stability maintenance), and the regulation and control targets are mostly focused on single dimensions of water quality or energy consumption, making it difficult to fully balance the multi-objective coordination relationship between energy consumption cost, system operation stability and water quality guarantee.

[0005] Based on the above industry status and technical pain points, there is an urgent need in the current sewage treatment field to develop an intelligent regulation and control system with multi-time scale learning capability, which can dynamically adapt to regulation and control needs under different conditions, cooperatively optimize multi-dimensional operation targets, solve the problems of poor adaptability and single optimization dimension of existing regulation and control methods, and ultimately achieve efficient, stable and energy-saving operation of sewage treatment systems. SUMMARY

[0006] To this end, the technical problem to be solved by the present application is to overcome the problems of the existing sewage treatment regulation system, such as insufficient self-adaptability, limited intelligent level, poor running stability, and the like, so as to provide a sewage treatment process intelligent regulation system and method based on multi-time scale reinforcement learning.

[0007] Specifically, the sewage treatment process intelligent regulation system based on multi-time scale reinforcement learning comprises a data acquisition module, a data preprocessing module, a short-time scale regulation module, a long-time scale configuration module, a process adaptation module, a multi-time scale fusion regulation module, and an execution module. The data acquisition module is configured to acquire real-time operation parameters in a sewage treatment process. The data preprocessing module is configured to perform anomaly detection and filtering processing on the real-time operation parameters to form a cleaning data set that can be used for model training and decision optimization. The short-time scale regulation module is configured to set a regulation period to a minute level to an hour level, perform real-time optimization control on short-term operation parameters in the cleaning data set, and obtain short-term optimization parameters. The long-time scale configuration module is configured to set a regulation period to a day level to a week level based on long-term operation parameters in the cleaning data set and a preset historical operation database, combine a preset evaluation index, perform periodic optimization on long-term configuration parameters of a sewage treatment process, consider a working condition change rule and a long-term operation result in the historical operation data in the optimization process, and obtain long-term configuration optimization parameters. The process adaptation module is configured to dynamically adjust a dimension of action parameters of the short-time scale regulation module and a dimension of state parameters of the long-time scale configuration module according to a type of a target sewage treatment process. The multi-time scale fusion regulation module is configured to dynamically adjust a weight coefficient according to a real-time working condition, fuse the short-term optimization parameters and the long-term configuration optimization parameters, and generate a comprehensive regulation instruction. The execution module is configured to drive an execution mechanism in a sewage treatment process to act through the comprehensive regulation instruction.

[0008] In an embodiment of the present application, the system further comprises a policy updating module configured to acquire regulation result data of the execution module, calculate reward feedback based on a preset reward function, update a policy network of the short-time scale regulation module and the long-time scale configuration module using an experience replay mechanism and a target network updating mechanism, and the reward function comprises a weighted item of at least one evaluation index.

[0009] In an embodiment of the present application, the short-time scale regulation module is configured to set the regulation period to the minute level to the hour level, perform real-time optimization control on the short-term operation parameters in the cleaning data set, and obtain short-term optimization parameters, including: Based on the regulation period, real-time operation parameters in a current observation window are collected at a preset time interval through a sliding window, and historical operation data of a preset length before the observation window are synchronously extracted; The collected real-time-historical joint data sequence is subjected to time series data preprocessing, and is converted into a high-dimensional feature vector that can be analyzed by a model; A short-time scale regulation model is constructed based on a reinforcement learning algorithm, the high-dimensional feature vector is taken as a model input, the short-time scale regulation model is iteratively trained through a first reward function, the training process takes the maximization of the reward function value as an optimization objective, and the short-time scale regulation model is converged until the short-time scale regulation model is finally generated.

[0010] In an embodiment of the present application, the first reward function is as follows: , wherein, is a water quality compliance rate in the current observation window, is a cumulative energy consumption in the current observation window, is a cumulative chemical consumption in the current observation window, , , is a weight coefficient.

[0011] In an embodiment of the present application, the short-time scale regulation model includes one or a combination of DQN, PG, and Actor-Critic reinforcement learning models, and is selected according to the process response complexity and the regulation target dimension.

[0012] In an embodiment of the present application, the long-time scale configuration module is configured to set the regulation period to the day level to the week level based on the long-term operation parameters in the cleaning data set and a preset historical operation database, combine a preset evaluation index, periodically optimize long-term configuration parameters of the sewage treatment process, consider the working condition change law and the long-term operation result in the historical operation data in the optimization process, and obtain long-term configuration optimization parameters, including:

[0013] Based on the preprocessed long-term operation parameters, system stability index values , COD removal rates , and operation costs are sequentially calculated; the system stability index values , the COD removal rates , and the operation costs , a second reward function is constructed as follows: wherein , , is a weight coefficient; A long-time scale configuration model is constructed, the long-time scale configuration model is iteratively trained through the second reward function, a training process takes maximizing a reward function value as an optimization objective, until the long-time scale configuration model converges, and finally a long-term configuration optimization parameter is generated.

[0014] In an embodiment of the present application, the system stability index value is calculated by using a weighted scoring method, and a formula is as follows: wherein represents a fluctuation degree of effluent indexes, and a calculation formula is as follows , represents a maximum deviation value of effluent COD in a set time unit; represents an upper limit value of effluent COD standard; represents a device fault-free operation rate, and a calculation formula is as follows , is a cumulative operation time length of the device in the set time unit; represents a cumulative fault time length of the device in the set time unit; represents a value assignment according to a process parameter adjustment frequency; , , is a weight coefficient.

[0015] In an embodiment of the present application, the calculation formula of the COD removal rate is as follows: , represents a total influent amount, represents an influent COD average value; represents a total effluent amount, represents an effluent COD average value.

[0016] In an embodiment of the present application, the calculation formula of the operation cost is as follows:

[0017] , is a cumulative energy consumption cost in the set time unit, represents a cumulative chemical consumption cost in the set time unit, represents a cumulative sludge treatment cost in the set time unit, represents a device maintenance cost in the set time unit.

[0018] Based on the same inventive concept, the application also provides a sewage treatment process intelligent regulation and control method based on multi-time scale reinforcement learning, comprising the following steps:

[0019] Real-time operation parameters in the sewage treatment process are collected, and the real-time operation parameters are subjected to anomaly detection and filtering processing to form a cleaning data set that can be used for model training and decision optimization;

[0020] The regulation and control period is set to be minute-level to hour-level, and the short-term operation parameters in the cleaning data set are subjected to real-time optimization control to obtain short-term optimization parameters;

[0021] Based on the long-term operation parameters in the cleaning data set and a preset historical operation database, and in combination with a preset evaluation index, the regulation and control period is set to be day-level to week-level, and the long-term configuration parameters of the sewage treatment process are subjected to periodic optimization, and in the optimization process, the working condition change law and long-term operation result in the historical operation data are considered to obtain long-term configuration optimization parameters;

[0022] According to the type of the target sewage treatment process, the action parameter dimension of the short-time scale regulation and control module and the state parameter dimension of the long-time scale configuration module are dynamically adjusted;

[0023] The weight coefficient is dynamically adjusted according to the real-time working condition, and the short-term optimization parameters and the long-term configuration optimization parameters are fused to generate comprehensive regulation and control instructions;

[0024] Through the comprehensive regulation and control instructions, the actuator in the sewage treatment process is driven to act.

[0025] The above technical scheme of the application has the following beneficial effects compared with the prior art:

[0026] The application fuses short-time scale (minute to hour level) real-time parameter optimization and long-time scale (day or week level) process configuration adjustment, relies on deep Q network, policy gradient method and other reinforcement learning algorithms and dynamic weighting fusion mechanism, realizes rapid response and accurate regulation and control of water quality and quantity fluctuations, effectively replaces artificial experience decision, avoids regulation lag and over-standard risk, balances global stability and long-term energy efficiency, adapts to SBR, A² / O, MBR and other processes, greatly shortens the deployment period with the aid of policy transfer and self-adaptive adjustment capability, continuously optimizes parameters such as aeration amount and dosing amount through a closed-loop feedback and policy updating mechanism, significantly reduces energy consumption and chemical consumption, improves water standard compliance rate, has intelligentization, high adaptability, synergistic optimization effect, wide practicality and long-term evolution potential, outstanding economic benefits and engineering promotion value. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to make the content of the present application more easily understood, the present application will be further described in detail below according to specific embodiments of the present application and in conjunction with the accompanying drawings.

[0028] Figure 1 is a hierarchical architecture diagram of a sewage treatment process intelligent regulation and control system based on multi-time scale reinforcement learning provided in an embodiment of the present application;

[0029] Figure 2 is a multi-time scale reinforcement learning strategy topology diagram provided in an embodiment of the present application;

[0030] Figure 3 is a parameter optimization and process configuration iteration flowchart provided in an embodiment of the present application.

[0031] The description of the figures of the drawings is as follows: 10, data acquisition module; 20, data preprocessing module; 30, short time scale regulation and control module; 40, long time scale configuration module; 50, process adaptation module; 60, multi-time scale fusion regulation and control module; 70, execution module; 80, strategy updating module. DETAILED DESCRIPTION

[0032] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present application and implement it, but the embodiments are not limiting to the present application.

[0033] Embodiment one: Referring to Figures 1-3 , the present application provides a sewage treatment process intelligent regulation and control system based on multi-time scale reinforcement learning, which follows the closed-loop control logic of "perception-decision-execution-feedback", and constructs a three-level architecture of data acquisition layer, learning decision layer and execution feedback layer, specifically including: data acquisition module 10, data preprocessing module 20, short time scale regulation and control module 30, long time scale configuration module 40, process adaptation module 50, multi-time scale fusion regulation and control module 60, execution module 70 and strategy updating module 80, to realize intelligent and self-adaptive regulation and control of the whole process of sewage treatment.

[0034] The data acquisition module 10 is used to acquire real-time operation parameters in the sewage treatment process; The data preprocessing module 20 is used to perform anomaly detection and filtering processing on the real-time operation parameters to form a cleaning data set that can be used for model training and decision optimization; The short time scale regulation and control module 30 is used to set the regulation and control period to minutes to hours, to perform real-time optimization control on the short-term operation parameters in the cleaning data set to obtain short-term optimization parameters; The long-term scale configuration module 40 is used to periodically optimize the long-term configuration parameters of the sewage treatment process based on the long-term operating parameters in the cleaning dataset and the preset historical operating database, combined with preset evaluation indicators, and set the control cycle to daily to weekly. In the optimization process, the operating condition change patterns in the historical operating data and the long-term operating results are considered to obtain the long-term configuration optimization parameters. The process adaptation module 50 is used to dynamically adjust the action parameter dimension of the short time scale control module and the state parameter dimension of the long time scale configuration module according to the type of the target wastewater treatment process. The multi-timescale fusion control module 60 is used to dynamically adjust the weight coefficients according to real-time operating conditions and to generate a comprehensive control command by fusing the short-term optimization parameters and the long-term configuration optimization parameters. The execution module 70 is used to drive the actuators in the wastewater treatment process to operate through the integrated control commands; The strategy update module 80 is used to obtain the control result data of the execution module, calculate the reward feedback based on the preset reward function, the reward function includes a weighted term of at least one evaluation index, and update the strategy network of the short time scale control module 30 and the long time scale configuration module 40 using an experience replay mechanism and a target network update mechanism.

[0035] Furthermore, the data acquisition module 10, as the core unit of the system's perception layer, deploys multi-dimensional sensors at key process nodes such as the inlet, anaerobic tank, aerobic tank, and outlet of the wastewater treatment system. These sensors collect key parameters in real time, including influent and effluent water quality (including COD, NH3-N, DO, pH, and sludge concentration), influent and effluent flow rates, equipment energy consumption (power of aeration and dosing equipment), and equipment operating status (fault duration and cumulative operating time). The data is collected and uploaded to the edge server at a frequency of once per minute by a PLC programmable logic controller, achieving full coverage and real-time perception of process data across the entire process scenario.

[0036] The data preprocessing module 20 receives the output data from the data acquisition module 10 and performs a standardized preprocessing process: First, outliers are identified through threshold judgment and trend analysis. A sudden increase in COD to 2000 mg / L, or DO < 0.5 mg / L or > 8 mg / L, is judged as a sensor malfunction. A difference between adjacent data exceeding 30% and occurring more than three times consecutively is judged as a fluctuation anomaly. The outlier data is completed using linear interpolation. Then, the state sequence of the time series data is reconstructed using a 10-minute sliding window. The data is mapped to the [0,1] interval through Min-Max normalization. Finally, a clean dataset that integrates real-time and historical data is output, providing high-quality data support for subsequent model training and decision optimization.

[0037] Subsequently, the system synchronously starts the short-timescale control module 30 and the long-timescale configuration module 40 in the learning decision layer for parallel processing. The short-timescale control module 30 uses a control cycle of minutes to hours, collects real-time operating parameters within the current observation window through an hourly sliding window, and extracts historical operating data for a preset duration (the first 30 minutes to 1 hour) before the window, and fuses them to generate a high-dimensional feature vector containing real-time status, short-term trends and equipment operating characteristics;

[0038] A short-timescale control model is constructed based on reinforcement learning algorithms. The model adaptively selects one or a combination of DQN (Deep Q-Network), PG (Policy Gradient Method), or Actor-Critic architectures according to the complexity of the process response and the dimensionality of the control target. Specifically: If the process response is simple and the state space is small, such as short-term DO concentration control in a single aeration tank, then a Deep Q-Network (DQN) is selected. This algorithm relies on an experience replay mechanism and a target network update strategy to achieve rapid convergence during the offline training phase and efficiently optimize short-term discrete control strategies. For multi-objective coupled optimization scenarios, and applicable to complex decision-making problems in high-dimensional state spaces, such as dynamic adjustment of sludge age and long-term carbon source addition strategy optimization, then the Policy Gradient (PG) method is selected. This algorithm uses a policy parameterization direct optimization framework, without relying on value function approximation, and can effectively solve nonlinear constraints and large-scale optimization problems in the process. To address the challenges of adapting to the long-term configuration parameter optimization needs of complex wastewater treatment processes, and for applications requiring rapid response and high long-term stability, such as transmembrane pressure differential control in MBR processes and aeration-sedimentation timing switching control in SBR processes, an Actor-Critic architecture is adopted. This architecture features a dual-network design with separate policy network (Actor) and value network (Critic). The Actor is responsible for action space exploration and decision output, while the Critic is responsible for real-time evaluation and feedback of policy value. This significantly improves the convergence stability and sample utilization efficiency of model training, while balancing real-time decision-making and long-term robustness. Using high-dimensional feature vectors as input, and the first reward function ( The current water quality compliance rate within the observation window. This represents the cumulative energy consumption within the current observation window. This represents the cumulative drug consumption within the current observation window. , , The model is iteratively trained with weighted coefficients as the optimization objective until convergence, at which point short-term optimization parameters such as aeration time, dissolved oxygen concentration, and reflux ratio are output.

[0039] The long-term scale configuration module 40 uses daily to weekly control cycles. Based on the long-term operating parameters of the cleaning dataset and the preset historical operating database, it optimizes the long-term configuration parameters in combination with preset evaluation indicators. During the optimization process, the operating condition change patterns in the historical operating data and the long-term operating results are fully coupled. According to the multi-objective optimization requirements of the process, reinforcement learning algorithms are selected. For complex multi-objective optimization problems such as sludge age adjustment and carbon source addition strategy, PG algorithm or PPO (proximal strategy optimization) algorithm is used. Constructing a second reward function ,in , , Weighting coefficients; system stability index values The weighted scoring method is used for calculation, and the formula is as follows: ,in The formula for calculating the fluctuation of effluent indicators is as follows: , This indicates the maximum deviation of effluent COD within the set time unit; Indicates the upper limit of the COD standard for effluent; The fault-free operating rate of equipment is expressed by the following formula: , This refers to the cumulative runtime of the device within a set time unit. This indicates the cumulative downtime of the device within the set time unit; This indicates the frequency of adjustment based on process parameters (≤1 time / day is 1, 2 times / day is 0.5, ≥3 times / day is 0). , , These are the weighting coefficients.

[0040] COD removal rate The calculation formula is as follows: , Indicates the total amount of water entering the system. This represents the average COD value of the influent. Indicates the total amount of water discharged. This represents the average COD value of the effluent.

[0041] Operating costs The calculation formula is as follows: , The cumulative cost of energy consumption within the set time unit. This indicates the cumulative cost of drug consumption within the set time unit. This represents the cumulative cost of sludge treatment within the set time unit. This indicates the equipment maintenance cost within the set time unit; The long-term configuration model is iteratively trained with the goal of maximizing the second reward function. After convergence, the long-term configuration optimization parameters such as sludge age and carbon source addition strategy are output.

[0042] The process adaptation module 50 dynamically adjusts the action parameter dimension of the short-timescale control module 30 and the state parameter dimension of the long-timescale configuration module 40 according to the target wastewater treatment process type (such as SBR, A² / O, MBR) through the built-in parameter template library. The template library presets a 5-10 dimensional state space and a 3-6 dimensional action space adaptation range to achieve rapid adaptation of different processes.

[0043] Subsequently, the multi-timescale fusion control module 60 dynamically adjusts the weighting coefficients based on real-time operating conditions. Through weighted fusion formula , ( Indicates short-term optimization parameters. (This refers to long-term configuration optimization parameters), which integrate the two types of parameters to generate a comprehensive control command that balances real-time response and long-term stability.

[0044] After receiving the comprehensive control command, the execution module 70 drives the wastewater treatment actuators such as blowers, return pumps, and dosing devices to perform precise actions, implement the adjustment requirements of various parameters, and realize the transformation of control commands into process operation.

[0045] The strategy update module 80 collects the control result data from the execution module 70, calculates reward feedback based on the first and second reward functions, uses an experience replay mechanism to store state-action-reward correlation data, and combines the target network update mechanism to iteratively update the strategy networks of the short-term control model and the long-term configuration model, thereby continuously improving the system's control performance and robustness.

[0046] The technical effects and adaptability of this system are further illustrated below through specific embodiments: In the application of the A² / O process wastewater treatment plant with a daily treatment capacity of 20,000 tons, the data acquisition module 10 is equipped with sensors such as pH (6-9), COD (0-1000mg / L), NH3-N (0-50mg / L), DO (0-8mg / L) and flow meters (0-3000m³ / h), and completes data acquisition and uploading at a frequency of 1 minute through the PLC system; the data preprocessing module 20 removes abnormal values ​​such as a sudden increase in COD of 2000mg / L, generates a state sequence through a 10-minute sliding window and normalizes it to the [0,1] interval; The short-timescale control module 30 constructs a short-term strategy network based on the Actor-Critic architecture. It takes a 10×5-dimensional state sequence as input and outputs aeration time (0-6 hours) and DO setpoint (1-3 mg / L). When the influent COD increases from 300 mg / L to 450 mg / L, the system adjusts the DO from 2.0 mg / L to 2.4 mg / L, extends the aeration time by 20 minutes, and stabilizes the effluent COD below 40 mg / L. The long-term configuration module 40 optimizes sludge age (5-20 days) and carbon source dosage (0-100 kg / d) based on the PPO algorithm with a 7-day cycle. After 3 months of operation, the sludge age was optimized from 15 days to 13 days, and the COD removal rate increased to 92%. The multi-timescale integrated control module 60 initially set α=0.7 and adjusted to 0.9 when the effluent fluctuation exceeded 10%. The execution module 70 executes commands through a 75kW blower and a 500m³ / h return pump. The strategy update module 80 updates the strategy on a daily cycle with an experience playback buffer capacity of 5000. The target network is updated every 100 iterations, and the daily energy consumption is reduced from 500kWh to 465kWh, and the effluent compliance rate is increased from 95% to 99%.

[0047] In the validation of the SBR process wastewater treatment plant with a daily treatment capacity of 5,000 tons, the system loaded a pre-trained DQN short-term model and a PPO long-term model and initialized the parameters; the data acquisition module 10 collected the influent COD (average 250 mg / L) and flow rate (200-300 m³ / h) in real time, and generated a state sequence after pretreatment; the short-timescale control module 30 optimized the aeration time and sedimentation time with a 1-hour cycle. When the influent flow rate increased to 350 m³ / h, the aeration time was extended from 3 hours to 3.5 hours, and the effluent COD decreased from 60 mg / L to 45 mg / L. The long-term configuration module 40 optimizes sludge age with a 7-day cycle, adjusting it from 10 days to 9 days, increasing COD removal rate from 88% to 91%; the fusion module generates comprehensive instructions with α=0.8, resulting in an energy consumption increase of ≤4% after execution and an effluent compliance rate of 98%; the strategy update module, based on a feedback update model, shortens the convergence time from 50 iterations to 30 iterations, and after one month of operation, daily energy consumption is reduced by 6% (200kWh→188kWh), chemical consumption is reduced by 10% (5kg→4.5kg), and the average effluent COD is stabilized below 40mg / L.

[0048] In the multi-process adaptability verification, for the SBR process (5000 tons / day), by adjusting the aeration time (3-5 hours) and sedimentation time (1-2 hours), when the influent NH3-N increased from 20 mg / L to 30 mg / L, the aeration time was extended to 4.5 hours, and the effluent NH3-N ≤ 5 mg / L, with an energy consumption increase of only 3%; for the A² / O process (20,000 tons / day), by adjusting the reflux ratio (50%-100%) and DO setpoint (1.5-3 mg / L), when the load fluctuated by 20%, the reflux ratio was optimized to 80% and DO to 2.2 mg / L, and the effluent COD ≤ 45 mg / L and NH3-N ≤ 4 mg / L, with an energy consumption saving of 5%; The MBR process (10,000 tons / day) added transmembrane pressure difference (TMP, 20-50 kPa) and membrane cleaning frequency (1-3 times / week) as controllable parameters. After optimization, TMP decreased from 40 kPa to 35 kPa, and the cleaning frequency decreased from 2 times / week to 1.5 times / week, extending membrane life by 10% and reducing energy consumption by 8%. Regarding process adaptation, after adding the TMP state dimension to the MBR process, the system automatically adjusted the model structure through the parameter template library, achieving 90% convergence in just 2 weeks of training, demonstrating highly efficient process transfer capabilities and fully validating the system's broad adaptability to mainstream wastewater treatment processes such as SBR, A² / O, and MBR.

[0049] Example 2: Based on the same inventive concept as Embodiment 1, the present invention also provides an intelligent control method for wastewater treatment processes based on multi-timescale reinforcement learning, comprising the following steps: High-precision sensors are deployed at the inlet, anaerobic / anoxic / aerobic tanks, and outlet to collect real-time operating parameters during the wastewater treatment process. Anomaly detection and filtering are performed on the real-time operating parameters to form a cleaning dataset that can be used for model training and decision optimization. By setting the control cycle to the minute to hour level, the short-term operating parameters in the cleaned dataset are optimized and controlled in real time to obtain short-term optimized parameters. Based on the long-term operating parameters in the cleaning dataset and the preset historical operating database, combined with the preset evaluation indicators, the control cycle is set to daily to weekly, and the long-term configuration parameters of the sewage treatment process are periodically optimized. In the optimization process, the operating condition change patterns in the historical operating data and the long-term operating results are considered to obtain the long-term configuration optimization parameters. Based on the type of the target wastewater treatment process, the action parameter dimension of the short-timescale control module and the state parameter dimension of the long-timescale configuration module are dynamically adjusted. The weighting coefficients are dynamically adjusted based on real-time operating conditions, and the short-term optimization parameters and the long-term configuration optimization parameters are combined to generate a comprehensive control command. The comprehensive control commands drive the actions of the actuators in the wastewater treatment process.

[0050] This embodiment proposes an intelligent control method for wastewater treatment processes based on multi-timescale reinforcement learning. This method is implemented through the aforementioned intelligent control system for wastewater treatment processes based on multi-timescale reinforcement learning. Therefore, the specific implementation methods of the intelligent control method for wastewater treatment processes based on multi-timescale reinforcement learning can be found in the embodiment section of the aforementioned intelligent control system for wastewater treatment processes based on multi-timescale reinforcement learning. To avoid redundancy, they will not be repeated here.

[0051] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0052] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0053] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0054] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0055] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A sewage treatment process intelligent regulation and control system based on multi-time scale reinforcement learning, characterized in that, The system comprises the following modules: a data acquisition module for acquiring real-time operation parameters in the sewage treatment process; a data preprocessing module for performing anomaly detection and filtering processing on the real-time operation parameters to form a cleaning data set that can be used for model training and decision optimization; a short-time scale regulation module for setting a regulation period to the minute level to the hour level, performing real-time optimization control on short-term operation parameters in the cleaning data set, and obtaining short-term optimization parameters; a long-time scale configuration module for setting a regulation period to the day level to the week level based on long-term operation parameters in the cleaning data set and a preset historical operation database, combining a preset evaluation index, performing periodic optimization on long-term configuration parameters of the sewage treatment process, and considering the working condition change law and long-term operation result in the historical operation data in the optimization process to obtain long-term configuration optimization parameters; a process adaptation module for dynamically adjusting the action parameter dimension of the short-time scale regulation module and the state parameter dimension of the long-time scale configuration module according to the type of the target sewage treatment process; a multi-time scale fusion regulation module for dynamically adjusting a weight coefficient according to a real-time working condition, fusing the short-term optimization parameters and the long-term configuration optimization parameters to generate a comprehensive regulation instruction; and an execution module for driving an execution mechanism in the sewage treatment process to act through the comprehensive regulation instruction.

2. The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process according to claim 1, wherein, The system further comprises a strategy updating module for obtaining regulation result data of the execution module, calculating reward feedback based on a preset reward function, the reward function containing a weighted item of at least one evaluation index, updating a strategy network of the short-time scale regulation module and the long-time scale configuration module using an experience replay mechanism and a target network updating mechanism.

3. The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process according to claim 1, wherein, The short-time scale regulation module is configured to set a regulation period to the minute level to the hour level, perform real-time optimization control on short-term operation parameters in the cleaning data set, and obtain short-term optimization parameters, and comprises: based on the regulation period, acquiring real-time operation parameters in a current observation window through a sliding window at a preset time interval, and synchronously extracting historical operation data of a preset time length before the observation window; performing time series data preprocessing on the acquired real-time-historical joint data sequence to convert it into a high-dimensional feature vector that can be analyzed by a model; based on a reinforcement learning algorithm, constructing a short-time scale regulation model, taking the high-dimensional feature vector as a model input, performing iterative training on the short-time scale regulation model through a first reward function, taking reward function value maximization as an optimization objective in the training process, until the short-time scale regulation model converges, and finally generating process short-term optimization parameters in the short-time scale.

4. The sewage treatment process intelligent regulation system based on multi-time scale reinforcement learning according to claim 3, wherein the first reward function is as follows: wherein, is the compliance rate of water quality in the current observation window, is the cumulative energy consumption in the current observation window, is the cumulative drug consumption in the current observation window, , , is the weight coefficient.

5. The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process according to claim 3, wherein, The short-time scale regulation model comprises one or a combination of DQN, PG, and Actor-Critic reinforcement learning models, and is selected according to the process response complexity and the regulation target dimension.

6. The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process according to claim 1, wherein, The long-time scale configuration module is configured to set a regulation period to a daily level to a weekly level based on long-term operation parameters in the cleaning data set and a preset historical operation database, in combination with a preset evaluation index, to periodically optimize long-term configuration parameters of the sewage treatment process, to consider a working condition change law and a long-term operation result in the historical operation data in the optimization process, and to obtain long-term configuration optimization parameters, including: Based on the pretreated long-term operation parameters, system stability index values are sequentially calculated , COD removal rate , and operation cost ; according to the system stability index values , the COD removal rate , and the operation cost , a second reward function is constructed as follows: wherein , , are weight coefficients; A long-time scale configuration model is constructed, and the long-time scale configuration model is iteratively trained by using the second reward function. The training process takes maximizing a reward function value as an optimization objective until the long-time scale configuration model converges, and finally generates long-term configuration optimization parameters.

7. The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process according to claim 6, wherein, The system stability index value is calculated by using a weighted scoring method, and a formula is as follows: wherein represents the water index fluctuation degree, and the calculation formula is , represents the maximum deviation value of the effluent COD in the set time unit; represents the upper limit value of the effluent COD standard; represents the equipment failure-free operation rate, and the calculation formula is , is the cumulative running time of the equipment in the set time unit; represents the cumulative failure time of the equipment in the set time unit; represents the assignment according to the process parameter adjustment frequency; , , is a weight coefficient. 8.The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process of claim 6, wherein, The COD removal rate The calculation formula is as follows: , represents the total amount of influent, represents the average value of influent COD; represents the total amount of effluent, represents the average value of effluent COD. 9.The multi-time scale reinforcement learning based intelligent control system for wastewater treatment process of claim 6, wherein, The running cost The formula for the calculation is as follows: , cumulative cost of energy consumption in the set time unit, cumulative cost of medicine consumption in the set time unit, cumulative cost of sludge treatment in the set time unit, cumulative cost of equipment maintenance in the set time unit.

10. A sewage treatment process intelligent regulation method based on multi-time scale reinforcement learning, characterized in that, The method comprises the following steps: Real-time operation parameters in a sewage treatment process are collected, and the real-time operation parameters are subjected to abnormality detection and filtering processing to form a cleaning data set that can be used for model training and decision optimization; A regulation period is set to a minute level to an hour level, real-time optimization control is performed on short-term operation parameters in the cleaning data set, and short-term optimization parameters are obtained; Based on long-term operation parameters in the cleaning data set and a preset historical operation database, in combination with a preset evaluation index, a regulation period is set to a daily level to a weekly level, long-term configuration parameters of the sewage treatment process are periodically optimized, a working condition change law and a long-term operation result in the historical operation data are considered in the optimization process, and long-term configuration optimization parameters are obtained; According to a type of a target sewage treatment process, action parameter dimensions of the short-time scale regulation module and state parameter dimensions of the long-time scale configuration module are dynamically adjusted; A weight coefficient is dynamically adjusted according to a real-time working condition, the short-term optimization parameters and the long-term configuration optimization parameters are fused to generate a comprehensive regulation instruction; An executing mechanism in the sewage treatment process is driven by using the comprehensive regulation instruction.