A water environment treatment process dynamic regulation system and method based on reinforcement learning

By using reinforcement learning models and meta-learning strategies, the rigidity and poor adaptability of traditional PID control are solved, enabling adaptive, safe and efficient control of water environment treatment processes, reducing energy consumption and the risk of water quality exceeding standards, and making it applicable to a variety of water treatment processes.

CN122362864APending Publication Date: 2026-07-10CHINA THREE GORGES UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2026-04-29
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional PID control methods suffer from problems such as rigid parameters, inability to balance multiple objectives, and poor adaptability in water environment treatment, resulting in lagging control, energy waste, and water quality exceeding standards, and failing to achieve intelligent full-process control.

Method used

A dynamic control system for water environment treatment processes based on reinforcement learning is adopted. Through data acquisition, reinforcement learning model, dynamic control and data storage modules, a state space and action space are constructed. Multi-head attention encoding and Bayesian Q-network uncertainty estimation are used to achieve adaptive dual-objective optimization and safety decision-making. Combined with distributed deployment and meta-learning strategies, rapid adaptation is achieved.

Benefits of technology

It achieves a dynamic balance between meeting water quality standards and minimizing energy consumption, reduces operating costs and equipment failure risks, improves system adaptability and control efficiency, and is suitable for various water environment treatment processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122362864A_ABST
    Figure CN122362864A_ABST
Patent Text Reader

Abstract

This invention discloses a dynamic control system and method for water environment treatment processes based on reinforcement learning. The system includes a data acquisition module, a reinforcement learning model module, a dynamic control module, and a data storage and interaction module, forming a closed-loop control system. The data acquisition module acquires water quality status, equipment operation, and environmental disturbance data in real time. The reinforcement learning model module extracts multi-scale modal features of time-series data through variational mode decomposition, generates weighted state representations through multi-head attention mechanism encoding, uses Bayesian Q-network uncertainty estimation to impose safety constraints on exploration actions, constructs a dual-objective reward function of achieving water quality standards and minimizing energy consumption, and adaptively adjusts the reward weights based on operating condition information entropy. The dynamic control module converts action decisions into control commands and issues them to the execution equipment. This invention solves the technical problems of rigid parameters, inability to simultaneously consider multiple objectives, and poor adaptability of traditional PID control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology for water environment treatment, and in particular relates to a dynamic control system and method for water environment treatment processes based on reinforcement learning. Background Technology

[0002] The stable and efficient operation of water environment treatment processes is crucial for ensuring effluent quality meets standards and reducing operational energy consumption. Currently, the vast majority of water environment treatment facilities in China use traditional PID control. This control method relies on fixed control parameters and can only be adjusted for a single control objective, presenting significant industry pain points and failing to meet the core requirements of "high efficiency, energy saving, and stability" in current water environment treatment.

[0003] Specifically, the limitations of traditional PID control are mainly reflected in three aspects: First, the control parameters are rigid and cannot adapt to dynamic interference factors such as fluctuations in water quality and quantity, changes in influent pollutant concentration, and equipment operating status degradation during the water environment treatment process, resulting in lagging regulation and easy problems such as water quality exceeding standards or energy waste. Second, multiple objectives cannot be taken into account. Traditional control can only prioritize ensuring a single water quality indicator, making it difficult to balance the two core objectives of "water quality compliance" and "lowest energy consumption." This often leads to a dilemma where excessive consumption of electricity and reagents is used to ensure water quality, or water quality exceeds standards in order to reduce energy consumption. Third, the adaptability is poor. Facilities with different treatment processes, different influent water quality, and different treatment scales require manual readjustment of PID parameters, resulting in weak versatility. Furthermore, manual readjustment relies on the experience of operation and maintenance personnel, leading to large errors and low efficiency, and failing to achieve intelligent control of the entire process.

[0004] In existing technologies, some water environment treatment and control attempts have introduced simple intelligent algorithms, but these are mostly single-objective optimizations, failing to achieve synergistic optimization of water quality and energy consumption, and failing to solve the problem of adaptive control under dynamic disturbances, thus still unable to overcome the limitations of traditional PID control. Therefore, developing an intelligent dynamic control system and method that can take into account dual objectives, adapt to dynamic disturbances, and has strong adaptability has become an urgent technical challenge in the field of water environment treatment, and is also a key breakthrough for improving the intelligence level of water environment treatment and promoting energy conservation and emission reduction in the industry; therefore, it is necessary to propose a dynamic control system and method for water environment treatment processes based on reinforcement learning to solve the above problems. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a dynamic control system and method for water environment treatment process based on reinforcement learning. It aims to solve the problems of rigidity, inability to take multiple objectives into account, and poor adaptability of traditional PID control in the prior art. Through the autonomous learning and dynamic optimization of the reinforcement learning model, the adaptive and precise control of water environment treatment process parameters can be achieved, overcoming the limitations of traditional control, while quantitatively improving the treatment effect and reducing operating costs.

[0006] To achieve the above technical solution, the technical solution adopted by the present invention is as follows: A dynamic control system for water environment treatment processes based on reinforcement learning, comprising: The data acquisition module is used to collect real-time status data of the entire water environment treatment process, including water quality status data, equipment operation data, and environmental disturbance data. Water quality status data includes COD, BOD5, NH3-N, TN, TP, DO, pH value, and SS at the inlet and outlet. Equipment operation data includes the operating power, speed, and valve opening of water pumps, blowers, aeration equipment, and chemical dosing equipment. Environmental disturbance data includes influent flow rate, influent concentration fluctuation, and ambient temperature. The reinforcement learning model module includes a state space construction unit, an action space construction unit, a dual-objective reward function unit, and a policy optimization unit. The state space construction unit normalizes the collected state data, extracts multi-scale modal features of the time-series state data through variational mode decomposition, and then encodes these features using a multi-head attention mechanism to generate weighted state representations, thus constructing the system state vector. The action space construction unit defines process-adjustable parameters as action vectors. The dual-objective reward function unit constructs a composite reward function. , This is a reward item for meeting water quality standards. Energy consumption optimization reward items, , The weighting coefficients and And based on the entropy of working condition information, and Adaptive adjustments are made; the policy optimization unit employs a deep reinforcement learning algorithm and uses Bayesian Q-network uncertainty estimation to impose safety constraints on the exploration actions, outputting the optimal action decision. The dynamic control module is used to receive the optimal action decision and convert it into control commands to be sent to each execution device, while feeding back the execution status of the device to the reinforcement learning model module; The data storage and interaction module is used to store various types of real-time data, training data, and historical data, and provides a human-computer interaction interface.

[0007] Preferably, the data acquisition module adopts a distributed deployment, with sensors and smart instruments set up in each process section. The sampling frequency is 1 to 5 minutes each time, and the data transmission adopts the OPC UA protocol. Noise data is filtered through edge nodes and multi-device timestamp synchronization is achieved at the millisecond level. The pre-processed data is synchronously sent to the reinforcement learning model module and the data storage and interaction module. The execution equipment includes water pumps, blowers, aeration equipment, dosing equipment, sedimentation tanks, and membrane modules. The control core adopts a PLC system.

[0008] Preferably, the state space construction unit includes a variational mode decomposition subunit, used to perform variational mode decomposition on each state variable in the time-series state data, and for the j-th state variable x j (t), solve the constrained variational model: ; ; in, For the k-th eigenmode function obtained from the decomposition, Let be the center frequency of the k-th mode. To preset the number of decomposition modes, For the Dirac function, This represents the convolution operation. This involves taking the partial derivative with respect to time and extracting the energy characteristics of each mode. and center frequency This constitutes the multi-scale feature vector of the state variable; after concatenating the multi-scale feature vectors of all state variables, the input is given to the multi-head attention encoding subunit.

[0009] Preferably, the multi-head attention encoding subunit uses a linear transformation matrix. Map the input multi-scale feature vector to a query matrix Q, a key matrix K, and a value matrix V, and compute the scaled dot product attention: ; in, The scaling factor is the column dimension of matrix K; multiple attention heads are used for parallel computation, and the output of the i-th attention head is... ,in Let be the projection matrix of the i-th attention head; after concatenating the outputs of all attention heads, a weighted state representation is obtained through linear transformation. : ; Where h is the number of attention heads. To output the projection matrix; The final system state vector is obtained by concatenating it with the original normalized state vector.

[0010] Preferably, the policy optimization unit includes a Bayesian uncertainty estimation subunit. When the deep reinforcement learning algorithm is the DQN algorithm, a Monte Carlo Dropout mechanism is introduced into the Q-network to perform estimation on the same state-action pair. After a random forward propagation, calculate the mean and variance of the Q-values: ; ; in, Output the Q-value for the m-th random forward propagation. The number of Dropout samples; the policy optimization unit adopts an uncertainty-guided exploration strategy when selecting actions: ; in, For uncertainty exploration coefficients; when Exceeding the preset safety threshold At this time, the strategy optimization unit abandons the current exploration action and instead executes a conservative action based on process safety rules to avoid excessive effluent or equipment failure due to exploration with high uncertainty.

[0011] Preferably, the dual-objective reward function unit includes an information entropy adaptive weight adjustment subunit, which calculates the working condition information entropy H(t) in real time: ; ; in, The number of state variables involved in the entropy calculation. Let be the normalized value of the deviation between the current value and the reference value of the i-th state variable. The current value, The baseline value is used; the weighting coefficients are dynamically adjusted based on the entropy of the operating condition information. ; ; in, As the initial water quality compliance weight, To adjust the coefficient, Use the reference entropy value; when Higher than When the time indicates an increase in operational complexity, the time will automatically increase. To ensure water quality safety; when Below When the operating condition is stable, the speed will automatically decrease. The focus is on energy consumption optimization.

[0012] Preferably, when the dynamic control module executes action decisions, it sets safety constraint ranges for each action parameter. When a certain action parameter in the optimal action decision exceeds the rated operating range of the corresponding execution device, the action parameter is automatically truncated to a safe boundary value. The dynamic control module also sets an emergency control mechanism. When the effluent water quality index exceeds 1.5 times the threshold of the compliance standard for two consecutive sampling cycles or the equipment operating parameters exceed the safe range, an emergency command is triggered to switch each execution device to a preset safe operating state and issue an alarm signal.

[0013] Preferably, a method for dynamic control system of water environment treatment process based on reinforcement learning includes the following steps: S1, System initialization, setting water quality standards, equipment operating baseline energy consumption, and initial parameters of the reinforcement learning model, and completing the debugging of each module and calibration of equipment parameters; S2, real-time data acquisition and preprocessing: The data acquisition module collects the status data of the entire process, performs noise reduction, normalization and outlier removal preprocessing, and then transmits it to the reinforcement learning model module and the data storage and interaction module. S3, State Space Construction: Variational mode decomposition is performed on the preprocessed temporal state data to extract multi-scale modal features. Weighted state representations are generated through multi-head attention mechanism to construct system state vectors and action vectors. The reward value of the current state is calculated. Bayesian Q network uncertainty estimation is used to impose safety constraints on exploration actions. The policy network parameters are updated through deep reinforcement learning algorithm to output the optimal action decision. S4, dynamic control execution, transforms the optimal action decision into control commands and sends them to the execution equipment to adjust process parameters, while simultaneously feeding back the equipment execution status to the reinforcement learning model module; S5, Model Iteration Optimization and Anomaly Handling: The strategy network is continuously iterated and updated based on real-time data and equipment feedback. The reward function weight coefficient is adaptively adjusted by calculating the entropy of working condition information. An emergency mechanism is triggered and an alarm is sounded when an abnormal situation occurs. S6, data storage and interaction, stores all data in the storage process and enables status monitoring, parameter adjustment and manual intervention through a human-machine interface.

[0014] Preferably, in step S3, variational mode decomposition decomposes the time series of each state variable into... Each intrinsic mode function (EMF) is used to extract the energy and center frequency features of each mode to construct a multi-scale feature vector. Multi-head attention encoding then concatenates the multi-scale feature vector by simultaneously calculating the attention weights of h attention heads to output a weighted state representation. Bayesian Q-network uncertainty estimation is performed in the DQN algorithm using... The variance of the Q-value is calculated using Monte Carlo Dropout sampling. When the variance exceeds a safety threshold... The conservative strategy is triggered at the time; in step S5, the working condition information entropy is determined according to... The degree to which each state variable deviates from the baseline value is calculated in real time, and the weighting coefficients are dynamically adjusted accordingly. and Adjust the formula to , .

[0015] Preferably, step S5 further includes a periodic offline training and optimization phase, in which the policy network is trained offline in batches using stored historical operating data. During the training process, a model-independent meta-learning strategy is introduced, and inner loop gradient updates and outer loop meta-gradient accumulations are performed on historical data subsets corresponding to multiple process scenarios to obtain cross-scenario universal initial network parameters. : ; in, For the outer loop learning rate, For batch quantity in the process scenario, This represents the training task corresponding to the i-th process scenario. For the network after the internal loop update, For the task The loss function is applied; when the system is deployed to a new process scenario, the loss function is applied. Rapid adaptation can be achieved by making a small number of interactive fine-tunings as initial parameters, with no more than 50 fine-tuning steps.

[0016] The beneficial effects of this invention are as follows: 1. This invention solves the technical problems of rigid parameters and inability to adapt to dynamic disturbances in traditional PID control by introducing variational mode decomposition and multi-head attention encoding mechanisms into the state space construction unit of the reinforcement learning model module. Specifically, the state space construction unit first performs variational mode decomposition on the time-series data of each state variable, adaptively decomposing the original signal into multiple intrinsic mode functions of different frequency bands, and extracting the energy characteristics and center frequency characteristics of each mode to achieve multi-scale representation of short-term abrupt changes and medium-to-long-term trend changes in influent water quality. Subsequently, the multi-head attention mechanism is used to automatically learn the complex nonlinear correlation weights between each state variable, and the multi-scale features are weighted and encoded to generate a weighted state representation that reflects the global dynamic characteristics of the process. The system state vector constructed in this way not only contains numerical information at a single moment, but also integrates the frequency domain distribution characteristics of the time-series signal and the correlation structure between variables, enabling the reinforcement learning model to keenly perceive the early signs and evolution trends of multi-source dynamic disturbances such as sudden fluctuations in influent pollutant concentration and slow degradation of equipment performance, and make timely forward-looking control decisions, avoiding the control lag and water quality exceedance problems caused by the single information representation in traditional control.

[0017] 2. This invention addresses the technical problems of traditional control systems failing to simultaneously address multiple objectives and the inherent safety risks in the exploration process of intelligent algorithms by introducing a Bayesian Q-network uncertainty estimation mechanism into the strategy optimization unit and combining it with an information entropy adaptive weight adjustment strategy. On one hand, the strategy optimization unit embeds a Monte Carlo Dropout mechanism into the Q-network of the DQN algorithm. By performing multiple random forward propagations on the same state-action pair to calculate the mean and variance of the Q-value, the uncertainty of each decision is quantified. When the variance of the Q-value of a certain exploration action exceeds a preset safety threshold, the system automatically abandons the high-risk action and instead executes a conservative action based on process safety rules. This constructs a safety barrier for the training process without sacrificing the reinforcement learning exploration capability, preventing severe exceedances of effluent indicators or equipment damage due to blind exploration. On the other hand, the dual-objective reward function unit quantifies the complexity of influent water quality fluctuations and process operation status by calculating the entropy of operating conditions in real time. When the complexity of operating conditions increases, it automatically increases the weight of water quality compliance to ensure the safety of effluent. When the operating conditions tend to be stable, it automatically decreases the weight of water quality compliance to focus on energy consumption optimization. This achieves a dynamic adaptive balance between the two objectives of water quality compliance and minimum energy consumption, completely getting rid of the dilemma of traditional control where one objective is neglected while the other is achieved.

[0018] 3. This invention solves the technical problems of poor adaptability of traditional control systems and the need for repeated manual adjustments for different processes by introducing a model-independent meta-learning strategy into the method and combining it with the distributed deployment of the data acquisition module and the closed-loop feedback mechanism of the dynamic control module. In the periodic offline training and optimization stage, the system uses stored subsets of historical data from multiple process scenarios and trains universal initialization network parameters across scenarios through a meta-learning mechanism of inner loop gradient update and outer loop meta-gradient accumulation. When the system is deployed to a new water environment treatment process scenario, it does not need to be trained from scratch; it can be quickly adapted by making only a few interactive fine-tunings based on these universal initialization parameters, significantly reducing the training cost and time investment of the model in new scenarios. Meanwhile, the data acquisition module adopts a distributed deployment and achieves millisecond-level synchronization of multi-sensor timestamps through the OPC UA protocol. The dynamic control module converts the optimal action decision into control commands and sends them to the execution equipment and transmits the execution status back in real time, forming a complete closed-loop control system. These components work together with the meta-learning strategy, enabling this solution to flexibly adapt to various water environment treatment processes such as A² / O process, MBR membrane bioreactor, and oxidation ditch, without the need to redesign control schemes for different processes. This fundamentally breaks through the technical bottlenecks of traditional PID control, which suffers from weak versatility and low efficiency of manual debugging. Attached Figure Description

[0019] Figure 1 This is a structural block diagram of the present invention; Figure 2 This is a schematic diagram of the method flow of the present invention; Figure 3 This is a schematic diagram of the distributed deployment and data transmission of the data acquisition module in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the principle and control points of an embodiment of the invention applied to the A2 / O process; Figure 5 This is a schematic diagram illustrating the calculation logic and training effect of the dual-objective reward function in an embodiment of the present invention; Figure 6 This is a schematic diagram of the human-computer interaction interface (HMI) in an embodiment of the present invention. Detailed Implementation

[0020] Example 1: like Figure 1 As shown, a dynamic control system for water environment treatment processes based on reinforcement learning includes: The data acquisition module is used to collect real-time status data of the entire water environment treatment process, including water quality status data, equipment operation data, and environmental disturbance data. Water quality status data includes COD, BOD5, NH3-N, TN, TP, DO, pH value, and SS at the inlet and outlet. Equipment operation data includes the operating power, speed, and valve opening of water pumps, blowers, aeration equipment, and chemical dosing equipment. Environmental disturbance data includes influent flow rate, influent concentration fluctuation, and ambient temperature. The reinforcement learning model module includes a state space construction unit, an action space construction unit, a dual-objective reward function unit, and a policy optimization unit. The state space construction unit normalizes the collected state data, extracts multi-scale modal features of the time-series state data through variational mode decomposition, and then encodes these features using a multi-head attention mechanism to generate weighted state representations, thus constructing the system state vector. The action space construction unit defines process-adjustable parameters as action vectors. The dual-objective reward function unit constructs a composite reward function. , This is a reward item for meeting water quality standards. Energy consumption optimization reward items, , The weighting coefficients and And based on the entropy of working condition information, and Adaptive adjustments are made; the policy optimization unit employs a deep reinforcement learning algorithm and uses Bayesian Q-network uncertainty estimation to impose safety constraints on the exploration actions, outputting the optimal action decision. The dynamic control module is used to receive the optimal action decision and convert it into control commands to be sent to each execution device, while feeding back the execution status of the device to the reinforcement learning model module; The data storage and interaction module is used to store various types of real-time data, training data, and historical data, and provides a human-computer interaction interface.

[0021] Preferably, the data acquisition module adopts a distributed deployment, with sensors and smart instruments set up in each process section. The sampling frequency is 1 to 5 minutes each time, and the data transmission adopts the OPC UA protocol. Noise data is filtered through edge nodes and multi-device timestamp synchronization is achieved at the millisecond level. The pre-processed data is synchronously sent to the reinforcement learning model module and the data storage and interaction module. The execution equipment includes water pumps, blowers, aeration equipment, dosing equipment, sedimentation tanks, and membrane modules. The control core adopts a PLC system.

[0022] Preferably, the state space construction unit includes a variational mode decomposition subunit, used to perform variational mode decomposition on each state variable in the time-series state data, and for the j-th state variable x j (t), solve the constrained variational model: ; ; in, For the k-th eigenmode function obtained from the decomposition, Let be the center frequency of the k-th mode. To preset the number of decomposition modes, For the Dirac function, This represents the convolution operation. This involves taking the partial derivative with respect to time and extracting the energy characteristics of each mode. and center frequency This constitutes the multi-scale feature vector of the state variable; after concatenating the multi-scale feature vectors of all state variables, the input is given to the multi-head attention encoding subunit.

[0023] Preferably, the multi-head attention encoding subunit uses a linear transformation matrix. Map the input multi-scale feature vector to a query matrix Q, a key matrix K, and a value matrix V, and compute the scaled dot product attention: ; in, The scaling factor is the column dimension of matrix K; multiple attention heads are used for parallel computation, and the output of the i-th attention head is... ,in Let be the projection matrix of the i-th attention head; after concatenating the outputs of all attention heads, a weighted state representation is obtained through linear transformation. : ; Where h is the number of attention heads. To output the projection matrix; The final system state vector is obtained by concatenating it with the original normalized state vector.

[0024] Preferably, the policy optimization unit includes a Bayesian uncertainty estimation subunit. When the deep reinforcement learning algorithm is the DQN algorithm, a Monte Carlo Dropout mechanism is introduced into the Q-network to perform estimation on the same state-action pair. After a random forward propagation, calculate the mean and variance of the Q-values: ; ; in, Output the Q-value for the m-th random forward propagation. The number of Dropout samples; the policy optimization unit adopts an uncertainty-guided exploration strategy when selecting actions: ; in, For uncertainty exploration coefficients; when Exceeding the preset safety threshold At this time, the strategy optimization unit abandons the current exploration action and instead executes a conservative action based on process safety rules to avoid excessive effluent or equipment failure due to exploration with high uncertainty.

[0025] Preferably, the dual-objective reward function unit includes an information entropy adaptive weight adjustment subunit, which calculates the working condition information entropy H(t) in real time: ; ; in, The number of state variables involved in the entropy calculation. Let be the normalized value of the deviation between the current value and the reference value of the i-th state variable. The current value, The baseline value is used; the weighting coefficients are dynamically adjusted based on the entropy of the operating condition information. ; ; in, As the initial water quality compliance weight, To adjust the coefficient, Use the reference entropy value; when Higher than When the time indicates an increase in operational complexity, the time will automatically increase. To ensure water quality safety; when Below When the operating condition is stable, the speed will automatically decrease. The focus is on energy consumption optimization.

[0026] Preferably, when the dynamic control module executes action decisions, it sets safety constraint ranges for each action parameter. When a certain action parameter in the optimal action decision exceeds the rated operating range of the corresponding execution device, the action parameter is automatically truncated to a safe boundary value. The dynamic control module also sets an emergency control mechanism. When the effluent water quality index exceeds 1.5 times the threshold of the compliance standard for two consecutive sampling cycles or the equipment operating parameters exceed the safe range, an emergency command is triggered to switch each execution device to a preset safe operating state and issue an alarm signal.

[0027] like Figure 2 and Figure 3 As shown, preferably, a method for a dynamic control system of a water environment treatment process based on reinforcement learning includes the following steps: S1, System initialization, setting water quality standards, equipment operating baseline energy consumption, and initial parameters of the reinforcement learning model, and completing the debugging of each module and calibration of equipment parameters; S2, real-time data acquisition and preprocessing: The data acquisition module collects the status data of the entire process, performs noise reduction, normalization and outlier removal preprocessing, and then transmits it to the reinforcement learning model module and the data storage and interaction module. S3, State Space Construction: Variational mode decomposition is performed on the preprocessed temporal state data to extract multi-scale modal features. Weighted state representations are generated through multi-head attention mechanism to construct system state vectors and action vectors. The reward value of the current state is calculated. Bayesian Q network uncertainty estimation is used to impose safety constraints on exploration actions. The policy network parameters are updated through deep reinforcement learning algorithm to output the optimal action decision. S4, dynamic control execution, transforms the optimal action decision into control commands and sends them to the execution equipment to adjust process parameters, while simultaneously feeding back the equipment execution status to the reinforcement learning model module; S5, Model Iteration Optimization and Anomaly Handling: The strategy network is continuously iterated and updated based on real-time data and equipment feedback. The reward function weight coefficient is adaptively adjusted by calculating the entropy of working condition information. An emergency mechanism is triggered and an alarm is sounded when an abnormal situation occurs. S6, data storage and interaction, stores all data in the storage process and enables status monitoring, parameter adjustment and manual intervention through a human-machine interface.

[0028] Preferably, in step S3, variational mode decomposition decomposes the time series of each state variable into... Each intrinsic mode function (EMF) is used to extract the energy and center frequency features of each mode to construct a multi-scale feature vector. Multi-head attention encoding then concatenates the multi-scale feature vector by simultaneously calculating the attention weights of h attention heads to output a weighted state representation. Bayesian Q-network uncertainty estimation is performed in the DQN algorithm using... The variance of the Q-value is calculated using Monte Carlo Dropout sampling. When the variance exceeds a safety threshold... The conservative strategy is triggered at the time; in step S5, the working condition information entropy is determined according to... The degree to which each state variable deviates from the baseline value is calculated in real time, and the weighting coefficients are dynamically adjusted accordingly. and Adjust the formula to , .

[0029] Preferably, step S5 further includes a periodic offline training and optimization phase, in which the policy network is trained offline in batches using stored historical operating data. During the training process, a model-independent meta-learning strategy is introduced, and inner loop gradient updates and outer loop meta-gradient accumulations are performed on historical data subsets corresponding to multiple process scenarios to obtain cross-scenario universal initial network parameters. : ; in, For the outer loop learning rate, For batch quantity in the process scenario, This represents the training task corresponding to the i-th process scenario. For the network after the internal loop update, For the task The loss function is applied; when the system is deployed to a new process scenario, the loss function is applied. Rapid adaptation can be achieved by making a small number of interactive fine-tunings as initial parameters, with no more than 50 fine-tuning steps.

[0030] Example 2: like Figures 4-6 As shown, this embodiment is applied to a municipal wastewater treatment plant using the A² / O process, with a designed treatment capacity of 50,000 tons / day. The influent water quality fluctuates significantly, and traditional PID control suffers from unstable water quality compliance and high energy consumption. The present invention employs a dynamic control system and method for water environment treatment processes based on reinforcement learning. The specific implementation process is as follows: S1: Construct the control system described in this invention. Data acquisition modules are deployed at the inlet, anaerobic tank, anoxic tank, aerobic tank, outlet, and each equipment room. Install online COD monitors, DO sensors, online ammonia nitrogen monitors, power sensors, etc., with a sampling frequency set to 2 minutes / time. Set the water quality standard to GB18918-2002-2025 revised Class A standard (COD≤50mg / L, NH3-N≤5mg / L, TP≤0.1mg / L, SS≤10mg / L). The baseline energy consumption for equipment operation is 1.2kWh / m³. The reinforcement learning model uses the DQN algorithm, with initial weight coefficients... The learning rate is 0.001 and the discount factor is 0.95. The Siemens S7-1500 series PLC system is selected as the control core, and the water quality analyzer is from the German brand E+H to ensure the stability of data acquisition and command issuance.

[0031] S2: The data acquisition module collects in real time the influent COD, NH3-N, TP, and flow rate; the DO and pH values ​​of the anaerobic, anoxic, and aerobic tanks; the effluent COD, NH3-N, TP, and SS; and the operating power, speed, and dosage of the blower, pump, and dosing equipment. The collected data is denoised and normalized, outliers such as jump data caused by sensor malfunctions are removed, and the processed data is transmitted to the reinforcement learning model module and the data storage module.

[0032] S3: The reinforcement learning model module constructs a state space vector S = [influent COD, influent NH3-N, aerobic tank DO, effluent COD, blower power, pump speed], and an action space vector A = [blower aeration rate, pump speed, chemical dosage, sludge discharge cycle]. The bi-objective reward function is R = 0.6·R1 + 0.4·R2, where R1 is the water quality compliance reward. When all effluent indicators meet the standards, R1 = 10. When an indicator exceeds the standard, R1 decreases by 2 for every 10% exceedance, and R1 = -20 when the exceedance is more than 50%. R2 is the energy consumption reward. When the actual energy consumption is lower than the baseline energy consumption of 1.2 kWh / m³, R2 increases by 3 for every 10% decrease. When the actual energy consumption is higher than the baseline energy consumption, R2 decreases by 2 for every 10% increase. The model continuously trains and optimizes the strategy network through real-time interaction with the process. After 15 days of training, the model reward value stabilizes, enabling it to output stable optimal action decisions.

[0033] S4: The dynamic control module transforms the optimal action decision output by the model into control commands and sends them to each execution device. For example, when the COD concentration of the influent increases, the model outputs commands to increase the aeration rate of the blower and increase the speed of the water pump to ensure that the DO in the aerobic tank is maintained at 2-4 mg / L and improve the organic matter removal efficiency. When the energy consumption is higher than the benchmark value, the model outputs commands to reduce the aeration rate of the blower and optimize the speed of the water pump to reduce energy consumption while ensuring that the water quality meets the standards. The dynamic control module provides real-time feedback on the execution status of the equipment, and the model continuously optimizes its decisions based on the feedback data.

[0034] S5: After operating the control system and method of this invention for 30 days, compared with the traditional PID control, the energy consumption decreased from 1.2 kWh / m³ to 0.9 kWh / m³, a reduction of 25%; the effluent quality compliance rate increased from 88% to 99.2%, with core indicators such as COD, NH3-N, and TP all stably meeting the Class A standard; the equipment operation stability was improved, and the manual operation and maintenance cost was reduced by 30%, completely solving the pain points of rigidity and inability to simultaneously achieve multiple objectives in traditional PID control, and achieving the expected technical effect.

[0035] Example 3: This embodiment is applied to an industrial wastewater treatment plant. The wastewater being treated is chemical wastewater with large fluctuations in influent pollutant concentrations, making traditional PID control difficult to adapt. The control system and method of this invention are adopted, and the specific implementation process is as follows: S1: The water quality standard is set as GB 18918-2002 2025 revised Class A standard, and the equipment operating baseline energy consumption is 1.5 kWh / m³; the reinforcement learning model adopts the PPO algorithm, with initial weight coefficients... The learning rate was 0.002 and the discount factor was 0.9. The sampling frequency of the data acquisition module was set to 1 min / time, focusing on collecting COD, BOD5, TN, TP of the influent and DO and pH values ​​of each reaction tank. The equipment operation data focused on collecting the energy consumption parameters of the aeration equipment and the dosing equipment.

[0036] S2: The reinforcement learning model learns and adapts to the regulation rules of chemical wastewater based on the collected industrial wastewater quality fluctuation data, and optimizes action decisions (dosage, aeration rate, reaction time, etc.). In view of the large fluctuation of COD concentration in chemical wastewater, the model dynamically adjusts the dosage and aeration rate to ensure that the COD of the effluent meets the standard stably, while optimizing equipment operating parameters and reducing energy consumption.

[0037] S3: After 20 days of operation, energy consumption decreased from 1.5 kWh / m³ to 1.05 kWh / m³, a reduction of 30%; the effluent quality compliance rate increased from 85% to 99.5%, achieving the dual goals of "water quality compliance + lowest energy consumption". It is adapted to the characteristics of large fluctuations in industrial wastewater quality, eliminating the need for frequent manual parameter adjustments and significantly improving treatment efficiency and stability.

[0038] The above embodiments are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic control system for water environment treatment processes based on reinforcement learning, characterized in that, include: The data acquisition module is used to collect real-time status data of the entire water environment treatment process, including water quality status data, equipment operation data, and environmental disturbance data. Water quality status data includes COD, BOD5, NH3-N, TN, TP, DO, pH value, and SS at the inlet and outlet. Equipment operation data includes the operating power, speed, and valve opening of water pumps, blowers, aeration equipment, and chemical dosing equipment. Environmental disturbance data includes influent flow rate, influent concentration fluctuation, and ambient temperature. The reinforcement learning model module includes a state space construction unit, an action space construction unit, a bi-objective reward function unit, and a policy optimization unit. The state space construction unit is used to normalize the collected state data, extract multi-scale modal features of the temporal state data through variational mode decomposition, and then encode the weighted state representation through a multi-head attention mechanism to construct the system state vector. The action space construction unit is used to define process controllable parameters as action vectors; the bi-objective reward function unit is used to construct the composite reward function. , This is a reward item for meeting water quality standards. Energy consumption optimization reward items, , The weighting coefficients and And based on the entropy of working condition information, and Adaptive adjustments are made; the policy optimization unit employs a deep reinforcement learning algorithm and uses Bayesian Q-network uncertainty estimation to impose safety constraints on the exploration actions, outputting the optimal action decision. The dynamic control module is used to receive the optimal action decision and convert it into control commands to be sent to each execution device, while feeding back the execution status of the device to the reinforcement learning model module; The data storage and interaction module is used to store various types of real-time data, training data, and historical data, and provides a human-computer interaction interface.

2. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 1, characterized in that, The data acquisition module adopts a distributed deployment, with sensors and smart instruments set up in each process section. The sampling frequency is 1 to 5 minutes each time. Data transmission adopts the OPC UA protocol. Noise data is filtered through edge nodes and multi-device timestamp synchronization is achieved at the millisecond level. The pre-processed data is synchronously sent to the reinforcement learning model module and the data storage and interaction module. The execution equipment includes water pumps, blowers, aeration equipment, dosing equipment, sedimentation tanks, and membrane modules. The control core adopts a PLC system.

3. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 1, characterized in that, The state space construction unit includes a variational mode decomposition subunit, used to perform variational mode decomposition on each state variable in the time-series state data, and to perform variational mode decomposition on the j-th state variable x. j (t), solve the constrained variational model: ; ; in, For the k-th eigenmode function obtained from the decomposition, Let be the center frequency of the k-th mode. To preset the number of decomposition modes, For the Dirac function, This represents the convolution operation. This involves taking the partial derivative with respect to time and extracting the energy characteristics of each mode. and center frequency This constitutes the multi-scale feature vector of the state variable; after concatenating the multi-scale feature vectors of all state variables, the input is given to the multi-head attention encoding subunit.

4. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 3, characterized in that, The multi-head attention encoding subunit is transformed by a linear transformation matrix. Map the input multi-scale feature vector to a query matrix Q, a key matrix K, and a value matrix V, and compute the scaled dot product attention: ; in, The scaling factor is the column dimension of matrix K; multiple attention heads are used for parallel computation, and the output of the i-th attention head is... ,in Let be the projection matrix of the i-th attention head; after concatenating the outputs of all attention heads, a weighted state representation is obtained through linear transformation. : ; Where h is the number of attention heads. To output the projection matrix; The final system state vector is obtained by concatenating it with the original normalized state vector.

5. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 1, characterized in that, The policy optimization unit includes a Bayesian uncertainty estimation subunit. When the deep reinforcement learning algorithm is DQN, a Monte Carlo Dropout mechanism is introduced into the Q-network to perform the same state-action pair estimation. After a random forward propagation, calculate the mean and variance of the Q-values: ; ; in, Output the Q-value for the m-th random forward propagation. The number of Dropout samples; the policy optimization unit adopts an uncertainty-guided exploration strategy when selecting actions: ; in, For uncertainty exploration coefficients; when Exceeding the preset safety threshold At this time, the strategy optimization unit abandons the current exploration action and instead executes a conservative action based on process safety rules to avoid excessive effluent or equipment failure due to exploration with high uncertainty.

6. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 1, characterized in that, The dual-objective reward function unit includes an information entropy adaptive weight adjustment subunit, which calculates the working condition information entropy H(t) in real time: ; ; in, The number of state variables involved in the entropy calculation. Let be the normalized value of the deviation between the current value and the reference value of the i-th state variable. The current value, The baseline value is used; the weighting coefficients are dynamically adjusted based on the entropy of the operating condition information. ; ; in, As the initial weight for water quality compliance, To adjust the coefficient, Use the reference entropy value; when Higher than When the time indicates an increase in operational complexity, the time will automatically increase. To ensure water quality safety; when Below When the operating condition is stable, the speed will automatically decrease. The focus is on energy consumption optimization.

7. The dynamic control system for water environment treatment process based on reinforcement learning according to claim 1, characterized in that, When executing action decisions, the dynamic control module sets safety constraint ranges for each action parameter. When an action parameter in the optimal action decision exceeds the rated operating range of the corresponding execution device, the action parameter is automatically truncated to a safe boundary value. The dynamic control module also sets an emergency control mechanism. When the effluent water quality index exceeds 1.5 times the standard threshold for two consecutive sampling cycles or the equipment operating parameters exceed the safe range, an emergency command is triggered to switch each execution device to a preset safe operating state and issue an alarm signal.

8. A method for a dynamic control system of a water environment treatment process based on reinforcement learning according to any one of claims 1-7, characterized in that, Includes the following steps: S1, System initialization, setting water quality standards, equipment operating baseline energy consumption, and initial parameters of the reinforcement learning model, and completing the debugging of each module and calibration of equipment parameters; S2, real-time data acquisition and preprocessing: The data acquisition module collects the status data of the entire process, performs noise reduction, normalization and outlier removal preprocessing, and then transmits it to the reinforcement learning model module and the data storage and interaction module. S3, State Space Construction: Variational mode decomposition is performed on the preprocessed temporal state data to extract multi-scale modal features. Weighted state representations are generated through multi-head attention mechanism to construct system state vectors and action vectors. The reward value of the current state is calculated. Bayesian Q network uncertainty estimation is used to impose safety constraints on exploration actions. The policy network parameters are updated through deep reinforcement learning algorithm to output the optimal action decision. S4, dynamic control execution, transforms the optimal action decision into control commands and sends them to the execution equipment to adjust process parameters, while simultaneously feeding back the equipment execution status to the reinforcement learning model module; S5, Model Iteration Optimization and Anomaly Handling: The strategy network is continuously iterated and updated based on real-time data and equipment feedback. The reward function weight coefficient is adaptively adjusted by calculating the entropy of working condition information. An emergency mechanism is triggered and an alarm is sounded when an abnormal situation occurs. S6, data storage and interaction, stores all data in the storage process and enables status monitoring, parameter adjustment and manual intervention through a human-machine interface.

9. The method for dynamic control of water environment treatment process based on reinforcement learning according to claim 8, characterized in that, In step S3, variational mode decomposition decomposes the time series of each state variable into... Each intrinsic mode function (EMF) is used to extract the energy and center frequency features of each mode to construct a multi-scale feature vector. Multi-head attention encoding then concatenates the multi-scale feature vector by simultaneously calculating the attention weights of h attention heads to output a weighted state representation. Bayesian Q-network uncertainty estimation is performed in the DQN algorithm using... The variance of the Q-value is calculated using Monte Carlo Dropout sampling. When the variance exceeds a safety threshold... Conservative strategies are triggered at certain times; In step S5, the working condition information entropy is based on The degree to which each state variable deviates from the baseline value is calculated in real time, and the weighting coefficients are dynamically adjusted accordingly. and Adjust the formula to , .

10. The method for dynamic control of water environment treatment process based on reinforcement learning according to claim 8, characterized in that, Step S5 also includes a periodic offline training and optimization phase, which uses stored historical operating data to perform offline batch training on the policy network. During the training process, a model-independent meta-learning strategy is introduced, and inner loop gradient updates and outer loop meta-gradient accumulations are performed on historical data subsets corresponding to multiple process scenarios to obtain cross-scenario generalized initial network parameters. : ; in, For the outer loop learning rate, For batch quantity in the process scenario, This represents the training task corresponding to the i-th process scenario. For the network after the internal loop update, For the task The loss function is applied; when the system is deployed to a new process scenario, the loss function is applied. Rapid adaptation can be achieved by making a small number of interactive fine-tunings as initial parameters, with no more than 50 fine-tuning steps.