Sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization

By combining reinforcement learning with particle swarm optimization algorithms, designing reward functions and dual experience pool architecture, the problems of low data utilization efficiency and difficulty in multi-objective coordination in the control of drug delivery in sewage treatment are solved, real-time dynamic optimization and multi-objective balance are achieved, and the adaptability and efficiency of the sewage treatment system are improved.

CN120802618APending Publication Date: 2025-10-17ZHEJIANG UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510944664.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing sewage treatment drug dosage control methods have problems such as low data utilization efficiency, difficulty in multi-objective coordination, local optimal traps and broken prediction-control links, making it difficult to achieve real-time dynamic optimization and multi-objective balance.

Method used

Combining reinforcement learning with particle swarm optimization algorithms, a reward function is designed and a dual experience pool architecture is constructed. Through hybrid sampling and real-time updating of particle inertia weights, dynamic optimization of drug delivery is achieved, and closed-loop feedback control is performed in combination with time series prediction information.

Benefits of technology

It achieves a balance between global search capability and real-time dynamic adjustment, improves sample utilization, enhances adaptability to water quality fluctuations, and achieves a multi-objective dynamic balance of turbidity compliance, cost savings and drug synergy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802618A_ABST
    Figure CN120802618A_ABST
Patent Text Reader

Abstract

The invention discloses a sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization, and the method achieves the precise control of turbidity and the optimization of drug cost through the dynamic adjustment of particle swarm parameters and the fusion of a reinforcement learning strategy. According to the technical scheme, the method comprises the following steps: combining a PSO algorithm with a DDPG algorithm to construct a drug control model; designing a reward function; the empirical data with the state coverage rate higher than a first preset threshold value are stored in a full-reservation empirical pool, the empirical data with the current policy access larger than a second preset threshold value are reserved in a first-in first-out mode through a policy empirical pool, and mixed sampling is carried out according to a preset proportion; training a drug control model; and inputting the current turbidity value, the drug delivery amount of the last time step, the water inlet flow, the parameters of the water quality detection camera and the future turbidity value into a trained drug control model, and outputting the current drug delivery amount. Through prediction-control closed-loop integration, a raw water turbidity prediction result is fed back to the optimizer, and the method is suitable for real-time adaptive control of a sewage treatment system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent control of sewage treatment, and particularly relates to a sewage treatment drug feeding control method based on reinforcement learning and particle swarm optimization (PSO). BACKGROUND

[0002] In the sewage treatment process, dynamic adjustment of drug feeding quantity is a core problem for ensuring that the effluent turbidity meets the standard and controlling the operation cost. Traditional control methods mainly fall into four categories:

[0003] (1) PID-based control method: relying on fixed parameter adjustment, it is difficult to adapt to water quality fluctuations and working condition changes, and is only suitable for steady-state scenarios.

[0004] (2) Predictive control method: it needs high-precision prediction model support, and model inaccuracy can easily lead to control failure, and the calculation complexity is high.

[0005] (3) Neural network control method: although it can learn nonlinear relationships, it has the problem of not extending optimization information, and lacks a knowledge transfer mechanism between control cycles.

[0006] (4) Intelligent optimization method (such as genetic algorithm, particle swarm algorithm): strong global search capability, but poor real-time performance, unable to interactively optimize with dynamic environment.

[0007] The existing technology has the following limitations:

[0008] (1) Low data utilization efficiency: the traditional experience replay mechanism discards historical data, resulting in insufficient sample utilization and limited strategy learning speed.

[0009] (2) Difficulty in multi-objective coordination: existing reward function design mostly focuses on a single objective (such as turbidity or cost), lacking comprehensive consideration of drug synergistic effects and long-term impacts.

[0010] (3) Local optimal trap: single algorithm framework (such as pure reinforcement learning) is prone to local optimal, and parameter adjustment relies on human experience.

[0011] In recent years, the integration of reinforcement learning and swarm intelligence algorithms has become a new direction for solving complex control problems. For example, the deep deterministic policy gradient (DDPG) algorithm performs well in continuous action space control, but its exploration efficiency is low and convergence speed is slow; the particle swarm algorithm (PSO) can globally optimize, but lacks dynamic environment adaptability. How to combine the advantages of the two and build a hybrid framework with real-time decision-making and global optimization has become a technical problem to be solved.

[0012] In addition, the existing sewage treatment control system adopts independent prediction and decision module, the prediction result does not participate in the optimization of the control strategy deeply, and the 'prediction-control' link is broken. An intelligent control method capable of fusing time series prediction information, dynamically adjusting the optimization target and realizing closed-loop feedback is urgently needed. SUMMARY

[0013] The present application aims at the defects of the existing sewage treatment drug feeding control method, and provides a sewage treatment drug feeding control method based on reinforcement learning and particle swarm optimization.

[0014] The purpose of the present application is achieved by the following technical scheme: a sewage treatment drug feeding control method based on reinforcement learning and particle swarm optimization, comprising the following steps:

[0015] (1) combine the PSO algorithm with the DDPG algorithm to construct a drug control model; the PSO algorithm comprises a plurality of particle swarms composed of actor networks; the DDPG algorithm comprises an actor network and a critic network;

[0016] (2) design a reward function: comprehensively consider turbidity penalty items, drug cost items and synergistic effect items;

[0017] (3) realize a double experience pool architecture: use a full reservation experience pool to store experience data with a state coverage higher than a first preset threshold, a strategy experience pool reserves experience data accessed by the current strategy greater than a second preset threshold in a first-in-first-out manner, and mixes sampling in a preset proportion;

[0018] (4) train the DDPG algorithm model based on the data sampled in step (3), and adaptively update the particle inertia weight and individual / social learning factor according to the real-time reward function; inject the trained actor network parameters of the DDPG into the particles ranked k after the fitness in the PSO population, to obtain a trained drug control model;

[0019] (5) input the current state, i.e. the current turbidity value, the drug feeding amount at the last time step, the inflow, the water quality detection camera parameters and the future turbidity value into the trained drug control model, and output the current drug feeding amount.

[0020] The present application has the following beneficial effects:

[0021] (1) efficient and synergistic optimization: through the mixed framework of PSO and DDPG, the global search ability and real-time dynamic adjustment demand are effectively balanced, and the local optimal trap is avoided.

[0022] (2) sample utilization rate is improved: the double experience pool architecture combined with the mixed sampling strategy significantly improves the reuse rate of historical key experience and accelerates the strategy convergence.

[0023] (3) Multi-objective Synergistic Control: The reward function design is integrated with future turbidity prediction, achieving a dynamic balance of turbidity compliance, cost savings, and drug synergistic effects.

[0024] (4) System Robustness Enhancement: Closed-loop feedback mechanism between prediction model and optimizer, improving adaptability to water quality fluctuations and operating condition changes. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0026] Figure 1 is a double experience replay framework;

[0027] Figure 2 is a drug dosing method flowchart based on reinforcement learning particle swarm algorithm;

[0028] Figure 3 is a drug dosing control model training flowchart based on reinforcement learning and particle swarm optimization;

[0029] Figure 4 is a hardware structure schematic diagram provided by the embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions of the present application will be described in detail below in combination with the drawings and embodiments. In the case of no conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0031] As shown in Figure 1 , a drug dosing control method based on reinforcement learning particle swarm algorithm comprises the following steps:

[0032] (I) Establish a water turbidity prediction model to predict future turbidity values in real time

[0033] (Note: This part briefly explains the relevance of the prediction model and the drug control model)

[0034] Step 1: Real-time acquisition of raw water turbidity, inflow, historical drug dosing amount (aluminum sulfate, Polymer), water quality monitoring camera parameters (such as floc size, brightness, mass fraction, etc.) and other time series data through sensors, data preprocessing including:

[0035] (1) Missing value filling: Linear interpolation method is used to complete the missing data caused by equipment failure;

[0036] (2) Time alignment: According to the time delay difference between dosing time and turbidity monitoring, the data is time-shifted in the time dimension.

[0037] (3) Normalization processing: Call the standardization function (such as MinMaxScaler) to normalize the input features.

[0038] Step two: input the preprocessed data into the trained time series prediction model (such as LSTM, TCN, etc. using the data set in Table 1), and output the predicted value of raw water turbidity in the next 1 hour. The predicted value is used as the state input parameter of the drug control model.

[0039] (II) Drug dosing control method based on reinforcement learning and particle swarm algorithm

[0040] Step one: build a reinforcement learning agent framework, including Actor network, Critic network, target Actor network and target Critic network

[0041] (1) Definition of state space:

[0042] Current turbidity value, predicted turbidity value, inflow, historical drug dosing amount (aluminum sulfate, Polymer), water quality monitoring camera parameters (such as maximum size of flocculation, sphericity).

[0043] (2) Definition of action space:

[0044] The action output is the adjustment value of the dosing amount of aluminum sulfate and Polymer.

[0045] The raw water turbidity prediction data set used in this embodiment is from a thermal power plant in Ningbo City. Its filtered water treatment system mainly uses Yaojiang surface water, municipal sewage and rainwater as water source, which is pumped to the coagulation tank after being treated by large industrial water and water company, and aluminum sulfate (coagulant) is added. After uniform mixing by fast mixing, small colloidal floc flows into the colloidal floc tank. The raw water containing small colloidal floc adds POLYMER (coagulant aid) in the colloidal floc tank, and after slow mixing, it forms a precipitated colloidal floc flow into the sedimentation tank at different speeds. The clarified water after sedimentation treatment is overflowed to the sand filter tank through the pipeline, and is filtered by gravity, and the filtered water that meets the inspection standard is delivered to other production processes. As can be seen from Table 1, the output is Turbidity (sedimentation tank turbidity), and the rest are inputs.

[0046] Table 1: Content of raw water data set

[0047]

[0048] (3) Reward function design:

[0049] The reward function can be decomposed into a turbidity penalty term, a drug cost term and a synergistic effect term.

[0050] The turbidity penalty term is formulated as:

[0051] R 浊度 = R 当前浊度 + R 未来浊度

[0052]

[0053] where R 当前浊度 is the current turbidity penalty term, R 未来浊度 is the future turbidity penalty term, Z represents the current turbidity value, Z' is the future turbidity value predicted by the time series prediction model. α represents the threshold value of turbidity, which is set to 2 NTU according to the water quality standard. λ1 and λ3 represent the over-standard penalty coefficients, which need to be much larger than other terms to ensure that the threshold is met first. λ2 and λ4 represent the excessive reduction penalty coefficients, which prevent drug waste. β represents the allowed turbidity reduction tolerance, which is set to 0.3 NTU.

[0054] The design goal is to penalize high-cost drugs polymer, encourage the use of less polymer, but need to ensure the reasonable use of aluminum sulfate, which is much less than polymer, but aluminum sulfate has a greater impact on turbidity. The drug cost term is formulated as:

[0055] R 成本 = -(polymer 成本 · polymer 投放 + Al2(SO4) 3成本 · Al2(SO4) 3投放 )

[0056] where polymer 成本 is the unit cost of Polymer, polymer 投放 is the dosage of Polymer, Al2(SO4) 3成本 is the unit cost of aluminum sulfate, Al2(SO4) 3投放 is the dosage of aluminum sulfate.

[0057] The synergistic effect term R 协同 : When the dosage ratio of aluminum sulfate to Polymer exceeds the dosage ratio threshold, a penalty is given.

[0058]

[0059] where γ (usually set to 3) represents the allowed dosage ratio threshold of Al2(SO4)3 / polymer, ∈ is a very small value to prevent division by zero error, and λ5 represents the ratio imbalance penalty coefficient.

[0060] If there are A, P, C three drugs, the formula is as follows:

[0061]

[0062] wherein γ AP , γ AC and γ PC respectively represent the threshold of the allowed ratio of drug A / P, the threshold of the allowed ratio of drug A / C and the threshold of the allowed ratio of drug P / C, λ5, λ6 and λ7 respectively represent the penalty coefficient of the imbalance of the ratio of drug A and drug P, the penalty coefficient of the imbalance of the ratio of drug A and drug C and the penalty coefficient of the imbalance of the ratio of drug P and drug C.

[0063] It should be noted that the unit costs of the three drugs A, P and C are A>P>C, and the allowed ratios of the three drugs A, P and C are A / P, A / C and P / C in the formula.

[0064] The R 协同 calculation formula for two and three drugs is given above. 协同 The R 协同 calculation formula for more than three drugs can be designed by those skilled in the art, which is not shown one by one here.

[0065] Step two: double experience pool architecture and mixed sampling

[0066] (1) Experience pool division (such as Figure 1 ):

[0067] Full retention experience pool D full : stores experience data with a high state coverage rate (such as at least 80% of the state space is covered or the single state access frequency exceeds 10% of the total training steps), and the capacity is fixed as N. The experience data includes the sewage turbidity value, the drug dosage, the inflow, the water quality detection camera parameters, the action output by the Actor network of the DDPG at each time step and all drug dosing strategies generated by the PSO.

[0068] Strategy experience pool D policy : stores experience data in the last L time steps (L=βM, β is the retention ratio coefficient) according to the first-in-first-out principle, and the capacity is M (M=αN, α is the retention ratio coefficient). The experience data includes the sewage turbidity value, the drug dosage, the inflow, the water quality detection camera parameters, the action output by the Actor network of the DDPG at each time step and all drug dosing strategies generated by the PSO.

[0069] The experience dataset records the water treatment process parameters of Ningbo Thermal Power Plant, including time, industrial water / middle water / rainwater flow, drug dosage (aluminum sulfate, polymer), 20 items of colloidal particle shape characteristics (diameter, sphericity, fractal dimension, etc.), and sedimentation tank turbidity, which are used to predict raw water turbidity.

[0070] (2) Mixed sampling strategy:

[0071] Randomly sample from R1 and R2 with proportions β and 1-β each time (e.g., sample with a 6:4 ratio), combine them into a training batch, and improve the utilization of experience.

[0072] Step three: collaborative optimization of reinforcement learning and particle swarm algorithm

[0073] (1) Algorithm initialization:

[0074] Initialize the Actor network (policy network) and Critic network (value network) and randomly generate the initial weights. Define a full-retention experience pool (Replay Buffer) to store interaction data.

[0075] Initialize a particle swarm containing K particles estimated by the actor network π ri , improve the number of strategies k, and each particle π ri represents a set of candidate dosing amounts (A i , B i ) (here A is polymer and B is aluminum sulfate). The particle position x i = [A i , B i ], and the fitness f i =R(x i ) is calculated by the reward function.

[0076] Initialize a noise generator and a random number generator.

[0077] Set the inertia weight ω, acceleration constants c1 and c2 in the PSO parameters, and the maximum number of iterations T max . Set the discount factor γ, learning rate α Actor and α Critic in the RL parameters. Set the collaborative weight η for the fusion of RL reward and PSO fitness.

[0078] The algorithm hyperparameter settings are divided into two parts: DDPG parameter settings and PSO parameter settings. The PSO parameters are shown in Table 2:

[0079] Table 2: PSO parameters

[0080]

[0081] DDPG parameters are shown in Table 3:

[0082] Table 3: DDPG parameters

[0083]

[0084] (2) Environment interaction and action generation:

[0085] At time step t, the current state s t is observed, including: current turbidity value, predicted turbidity value, inflow flow, camera parameter value, and last time step drug dosage.

[0086] At the RL module input current state s t , according to the current policy noise to select action output action a t =[ΔA,ΔB]. In the PSO module, the highest fitness particle x best =[A best ,B bets ] is selected from the subgroup.

[0087] Combine RL action and PSO global optimal solution:

[0088] A t =A t-1 +ΔA+η(A best -A t-1 )

[0089] B t =B t-1 +ΔB+η(B best -B t-1 )

[0090] Where A t represents the actual dosage of polymer (Polymer) added at the current time step t, b t represents the actual dosage of aluminum sulfate (Al2(SO4)3) added at the current time step t, both of which are determined by the adjustment amount (ΔA,ΔB) output by the reinforcement learning module and the global optimal solution (A best ,B best ) provided by the particle swarm optimization module. Specifically, ΔA and ΔB are the incremental correction values calculated by the Actor network according to the current state s t , while A best and B best are the optimal dosage combination represented by the highest fitness particle in the particle swarm. η is used to fuse the synergistic weight of RL reward and PSO fitness.

[0091]

[0092] where σ is the maximum weight value (set to 0.5 here), sen is the response sensitivity (set to 0.2), r t is the current reward value, r ave is the historical average reward (sliding window 100-step mean).

[0093] According to A t and B t , adjust the drug dosage and apply the time series prediction model mentioned above to predict the turbidity in the next hour.

[0094] Calculate the reward:

[0095] r t = R 浊度 + R 成本 + R 协同

[0096] Store {(s t , a t , r t , s t+1 )} in the full- retention experience pool D full . Sort the cumulative returns of each particle and store the optimal k strategies in the strategy experience pool D policy .

[0097] (3) Reinforcement learning network update:

[0098] Randomly sample η full · n samples from D s , randomly sample (1-η s ) n sample experiences from D policy , and combine them into a batch of data {(s i , a i , r i , s i+1 )}, update the Critic network, and calculate the target Q value:

[0099] y i = r i + γ · Q'(s i+1 , μ'(s i+1 ))

[0100] where y i is the target Q value, indicating the expected cumulative return of performing action a i in state s i ; r i is the immediate reward, reflecting the state-action pair (s i , a iimmediate reward; Q' is the target Critic network, μ' is the target Actor network; γ ∈ [0, 1] is the discount factor to weigh the importance of current and future rewards. Minimize the Critic loss function:

[0101]

[0102] where N is the total number of batch sampled samples; Q(s i ,a i, ) is the action value estimation of the Critic network for state s i and action a i ; s i represents the state of the i-th sample; a i represents the action of the i-th sample. Maximize the expected return by policy gradient:

[0103]

[0104] where τ is the target network update rate (soft update coefficient), which is a preset hyperparameter (set to 0.001), used to control the speed of updating the parameters of the target network; θ is the weight parameter of the current Actor or Critic network. Update the target Actor and Critic networks:

[0105] θ' <- τθ + (1 - τ)θ'

[0106] (4) Particle swarm optimization update

[0107] For each particle i, calculate its fitness f i :

[0108] f i = R 浊度 (x i ) + R 成本 (x i ) + R 协同 (x i )

[0109] f pbest,i is the fitness of the individual optimal of particle i, if f i > f pbest,i , update the individual optimal position x pbest,i of the particle <- x i ; f gbest is the fitness of the swarm optimal, if f i > f gbest , update the swarm optimal position x gbest <- x i .

[0110] Adjusting PSO parameters according to RL reward dynamics:

[0111]

[0112] where c1 base and c2 base are the baseline values of c1 and c2 (set to 1.4 and 1.7 in Table 1); r ave is the historical average reward (sliding window of 100 steps mean); p is the PSO parameter adjustment coefficient, used to control the degree of influence of the immediate reward r t on the adjustment range of the acceleration factor.

[0113] Improving the Sigmoid function to map the real-time reward value r t to the inertia weight interval [ω min , ω max ] (set to 0.4 and 0.9 in Table 1), achieving adaptive updating of particle inertia weight according to real-time reward signal:

[0114]

[0115] The formula is designed as a monotonically decreasing Sigmoid function, dynamically adjusting the inertia weight ω by the real-time reward r t . When the control system obtains negative reward (water quality not up to standard or excessive drug consumption), ω automatically increases to close to ω max , enhancing the global exploration ability of the particle swarm, helping the algorithm to jump out of the local optimum. When obtaining positive reward (water quality up to standard and reasonable drug consumption), ω decreases to close to ω min , strengthening the local development ability and improving the convergence accuracy. The kap parameter controls the sensitivity of the weight to the change of the reward, ensuring the continuity of the search strategy of the algorithm under stable working conditions and rapid response when the working conditions change suddenly.

[0116] (5) Synergistic mechanism:

[0117] Injecting the Actor network parameters of DDPG into the particles ranked k in the PSO population after fitness, improving the diversity of the population. The fitness ranking of particles is determined in descending order of the current reward function value.

[0118] Step four: Real-time control and decision making (e.g. Figure 2 , Figure 3 )

[0119] Input the predicted future turbidity value and data set into the sewage treatment drug control model based on reinforcement learning and particle swarm optimization, and output the adjustment instructions for the dosage of aluminum sulfate and Polymer.

[0120] Display the predicted turbidity and recommended dosage through the visualization platform, supporting manual review or automatic execution of control instructions.

[0121] (III) System integration and deployment

[0122] (1) Hardware deployment:

[0123] Integrate turbidity sensor, flow meter, dosing equipment controller with algorithm server in real-time communication in the sewage treatment control system.

[0124] (2) Software modules:

[0125] Prediction module: run raw water turbidity time series prediction model, real-time output future turbidity value.

[0126] Control module: execute sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization, dynamically optimize drug delivery strategy.

[0127] Interactive interface: display real-time data, prediction curve, control instruction and history record.

[0128] Corresponding to the foregoing embodiment of the sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization, the application also provides an embodiment of a drug delivery control device based on reinforcement learning and particle swarm optimization.

[0129] Referring to Figure 4 , the drug delivery control device based on reinforcement learning and particle swarm optimization provided by the embodiment of the application comprises one or more processors for implementing the sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization in the foregoing embodiment.

[0130] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0131] The embodiment of the drug delivery control device based on reinforcement learning and particle swarm optimization can be applied to any device with data processing capability, which can be a device or apparatus such as a computer. The device embodiment can be realized by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logical device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory for running by the processor of the device with data processing capability. From the hardware level, as shown in Figure 4 The processor, memory, network interface, and non-volatile memory shown in Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown in

[0132] The implementation process of the functions and roles of each unit in the above device is specifically described in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0133] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the part of the method embodiment. The above described device embodiment is only schematic, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. According to the actual needs, part or all of the modules can be selected to achieve the purpose of the present application. Those skilled in the art can understand and implement without creative labor.

[0134] The embodiment of the present application also relates to a computer readable storage medium, which stores a computer program, and when the program is executed by a processor, the aforementioned medicine dispensing control method can be realized. The computer readable storage medium includes but is not limited to a fixed storage device (such as a hard disk, a solid state disk, a read-only memory or a random access memory) integrated in a computing device, a removable storage device (such as a plug-in hard disk, a smart memory card, a secure digital card, a micro memory card, a flash memory card or an optical disk) and a distributed storage system combined by a local storage device and a cloud storage service, wherein the distributed storage system supports synchronization and backup of data. The storage medium is used for non-transiently storing program codes of the medicine dispensing control method, related configuration parameters and intermediate data generated in the running process, and simultaneously realizes temporary caching of real-time data and persistent storage of historical records. It should be understood by those skilled in the art that any physical medium capable of carrying executable codes and read by a processor belongs to the category of the computer readable storage medium defined in the embodiment.

[0135] The above embodiment is only used for illustrating the design idea and characteristics of the present application, and its purpose is to enable those skilled in the art to understand the present application and to implement it, and the protection scope of the present application is not limited to the above embodiment. Therefore, any equivalent changes or modifications made according to the disclosed principles and design ideas of the present application are within the protection scope of the present application.

Claims

1. A sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization, characterized in that: The following steps are involved: (1) Combining the PSO algorithm with the DDPG algorithm to construct a drug control model; the PSO algorithm includes a particle swarm composed of several actor networks; the DDPG algorithm includes an actor network and a critic network; (2) Design reward function: comprehensive turbidity penalty term, drug cost term, and synergistic effect term; (3) Implementing a dual experience pool architecture: using a full-retention experience pool to store experience data with a state coverage rate higher than a first preset threshold, and a policy experience pool to retain experience data with a current policy access rate higher than a second preset threshold in a first-in-first-out manner, and mixed sampling according to a preset ratio; (4) Train the DDPG algorithm model based on the data sampled in step (3), and adaptively update the particle inertia weight and individual / social learning factor according to the real-time reward function; reversely inject the trained DDPG Actor network parameters into the k particles with the lowest fitness ranking in the PSO population to obtain the trained drug control model; (5) The current state, i.e., the current turbidity value, the drug dosage in the previous time step, the water flow rate, the water quality detection camera parameters, and the future turbidity value are input into the trained drug control model, and the current drug dosage is output.

2. The method according to claim 1, characterized in that Reward function r t The calculation formula is: r t =R 浊度 +R 成本 +R 协同 Among them, the turbidity penalty term R 浊度 Calculated based on current turbidity penalty and future turbidity penalty; drug cost item distinguishes the cost difference of each drug; synergistic effect item constrains the proportion of each drug release.

3. The method according to claim 2, characterized in that Current turbidity penalty term R 当前浊度 and is the future turbidity penalty term R 未来浊度 The calculation formula is: Among them, Z represents the current turbidity value, Z ′ represents the future turbidity value; α represents the turbidity threshold; λ1 and λ3 represent the penalty coefficients for exceeding the standard; λ2 and λ4 represent the penalty coefficients for excessive reduction; β represents the allowable turbidity reduction tolerance.

4. The method according to claim 1, wherein According to the real-time reward function r t Adaptively update particle inertia weights and individual / social learning factors: Among them, c1 is the individual learning factor, c2 is the social learning factor; c1 base and c2 base are the benchmark values ​​of c1 and c2 respectively; ρ is the PSO parameter adjustment coefficient; r ave is the historical average reward; ω is the inertia weight, and the kap parameter controls the sensitivity of the weight to the reward change. max and ω min are the upper and lower limits of the inertia weight ω respectively.

5. The method according to claim 1, wherein The drug control model outputs the current drug delivery amount based on the actions output by the Actor network and the PSO global optimal solution: IN t = Yes t-1 +ΔA+η(A best -IN t-1 ) Among them, A t is the actual dosage of a drug at the current time step t, ΔA is the action output by the Actor network; A best is the global optimal solution given by the PSO algorithm, and η is the weight; Among them, σ is the maximum weight value, sen is the response sensitivity, r t is the current reward value, r ave The historical average reward.

6. The method according to claim 1, characterized in that Future turbidity value acquisition includes: The current turbidity value, the amount of medicine added in the previous time step, the water flow rate and the water quality detection camera parameters are collected and preprocessed. The preprocessed data is input into the water turbidity prediction model, and the output is the future turbidity value.

7. The method according to claim 6, characterized in that The raw water turbidity prediction model is a time series prediction model.

8. The method according to claim 1, characterized in that The preprocessing includes missing value filling, time series alignment and normalization.

9. A drug delivery control device based on reinforcement learning and particle swarm optimization, characterized in that: It includes one or more processors for implementing a sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization as described in any one of claims 1-8.

10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, it is used to implement a sewage treatment drug delivery control method based on reinforcement learning and particle swarm optimization as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Visual identification-based alumen ustum detection dosing control method, system and equipment

    CN121635499A

  • Mine water treatment agent ratio intelligent recommendation method based on multi-objective optimization

    CN121983181A

  • Intelligent recommendation method for mine water treatment agent proportioning based on multi-objective optimization

    CN121983181B

  • Chemical Dosing System and Method for Water Treatment Facilities

    KR102998990B1