Double-capacity water tank control method, system and equipment based on multi-target reward Q learning and medium

Through the multi-objective reward Q learning method, the problems of insufficient adaptability and real-time performance of traditional PID control and reinforcement learning in the liquid level control of double-capacity water tanks are solved, and high-precision, low-complexity adaptive control is achieved, which is suitable for nonlinear and strongly coupled systems.

CN120631069APending Publication Date: 2025-09-12LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510772400.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional PID control has difficulty coping with dynamic changes in dual-capacity water tank level control. Existing multi-objective optimization methods rely on precise models and lack adaptability. The single-objective reward function in reinforcement learning easily leads to policy oscillation, the experience replay mechanism is inefficient, and Q learning is difficult to discretize effectively, affecting control stability and real-time performance.

Method used

A multi-objective reward Q-learning method is adopted to construct an adaptive control architecture through discretized mathematical model, multi-objective reward function, event-triggered priority experience replay and dynamic PID parameter adjustment to achieve coordinated optimization of liquid level tracking error, control quantity fluctuation and parameter stability.

Benefits of technology

It achieves high-precision real-time control in nonlinear and strongly coupled systems, reduces computational complexity and memory usage, improves learning efficiency and system stability, and adapts to dynamic working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005443326910000021
    Figure BDA0005443326910000021
  • Figure BDA0005443326910000022
    Figure BDA0005443326910000022
  • Figure BDA0005443326910000061
    Figure BDA0005443326910000061
Patent Text Reader

Abstract

The invention discloses a double-capacity water tank control method based on multi-target reward Q learning, and belongs to the technical field of machine learning and automatic control. The method comprises the following steps: firstly, establishing a discretization mathematical model of a double-capacity water tank liquid level control system, and designing a multi-target reward function; secondly, empirical data are stored only when the liquid level tracking error exceeds the limit or the control quantity suddenly changes, and sampling weights are distributed according to award absolute values; and finally, updating the Q table based on the time difference error, and performing closed-loop real-time control on the double-tank water tank according to the dynamically adjusted PID parameter. According to the method, a five-dimensional reward function is designed, and multi-index dynamic balance is realized through linear weighting; the empirical tuple is stored only when the liquid level tracking error exceeds the limit, and the quick response requirement of the dynamic time-varying scene of the double-capacity water tank is met; dimensionality reduction is performed on a continuous state space by adopting a non-uniform grading strategy, so that the calculation complexity is remarkably reduced; a perception-decision-execution integrated framework is constructed, and the dependence of a traditional control method on a mathematical model of the water tank is abandoned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning and automatic control, and specifically relates to a dual-capacity water tank control method, system, equipment and medium based on multi-objective reward Q learning. Background Art

[0002] In the field of industrial process control, dual-tank level control systems, due to their nonlinearity, time lag, and strong coupling, place extremely high demands on the dynamic adaptability and robustness of control algorithms. Traditional PID control relies on fixed parameter adjustments, making it difficult to cope with dynamic changes in tank level. This is especially true when valve openings fluctuate or tank parameters drift, often leading to problems such as regulation lag and large overshoot. While existing multi-objective optimization methods (such as ant colony algorithms, genetic algorithms, and fuzzy control) attempt to improve control performance, they generally suffer from strong model dependence. This requires the establishment of precise mathematical models (such as differential equations or transfer functions) for tank flow and level. This significantly degrades control performance when system parameters (such as tank cross-sectional area and fluid resistance) change due to scale deposition or equipment aging. Furthermore, methods like fuzzy control rely on manually designed rule tables, making it difficult to dynamically adapt to the multi-stage nature of tank level changes (from sudden change to regulation to steady state). The parameter tuning process is cumbersome and lacks adaptability.

[0003] In recent years, reinforcement learning-based control methods have shown promising potential in complex system control. However, when applied to dual-tank level control systems, these methods also face several pressing challenges. First, a single-objective reward function (e.g., optimizing only the level tracking error) cannot balance tracking accuracy, control smoothness, and parameter stability, leading to policy oscillation. For example, frequent adjustments to the control variable can exacerbate valve wear or cause system overshoot. Second, traditional experience replay mechanisms store large amounts of redundant data, which interferes with the learning process due to low-value information, resulting in slow convergence and difficulty meeting real-time control requirements. Third, the continuous state and action space of dual-tank systems makes it difficult to effectively discretize classical Q learning. The uniform binning strategy suffers from insufficient resolution in high-error regions and lacks physical constraints on PID parameters, making it prone to problems such as out-of-bounds differential coefficients, further impacting control stability.

[0004] Existing technologies in the multi-objective control of dual-capacity water tanks have problems such as model dependence, high manual intervention cost, and inefficient strategy optimization. There is an urgent need for a new control framework that combines adaptability, data efficiency, and physical constraints to achieve high-precision real-time control in complex industrial scenarios. Summary of the Invention

[0005] To address the shortcomings of the existing technology, this paper proposes a dual-capacity water tank control method based on Multi-Objective Reward Q-Learning (MOR-QL), which is implemented through the following steps:

[0006] The dual-capacity water tank control method based on multi-objective reward Q learning includes the following steps:

[0007] Step 1: Establish a discretized mathematical model of the dual-tank liquid level control system, and construct a three-dimensional continuous state space based on the liquid level tracking error, error integral, and previous time step error;

[0008] Step 2: Define the incremental adjustment action space of PID parameters and design a multi-objective reward function to collaboratively optimize the liquid level tracking error, control variable fluctuation, and parameter stability;

[0009] Step 3: Adopt an event-triggered priority experience replay strategy, store experience data only when the level tracking error exceeds the limit or the control amount changes suddenly, and assign sampling weights based on the absolute value of the reward;

[0010] Step 4: Update the Q table based on the time difference error and dynamically adjust the PID parameters;

[0011] Step 5: Substitute the adjusted PID parameters into the discrete PID formula to generate the control quantity, and perform closed-loop real-time control on the double-capacity water tank.

[0012] Preferably, in step 1, the discretized mathematical model of the dual-capacity water tank level control system is:

[0013]

[0014] The transfer function of the discretized system is:

[0015]

[0016] Among them, M1=T1T2, M2=T1+T2, T1=A1R1, T2=A2R2, H2(s) is the function of the liquid level in the lower water tank after Laplace transform in the complex frequency domain, q1 is the flow rate of the first water tank, A1 and A2 are the cross-sectional areas of the upper and lower water tanks respectively, R1 and R2 are the liquid resistances of valves 1 and 2 respectively, T is the sampling period, and z is the complex variable in the discrete system z transform.

[0017] Preferably, in step 1, the three-dimensional continuous state space is discretized by the following steps: performing a clipping operation on each state component, binning the clipped state component into n bins, generating a discrete state index, combining the three state components to obtain n 3 dimensional discrete state space.

[0018] Preferably, in step 2, the multi-objective reward function includes: error and control amount fluctuation reward r1, PID parameter change reward r2, control amount change reward r3, steady-state reward r4 and parameter out-of-bounds reward r5;

[0019] The steady-state reward r4 is: setting a continuous step threshold and a liquid level tracking error threshold. When the system's continuous steps reach the threshold and the absolute value of the liquid level tracking error of each step is less than the set threshold, a steady-state reward is given to the agent.

[0020] Preferably, in step 3, the priority experience replay strategy includes:

[0021] Experience storage strategy: When the absolute value of the liquid level tracking error is less than the set threshold or the absolute value of the control amount change is greater than the set threshold, the experience data is stored;

[0022] Priority sampling strategy: obtain the weight of each experience sample in the experience pool, divide high-reward samples and low-reward samples according to the set weight threshold, and sample at a ratio of 70% high-reward samples and 30% low-reward samples.

[0023] Preferably, in step 4, updating the Q table based on the time difference error specifically includes:

[0024] From the batch samples obtained by sampling, the Q value is corrected according to the TD (time difference) error. The TD error formula is:

[0025]

[0026] Among them, δ k represents the time difference (TD) error at the current time k, γ is the discount factor used to weigh the importance of current rewards and future rewards, and a′ represents the time difference (TD) error at the current time k+1 from the state v k+1 Starting from the optimal action that the agent may take, Q is the Q function, which is a state-action value function used to estimate the long-term cumulative reward that can be obtained by taking a specific action in a certain state, Q(v k ,a k ) means that at the current time k, the state is v k Take action a k Q value, r k Execute action a for agent k at the current moment k The reward value obtained after

[0027] The updated Q value is:

[0028] Q(v k ,a k )←Q(v k ,a k )+αδk (26)

[0029] Among them, α is the learning rate, which controls the magnitude of the change in Q value at each update.

[0030] Preferably, in step 4, dynamically adjusting the PID parameters specifically includes:

[0031] (1) Action selection: Generate a random number μ. If μ is less than the current exploration rate ε, randomly select the PID parameter adjustment amount; if μ is greater than or equal to the current exploration rate ε, select the PID parameter adjustment amount with the largest Q value;

[0032] (2) Exploration rate decay: The exploration rate ε decays dynamically with the number of training iterations ψ, and the formula is ε=max(ε0,λ ψ ,ε min ), where ε0 is the initial exploration rate, λ is the decay coefficient, and ε min is the minimum exploration rate;

[0033] (3) Parameter update and constraint: Update the PID parameters according to the selected action and limit the parameter range through the limit function.

[0034] The system adopts a dual-capacity water tank control method based on multi-objective reward Q learning, including:

[0035] Acquisition module: includes at least two liquid level sensors, used to collect the liquid level of the upper water tank and the lower water tank of the double-capacity water tank in real time and send it to the control module;

[0036] Control module: including an industrial controller, configured to execute any of the above-mentioned dual-capacity water tank control methods based on multi-objective reward Q learning, output PID parameter control quantities and send them to the execution module;

[0037] Execution module: including at least two electric regulating valves, used to receive PID parameter control quantity, and adjust the liquid inlet flow of the first water tank and the flow between the two water tanks according to the control quantity;

[0038] Communication module: used to realize data transmission between acquisition module, control module and execution module.

[0039] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of any of the above-mentioned dual-capacity water tank control methods based on multi-objective reward Q learning.

[0040] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned dual-capacity water tank control methods based on multi-objective reward Q learning are implemented.

[0041] Advantages of the present invention:

[0042] (1) Dynamic coordination mechanism of multi-objective reward function

[0043] The present invention designs a five-dimensional reward function that includes error penalty, control quantity fluctuation penalty, PID parameter stability reward, steady-state reward and parameter out-of-bounds constraint, and realizes multi-index dynamic balance through linear weighting. For example, the error penalty term and the control quantity fluctuation penalty term work together to preferentially suppress errors in the liquid level tracking stage, and automatically reduce the control frequency in the steady-state stage to reduce valve wear; the parameter out-of-bounds penalty term is directly related to the differential coefficient limit, and the parameter oscillation is suppressed through the negative feedback mechanism, realizing the physical constraint of PID parameters in the reinforcement learning framework for the first time. This function breaks through the limitations of traditional single-objective optimization, enabling the intelligent agent to autonomously switch the optimization target according to the real-time working conditions, and balance the tracking accuracy, control smoothness and system stability without human intervention;

[0044] (2) Event-triggered and reward-oriented efficient experience reuse

[0045] This invention proposes a priority experience replay mechanism based on key events: experience tuples are only stored when the liquid level tracking error exceeds the limit or the control variable suddenly changes, filtering out large amounts of steady-state redundant data and significantly reducing memory usage. At the same time, experience samples are weighted according to the absolute value of the reward, increasing the sampling probability of high-reward samples to 70%, accelerating the agent's learning of key control strategies. This mechanism solves the problem of low data utilization in traditional experience replay in complex systems, greatly improving the sampling efficiency of effective experience and significantly optimizing the convergence speed compared to uniform sampling strategies. It is suitable for the rapid response requirements of dynamic and time-varying scenarios such as double-capacity water tanks.

[0046] (3) Non-uniform discretization of three-dimensional states and model-free action mapping

[0047] The present invention adopts a non-uniform binning strategy to reduce the dimension of the continuous state space: by limiting the operation, generating bin boundaries and mapping bin indexes, the three-dimensional continuous state space is discretized into n^3 dimensional space, reducing the Q table dimension and computational complexity. Combined with the discretization constraint of the PID parameter adjustment amount, a closed-loop architecture of "state discretization-action constraint-online learning" is formed. It does not need to rely on the precise mathematical model of the double-capacity water tank, and end-to-end control can be achieved only through liquid level feedback. This solution compresses the state space dimension to n^3. 3 The number of action combinations is controlled at the level of o×l×m, which significantly reduces the computational complexity and adapts to the limited resources of the embedded controller.

[0048] (4) Lightweight model-free control architecture

[0049] An integrated "perception-decision-execution" framework was constructed, eliminating the traditional control method's reliance on the mathematical model of the water tank. The state space consists only of the three-dimensional components of level tracking error, error integral, and historical error, which are directly mapped to PID parameter adjustments through discretization, eliminating pre-processing steps such as model prediction and rule design. Q-table updates rely on real-time reward signals and state transition data, with control latency less than 10% of the sampling period, meeting the real-time requirements of industrial scenarios. This architecture breaks through the traditional "modeling-optimization-control" process, achieving adaptive regulation through data-driven control, and providing a plug-and-play solution for nonlinear and strongly coupled systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flow chart of a dual-capacity water tank control method based on multi-objective reward Q learning according to an embodiment of the present invention;

[0051] Figure 2 This is a structural diagram of a double-capacity water tank liquid level control system according to an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of a dual-capacity water tank control method based on multi-objective reward Q learning according to an embodiment of the present invention;

[0053] Figure 4 This is a comparison diagram of the parameter optimization process of an embodiment of the present invention, where (a) is the scale parameter k p Comparison chart of changes, (b) is the integral parameter k i The comparison chart of the changes, (c) is the differential parameter k d Comparison chart of changes;

[0054] Figure 5 A comparison diagram of liquid level tracking responses of three algorithms according to an embodiment of the present invention;

[0055] Figure 6 A comparison diagram of liquid level tracking errors of three algorithms according to an embodiment of the present invention;

[0056] Figure 7 Figure 1 is a comparison chart of the performance indicators of three algorithms according to an embodiment of the present invention, wherein (a) is a comparison chart of the average tracking error (AE), (b) is a comparison chart of the average overshoot (AO), and (c) is a comparison chart of the ITAE.

[0057] Figure 8 A comparison chart of average cumulative rewards between the QL algorithm of one embodiment of the present invention and the MOR-QL algorithm of the present invention;

[0058] Figure 9This is a stability comparison matrix bar chart of the QL algorithm of an embodiment of the present invention and the MOR-QL algorithm of the present invention, where (a) is the variance and mean graph of 5 experiments, (b) is the variance and mean graph of 10 experiments, (c) is the variance and mean graph of 15 experiments, and (d) is the variance and mean graph of 20 experiments. DETAILED DESCRIPTION

[0059] An embodiment of the present invention will be further described below with reference to the accompanying drawings.

[0060] In the embodiment of the present invention, a dual-capacity water tank control method based on multi-objective reward Q learning is shown in the flowchart of the method. Figure 1 As shown, the following steps are included:

[0061] Step 1: Modeling and state space construction of the dual-tank level control system:

[0062] The modeling of the double-capacity water tank level control system specifically includes:

[0063] The structural diagram of the double-capacity water tank level control system is as follows Figure 2 As shown, the system consists of two water tanks that are uniform in height and width and a liquid storage tank, which are used to store liquid and the storage volume is reflected by heights h1 and h2 respectively; it is equipped with connecting pipes to realize liquid circulation; three valves are set to control the flow rate Q1 flowing into the first water tank, the circulation flow rate Q2 between the two water tanks, and the outflow flow rate Q3 of the second water tank respectively; each water tank is also equipped with a liquid level detection device LT to monitor the liquid level height in real time, provide a liquid level feedback signal for system operation, and ensure effective monitoring and control of the water tank status. It is assumed here that both water tanks are uniform in height and width, and the cross-sectional areas of the upper and lower water tanks are A1 and A2 respectively. When the valve 3 remains open, the dynamic characteristics of the liquid level h2 when the flow rate Q1 changes are analyzed. Under the premise of ignoring the evaporation amount of the two water tanks, according to the material balance equation, the following differential equation can be obtained:

[0064]

[0065] Among them, t is a continuous time variable, representing any moment, is the rate of change of liquid level with time;

[0066] Performing Laplace transform on equations (1) and (2) yields:

[0067] A1sH1(s)=Q1(s)-Q2(s) (3)

[0068] A2sH2(s)=Q2(s)-Q3(s) (4)

[0069] Among them, H1(s) is the function of the liquid level of the upper water tank in the complex frequency domain after Laplace transform, and H2(s) is the function of the liquid level of the lower water tank in the complex frequency domain after Laplace transform;

[0070] Let the hydraulic resistance of valves 1, 2, and 3 be R1, R2, and R3 respectively, then:

[0071]

[0072] Substituting equations (5) and (6) into equations (3) and (4) respectively, the mathematical model of the double-capacity water tank level control system can be obtained as follows:

[0073]

[0074] The transfer function G(s) of the system is:

[0075]

[0076] Among them, T1=A1R1, T2=A2R2.

[0077] The formula for bilinear transformation is:

[0078]

[0079] Where T is the sampling period, z is the complex variable in the discrete system z-transform;

[0080] Substituting Equation (9) into Equation (8), the transfer function G(z) of the discretized system can be obtained as follows:

[0081]

[0082] Among them, M1=T1T2, M2=T1+T2.

[0083] The mathematical model is used to describe the dynamic response characteristics of the double-capacity water tank liquid level to the inlet flow rate, providing a theoretical basis for state space construction.

[0084] Constructing a three-dimensional state space specifically includes:

[0085] The liquid level tracking error e(k) of the system at the current moment, the integral of the error (∑e(k)) and the error of the previous time step (e(k-1)) are regarded as the system state, and the system state space is defined as:

[0086] V=[e(k),∑e(k),e(k-1)] (11)

[0087] Where k is the discrete time index, corresponding to t = kT, ξ is the index variable for summation, which starts from 0 and traverses to k in sequence. It is used to accumulate and sum the liquid level tracking errors at different times. e(ξ) is the liquid level tracking error of the system at the ξth time step and is the basic element involved in the summation operation.

[0088] In order to reduce the dimension and computational complexity of the Q table, a discretization strategy is used to reduce the dimension of the continuous state space. The discretization design is as follows:

[0089] (1) First, for each state component v, a limiting operation is performed:

[0090] v clipped =clip(v,-1,1) (12)

[0091] (2) Secondly, according to the set number of bins n, generate n+1 bin boundaries:

[0092] bin edges =linspace(-1,1,n+1) (13)

[0093] Among them, bin edges is the bin boundary;

[0094] (3) Finally, the state component v after limiting clipped Map to the corresponding bin index and force the index to be within the range of [0, n-1]:

[0095] idx=digitize(v clipped ,bin edges )-1 (14)

[0096] Among them, idx is the bin index;

[0097] By discretization, each state component is discretized into n levels, so the number of discrete states after the three state components are combined is n 3 .

[0098] Step 2: Define the incremental adjustment action space of PID parameters and construct a multi-objective reward function that includes error and control amount fluctuation rewards, PID parameter change rewards, control amount change rewards, steady-state rewards, and parameter out-of-bounds rewards. This function collaboratively optimizes tracking error, control amount fluctuation, and parameter stability. Specifically, it includes:

[0099] First, define the incremental adjustment action space of PID parameters:

[0100] The action space A is defined as the combination of PID parameter adjustments, namely:

[0101] A=[ΔK p ,ΔK i ,ΔKd ] (15)

[0102] Where ΔK p is the proportional coefficient K p Adjustment amount, ΔK i is the integration coefficient K i Adjustment amount, ΔK d is the differential coefficient K d The amount of adjustment.

[0103] At the same time, in order to ensure the rationality of PID parameters and the stability of the system, the adjustment range of each parameter will be restricted as follows:

[0104]

[0105] Among them, o, l, m are finite positive integers, by ΔK p ,ΔK i ,ΔK d Combining all possible values ​​of , we obtain a total of o × l × m action combinations. These different action combinations constitute the agent's action space. This incremental action design avoids the drastic fluctuations associated with full adjustment of traditional PID parameters and is more suitable for the hysteresis characteristics of dual-capacity water tanks.

[0106] Secondly, construct a multi-objective optimization reward function:

[0107] To address the problem of traditional reinforcement learning PID control with a single reward function and its difficulty in balancing dynamic tracking and system stability, this invention achieves multi-objective collaborative control under dynamic conditions by collaboratively optimizing indicators such as liquid level tracking error, control variable change, PID parameter changes, and the system's steady-state performance. The specific reward function r consists of five parts: error and control variable fluctuation reward r1, PID parameter change reward r2, control variable change reward r3, steady-state reward r4, and parameter out-of-bounds reward r5.

[0108] In order to enable the agent to minimize the error and fluctuation of the control amount, r1(k) is calculated based on the liquid level tracking error at the current moment k, the liquid level tracking error at the previous moment, and the change in the control amount. The calculation formula is as follows:

[0109] r1(k)=-[ω1e(k) 2 +ω2(e(k-1)) 2 +ω3(Δu) 2 ] (17)

[0110] Among them, ω1∈[0.5,1.5] represents the weight of the error penalty term. Increasing ω1 can strengthen the error penalty and improve tracking accuracy, while reducing it allows larger fluctuations to optimize other indicators. ω2∈[0.05,0.15] represents the weight of the error change trend penalty term. If the value exceeds 0.15, the system will be overly sensitive to small fluctuations in the error, which may cause frequent adjustments to the control variable and lead to response oscillation. If the value is lower than 0.05, the role of the differential link (error trend suppression) is weakened, making it difficult to effectively suppress error divergence. ω3∈[0.025,0.075] represents the weight of the control variable change penalty term. Increasing ω3 can reduce frequent valve operation and reduce actuator wear. If the value is lower than 0.025, the control variable is prone to high-frequency oscillation, destroying the system steady state and even causing equipment failure. Δu=u(k)-u(k-1) represents the change in the control variable between two adjacent moments, where u(k) is the control variable at the current moment k.

[0111] In order to encourage the agent to maintain the stability of the PID parameters, r2(k) is calculated according to the change of the PID parameters at the current moment k. The calculation formula is as follows:

[0112] r2(k)=-ω4Δk t (18)

[0113] Among them, ω4∈[0.075,0.225] represents the penalty weight of PID parameter change. If the value exceeds 0.225, the parameter adjustment is overly suppressed, which may cause the system response to lag. If the value is lower than 0.075, the parameter change penalty is insufficient, which may easily cause proportional / integral / differential coefficient oscillation and destroy control stability. t =|ΔK p |+|ΔK i |+|ΔK d | represents the sum of the absolute values ​​of the changes in the three parameters of the PID controller at the current moment k, which is used to measure the overall change amplitude of the PID parameters at that moment.

[0114] In order to avoid large fluctuations in the control amount, it is necessary to further penalize the change in the control amount. The calculation formula is as follows:

[0115] r3(k)=-ω5|Δu| (19)

[0116] Among them, ω5∈[0.04,0.12] indicates that the control amount further punishes the weight, and its value needs to be coordinated with ω3 to avoid repeated punishment or failure.

[0117] In order to encourage the agent to make the system reach a steady state faster, it is specially stipulated that when the continuous The level tracking errors of each step are all smaller than the set threshold, i.e., |e(k)| <e σWhen this condition is met, the agent will be given a steady-state reward, namely:

[0118]

[0119] in, If the number of steps is too small, the reward will be triggered frequently and the steady state cannot be accurately reflected. If the number of steps is too large, it will increase the difficulty of the agent to reach the steady state. It can not only effectively motivate the agent to pursue steady state, but also avoid excessive rewards that cause the agent to focus too much on steady state and ignore other performance indicators.

[0120] Finally, in order to avoid the PID parameters from exceeding the reasonable range, it is specially stipulated that when the differential coefficient K d Reaching the limit (K d K d,min or K d,max ), a parameter out-of-bounds penalty will be given, namely:

[0121] r5=r p (twenty one)

[0122] Among them, r p ∈[-1,-0.1], which indicates the penalty value for parameter out-of-bounds. This range can give reasonable penalties when the parameters are out of bounds, avoiding penalties that are too light to constrain or too heavy to cause imbalance in the agent's learning.

[0123] In summary, the reward r(k) obtained by the agent at the current time k can be calculated as follows:

[0124] r(k)=r1(k)+r2(k)+r3(k)+r4+r5 (22)

[0125] By collaboratively designing a multi-objective reward function, the intelligent agent is able to achieve a dynamic balance between tracking accuracy, control smoothness, parameter stability, and steady-state performance. Compared to single-objective reward functions, this approach significantly improves the adaptive control capabilities of complex industrial systems under dynamic disturbances through weight allocation and multi-objective collaborative optimization, providing a new technical path for the real-time optimization of highly coupled, nonlinear systems.

[0126] Step 3: Use an event-triggered priority experience replay strategy to store experience data only when the level tracking error exceeds the limit or the control variable changes suddenly, and assign sampling weights based on the absolute value of the reward, specifically including:

[0127] To address the problems of low data efficiency and policy oscillation in traditional Q-learning in complex dynamic systems, this paper introduces a priority experience replay mechanism, which improves the optimization efficiency of multi-objective reward functions through key event-triggered storage and reward-guided sampling.

[0128] Experience storage strategy:

[0129] In order to avoid storing a large amount of redundant experience data, only when the level tracking error exceeds the limit (|e(k)|>θ e Or the control quantity mutation (|Δu(k)|>θ u ) when storing the experience tuple (v k ,a k ,r(k),v k+1 ), where θ e is the liquid level tracking error exceeding the limit, θ u is the control quantity mutation value, v k is the state of the system k at the current moment, a k The action taken by agent k at the current moment corresponds to a set of parameter adjustment combinations ΔK = [ΔK p ,ΔK i ,ΔK d ]. Low-value data under steady-state conditions is filtered out through an event-triggered mechanism, significantly reducing memory usage.

[0130] Priority sampling strategy:

[0131] For each sample τ in the experience replay pool, according to the absolute value of the reward |r τ |Give each experience sample a weight w τ ,Right now:

[0132]

[0133] Among them, r τ is the reward value of sample τ, r j is the reward value of the jth sample in the experience replay pool, where j is an index variable (j = 1, 2, ..., N, N is the total number of samples in the experience replay pool), which is used to traverse each sample in the experience replay pool and calculate the reward value r for all samples. j Sum to determine the weight w of each sample τ τ The normalized denominator of , thus giving each sample a corresponding weight according to the reward size, and achieving priority sampling. Through this weight setting, samples with high rewards are given priority. τ | is larger, its weight w τ Higher, thus having a higher probability of being selected for updating, achieving priority learning of important experiences.

[0134] Specifically, according to the weight threshold w τh , divide the experience samples into high reward samples (w τ ≥w τh ) and low reward samples (w τ <w τh ),Right now:

[0135]

[0136] Among them, high is a high-reward experience sample, and low is a low-reward experience sample.

[0137] Each time sampling is performed, the ratio of 70% high reward samples and 30% low reward samples is used to ensure that the agent can learn more strategies to obtain high rewards, thereby accelerating the convergence of the strategy.

[0138] Step 4: Update the Q table based on the time difference error, dynamically adjust the PID parameters and constrain their range through the limit function, specifically including:

[0139] (1) Q-table update and batch learning

[0140] When the number of samples in the experience replay pool reaches 32, a Q-table update is triggered. From the batch of samples obtained by sampling, the Q value is corrected according to the TD (time difference) error. The TD error formula is:

[0141]

[0142] Among them, δ k represents the time difference (TD) error at the current time k, γ is the discount factor used to weigh the importance of current rewards and future rewards, and a′ represents the time difference (TD) error at the current time k+1 from the state v k+1 Starting from the optimal action that the agent may take, Q is the Q function, which is a state-action value function used to estimate the long-term cumulative reward that can be obtained by taking a specific action in a certain state, Q(v k ,a k ) means that at the current time k, the state is v k Take action a k Q value;

[0143] The updated Q value is:

[0144] Q(v k ,a k )←Q(v k ,a k )+αδ k (26)

[0145] Among them, α is the learning rate, which controls the amplitude of the Q value change at each update. It takes a value between 0 and 1. The larger the learning rate, the larger the update amplitude and the faster the response to new information. The smaller the learning rate, the slower the update and the better the stability.

[0146] Through this batch update mechanism, the correlation between consecutive experiences can be alleviated, thereby improving the stability of the algorithm.

[0147] (2) Strategy Optimization: Exploring and Utilizing Adaptive Mechanisms

[0148] To address the issues of slow convergence and easy local optima associated with traditional fixed exploration strategies, this invention designs an exploration-utilization adaptive mechanism that dynamically balances the need to explore new parameter combinations with the need to utilize known optimal strategies, achieving efficient optimization of PID parameters. Its core logic consists of two parts: action selection logic and dynamic adjustment of the exploration rate.

[0149] Action selection logic:

[0150] When the agent selects the PID parameter adjustment action at each moment k, it generates a uniform random number μ between 0 and 1 and compares it with the current exploration rate ε to select the action:

[0151] When μ < ε, perform random exploration with probability ε: randomly select ΔK from the action space A (Formula 15) p / ΔK i / ΔK d Combination,parameter adjustment is possible to cover high error areas;

[0152] When μ ≥ ε, perform greedy exploitation with probability 1-ε: retrieve the current state v k Under this condition, the Q values ​​corresponding to all possible actions are calculated by argmaxQ(v k ,a k ) Select the action with the greatest value in the current state and accelerate the convergence to the known optimal strategy.

[0153] Exploration rate dynamically decays:

[0154] In order to achieve “extensive exploration in the early stage and focused convergence in the later stage”, the exploration rate ε is dynamically adjusted with the number of training iterations ψ, that is, ε=max(ε0,λ ψ ,ε min ), where ε0 is the initial exploration rate, λ is the decay coefficient, and ε min The minimum exploration rate is set. In this way, new strategies can be explored with high probability in the early stage of training, covering more parameter combinations. The exploration rate gradually decreases in the later stage of training, and strategy optimization focuses on "using existing experience to converge to the global optimum", taking into account both comprehensive exploration and efficient convergence.

[0155] Collaboration with Q-table update: After the actions generated by random exploration are evaluated by the multi-objective reward function (Formula 22), high-reward samples are preferentially sampled at a ratio of 70% (the priority replay strategy in step 3), driving the Q-table to update in the optimization direction, forming a closed loop of "exploration → reward evaluation → Q-value correction → strategy iteration".

[0156] Furthermore, PID parameters are adjusted, and the agent will adjust the current state v at each moment k. k Select an action a from the action space A k, and then update the PID parameters according to the action. The formula is as follows:

[0157]

[0158] Among them, clip(x,x min ,x max ) is a limiting function, which is used to limit the parameters to a reasonable range.

[0159] Step 5: Substitute the adjusted PID parameters into the discrete PID formula to generate the control quantity, and perform closed-loop real-time control on the double-capacity water tank, specifically including:

[0160] In each time period, the updated PID parameters are substituted into the control quantity calculation:

[0161]

[0162] The output control quantity u(k) is applied to the actuator, and the liquid inlet flow Q1 is adjusted through real-time feedback from the sensor to achieve adaptive adjustment of the double-capacity water tank.

[0163] Example 1

[0164] In the embodiment of the present invention, Figure 3 As shown in the figure, in order to more intuitively demonstrate the effectiveness of the dual-capacity water tank control method based on multi-objective reward Q learning proposed in this invention, Python software was used to simulate and verify the method proposed in this invention. The agent performed 400 training rounds, each round containing 60 steps, and the total number of training times was 24,000. The experimental parameter settings are shown in Table 1:

[0165] Table 1

[0166]

[0167] Step 1: Modeling and state space construction of the dual-tank level control system:

[0168] Substituting the parameters in Table 1, the description of the double-tank level control system is as follows:

[0169]

[0170] The discrete transfer function is:

[0171]

[0172] The constructed three-dimensional state space is: Using formulas (11)-(14), setting the number of bins to 5, generating 6 bin boundaries, and forcibly limiting the index to the range of [0,4], each state component is discretized into 5 bins through discretization, so the number of discrete states after the three state components are combined is 53 .

[0173] Step 2: Define the incremental adjustment action space of PID parameters and construct a multi-objective reward function that includes error and control amount fluctuation rewards, PID parameter change rewards, control amount change rewards, steady-state rewards, and parameter out-of-bounds rewards. This function collaboratively optimizes tracking error, control amount fluctuation, and parameter stability. Specifically, it includes:

[0174] First, define the incremental adjustment action space of PID parameters:

[0175] The action space A is defined as the combination of PID parameter adjustment quantities, expressed as formula (15);

[0176] At the same time, in order to ensure the rationality of PID parameters and the stability of the system, the adjustment range of each parameter will be restricted as follows:

[0177]

[0178] By ΔK p ,ΔK i ,ΔK d By combining all possible values ​​of , the total number of action combinations is 7×7×5. These different action combinations constitute the action space of the intelligent agent.

[0179] Secondly, construct a multi-objective optimization reward function:

[0180] In order to enable the agent to minimize the error and fluctuation of the control amount, r1(k) is calculated based on the liquid level tracking error at the current moment k, the liquid level tracking error at the previous moment, and the change in the control amount. The calculation formula is as follows:

[0181] r1(k)=-[e(k) 2 +0.1(e(k-1)) 2 +0.05(Δu) 2 ] (32);

[0182] In order to encourage the agent to maintain the stability of the PID parameters, r2(k) is calculated according to the change of the PID parameters at the current moment k. The calculation formula is as follows:

[0183] r2(k)=-0.15Δk t (33); In order to avoid large fluctuations in the control amount, it is necessary to further penalize the change in the control amount. The calculation formula is as follows:

[0184] r3(k)=-0.08|Δu| (34);

[0185] In order to encourage the agent to reach a steady state faster, it is stipulated that when the level tracking error for 8 consecutive steps, i.e., |e(k)| < 0.05, is met, a steady-state reward will be given to the agent, i.e.:

[0186] r4=4.0 (35);

[0187] Finally, in order to avoid the PID parameters from exceeding the reasonable range, it is specially stipulated that when the differential coefficient K d Reaching the limit (K d When it is 0 or 0.1), a parameter out-of-bounds penalty will be given, namely:

[0188] r5=0.5 (36);

[0189] The reward r(k) obtained by the agent at the current time k can be calculated by formula (22).

[0190] Step 3: Use an event-triggered priority experience replay strategy to store experience data only when the level tracking error exceeds the limit or the control variable changes suddenly, and assign sampling weights based on the absolute value of the reward, specifically including:

[0191] Experience storage strategy: The experience tuple (v is stored only when the level tracking error exceeds the limit (|e(k)|>0.1) or the control amount suddenly changes (|Δu(k)|>0.2). k ,a k ,r k ,v k+1 );

[0192] Priority sampling strategy: For each sample τ in the experience replay pool, according to the absolute value of the reward |r τ |Give each experience sample a weight w τ , the expression is formula (23);

[0193] Specifically, according to the weight threshold of 0.01, the experience samples are divided into high reward samples (w τ ≥0.01) and low reward samples (w τ <0.01), that is:

[0194]

[0195] Each time sampling is performed, the ratio of 70% high-reward samples and 30% low-reward samples is used.

[0196] Step 4: Update the Q table based on the time difference error, dynamically adjust the PID parameters and constrain their range through the limit function, specifically including:

[0197] When the number of samples in the experience replay pool reaches 32, a Q-table update is triggered. From the batch of samples obtained by sampling, the Q value is corrected according to the TD (time difference) error. The TD error formula is formula (25), and the updated Q value expression is formula (26);

[0198] Furthermore, PID parameters are adjusted, and the agent will adjust the current state v at each moment k. k Select an action a from the action space A k , and then update the PID parameters according to the action. The formula is as follows:

[0199]

[0200] Step 5: Substitute the adjusted PID parameters into the discrete PID formula to generate the control quantity, and perform closed-loop real-time control on the double-capacity water tank, specifically including:

[0201] In each time period, the updated PID parameters are substituted into the control quantity calculation, and the expression is formula (28); the output control quantity u(k) is applied to the actuator, and the liquid inlet flow Q1 is adjusted through real-time feedback from the sensor to achieve adaptive regulation of the double-capacity water tank.

[0202] In order to verify the performance of the present invention in the tracking control of the dual-capacity water tank level control system, a comparative analysis is carried out on the average tracking error, average overshoot, time multiplied absolute error and average cumulative reward of the traditional PID algorithm, QL (Q-learning) algorithm and the MOR-QL algorithm of the present invention.

[0203] Average Tracking Error (AE): This metric reflects the average level of error between the actual liquid level and the set value during the entire training process. It is used to measure the overall accuracy of the system in tracking the set value. A smaller value indicates a higher tracking accuracy.

[0204]

[0205] Where, Υ is the total number of time steps;

[0206] Average overshoot (AO): It is used to measure the maximum deviation of the liquid level response from the set value, reflecting the stability of the system's dynamic response. The smaller the overshoot, the smoother the system response and the weaker the oscillation phenomenon:

[0207]

[0208] Among them, E is the total number of training rounds, It is a counting variable, which is used to indicate the number of rounds currently being calculated in the process of traversing and summing the total training rounds. training rounds;

[0209] Time-multiplied integral of absolute error (ITAE): This value comprehensively considers the size and duration of the error and is used to evaluate the overall performance of the system in the dynamic adjustment phase and the steady-state phase. The smaller the value, the better the overall performance of the system in terms of dynamic adjustment and steady-state maintenance:

[0210]

[0211] Average Cumulative Reward (AR): This represents the average level of rewards received by the agent during training, reflecting the overall quality of the agent's strategy. A larger value indicates a better strategy.

[0212]

[0213] in, For the The cumulative reward of each training round.

[0214] In the embodiment of the present invention, the performance indicators of the three different algorithms are shown in Table 2. The results obtained after multiple experiments using the QL algorithm and the MOR-QL algorithm of the present invention are shown in Table 3.

[0215] Table 2

[0216]

[0217] Table 3

[0218]

[0219] In the embodiment of the present invention, Figure 4 It can be seen that the parameter optimization process of the present invention is more stable and the parameter fluctuation is smaller. Figure 5 It can be seen that the liquid level curve response of the present invention is closest to the set value and has the smallest fluctuation range compared to the other two. Figure 6 It can be seen that the tracking error curve of the present invention is the most stable. Figure 7 As can be seen from Table 2, the performance index of the present invention is the best. Figure 8 It can be seen that the present invention has a higher cumulative reward and can reach convergence faster. Figure 9 As can be seen from Table 3, the proposed method exhibits excellent stability across multiple repeated tests, with its mean and variance consistently maintained within extremely low numerical ranges, demonstrating the algorithm's strong robustness. Simulation examples demonstrate the effectiveness of this approach. The simulation results show that the proposed dual-tank control method based on multi-objective reward Q-learning can accurately and efficiently achieve stable control of the system, with excellent performance in key indicators such as dynamic response and steady-state accuracy.

[0220] This application also provides a dual-capacity water tank control system based on multi-objective reward Q learning, including:

[0221] (1) Acquisition module: includes at least two liquid level sensors for real-time acquisition of the liquid level of the upper and lower water tanks of the double-capacity water tanks and sends the data to the control module; the liquid level sensor uses an ultrasonic level meter with a measurement accuracy of ≤±1mm;

[0222] (2) Control module: including an industrial controller for executing the dual-capacity water tank control method based on multi-objective reward Q learning, outputting PID parameter control quantities and sending them to the execution module; the industrial controller is an embedded system with an FPGA acceleration module, supporting parallel updates of the Q table, and a single-step calculation delay of ≤10ms;

[0223] (3) Execution module: including at least two electric regulating valves, which are used to receive PID parameter control quantities and adjust the liquid inlet flow rate Q1 and the water tank circulation flow rate Q2 according to the control quantities. The opening adjustment resolution of the electric regulating valve is 0.1%, and the response time is ≤500ms;

[0224] (4) Communication module: used to realize data transmission between the acquisition module, control module and execution module. The transmission cycle is synchronized with the sampling cycle.

[0225] The present application also provides an electronic device that may include a memory and a processor, wherein the memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the electronic device may also include various network interfaces, a power supply, and other components.

[0226] The present application also provides a readable storage medium having a computer program stored thereon, which, when executed, can implement the steps provided in the above embodiments. The storage medium may include: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0227] The above is a detailed introduction to the methods, systems, devices, and media provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the core ideas of the present invention. It should be noted that, for those skilled in the art, without departing from the principles of the present invention, several improvements and modifications may be made to the present invention, and such improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A dual-capacity water tank control method based on multi-objective reward Q learning, characterized in that: The following steps are involved: Step 1: Establish a discretized mathematical model of the dual-tank liquid level control system, and construct a three-dimensional continuous state space based on the liquid level tracking error, error integral, and previous time step error; Step 2: Define the incremental adjustment action space of PID parameters and design a multi-objective reward function to collaboratively optimize the liquid level tracking error, control variable fluctuation, and parameter stability; Step 3: Adopt an event-triggered priority experience replay strategy, store experience data only when the level tracking error exceeds the limit or the control amount changes suddenly, and assign sampling weights based on the absolute value of the reward; Step 4: Update the Q table based on the time difference error and dynamically adjust the PID parameters; Step 5: Substitute the adjusted PID parameters into the discrete PID formula to generate the control quantity, and perform closed-loop real-time control on the double-capacity water tank.

2. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 1, the discretized mathematical model of the dual-capacity water tank level control system is: The transfer function of the discretized system is: Among them, M1=T1T2, M2=T1+T2, T1=A1R1, T2=A2R2, H2(s) is the function of the lower water tank level in the complex frequency domain after Laplace transform, Q1 is the flow rate of the first water tank, A1 and A2 are the cross-sectional areas of the upper and lower water tanks respectively, R1 and R2 are the liquid resistances of valves 1 and 2 respectively, T is the sampling period, and z is the complex variable in the discrete system z transform.

3. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 1, the three-dimensional continuous state space is discretized by the following steps: performing a clipping operation on each state component, binning the clipped state component into n bins, generating a discrete state index, and combining the three state components to obtain n 3-dimensional discrete state space.

4. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 2, the multi-objective reward function includes: error and control amount fluctuation reward r1, PID parameter change reward r2, control amount change reward r3, steady-state reward r4, and parameter out-of-bounds reward r5; The steady-state reward r4 is: setting a continuous step threshold and a liquid level tracking error threshold. When the system's continuous steps reach the threshold and the absolute value of the liquid level tracking error of each step is less than the set threshold, a steady-state reward is given to the agent.

5. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 3, the priority experience replay strategy includes: Experience storage strategy: When the absolute value of the liquid level tracking error is less than the set threshold or the absolute value of the control amount change is greater than the set threshold, the experience data is stored; Priority sampling strategy: obtain the weight of each experience sample in the experience pool, divide high-reward samples and low-reward samples according to the set weight threshold, and sample at a ratio of 70% high-reward samples and 30% low-reward samples.

6. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 4, updating the Q table based on the time difference error specifically includes: From the batch samples obtained by sampling, the Q value is corrected according to the TD (time difference) error. The TD error formula is: Among them, δ k represents the time difference (TD) error at the current time k, γ is the discount factor used to weigh the importance of current rewards and future rewards, and a′ represents the time difference (TD) error at the current time k+1 from the state v k+1 Starting from the optimal action that the agent may take, Q is the Q function, which is a state-action value function used to estimate the long-term cumulative reward that can be obtained by taking a specific action in a certain state, Q(v k , α k ) means that at the current time k, the state is v k Take action α k Q value, r k Execute action a for agent k at the current moment k The reward value obtained after The updated Q value is: Q(v k ,a k )←Q(v k ,a k )+αδ k (26) Among them, α is the learning rate, which controls the magnitude of the change in Q value at each update.

7. The dual-capacity water tank control method based on multi-objective reward Q learning according to claim 1 is characterized in that: In step 4, dynamically adjusting the PID parameters specifically includes: (1) Action selection: Generate a random number μ. If μ is less than the current exploration rate ε, randomly select the PID parameter adjustment amount. If μ is greater than or equal to the current exploration rate ε, select the PID parameter adjustment amount with the largest Q value. (2) Exploration rate decay: The exploration rate ε decays dynamically with the number of training iterations ψ, and the formula is ε = max(ε0, λ ψ , ε min ), where ε0 is the initial exploration rate, λ is the decay coefficient, and ε min is the minimum exploration rate; (3) Parameter update and constraint: Update the PID parameters according to the selected action and limit the parameter range through the limit function.

8. The system adopts a dual-capacity water tank control method based on multi-objective reward Q learning, characterized by: include: Acquisition module: includes at least two liquid level sensors, used to collect the liquid level of the upper water tank and the lower water tank of the double-capacity water tank in real time and send it to the control module; Control module: comprising an industrial controller, configured to execute the dual-capacity water tank control method based on multi-objective reward Q learning according to any one of claims 1 to 7, output PID parameter control quantities and send them to the execution module; Execution module: including at least two electric regulating valves, used to receive PID parameter control quantity, and adjust the liquid inlet flow of the first water tank and the flow between the two water tanks according to the control quantity; Communication module: used to realize data transmission between acquisition module, control module and execution module.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the dual-capacity water tank control method based on multi-objective reward Q learning according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the dual-capacity water tank control method based on multi-objective reward Q learning according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Inverter PID parameter adaptive optimization method and system based on Q learning

    CN120972502A

  • Thermal power generation boiler water level control optimization method based on particle swarm optimization

    CN121165812A

  • An Optimization Method for Water Level Control in Thermal Power Boilers Based on Particle Swarm Optimization

    CN121165812B