Fault-tolerant h-infinity tracking control method for oil-water separation system based on reinforcement learning
Patent Information
- Application Number
- CN202211606347.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-12-14
AI Technical Summary
系统运行过程中容易收到执行器故障的影响,阀门生锈、磨损、卡住等非常常见,而执行器故障会导致系统性能严重下降,导致排放物的OiW浓度过高,因此需要在控制器设计中进行处理
[0043]1、本发明提出的基于强化学习的H∞控制,系统从运行数据中自主学习最优控制策略,既具有H∞控制的优点,又不依赖于模型;工业中常用的PID控制,液位跟踪精准,不需要模型,但是不能实现液位控制和PDR控制的协调PDR无法保证1.5~3之间;H∞控制放宽了液位跟踪性能,提高PDR跟踪精度;但是依赖精确模型,工业系统维护成本高。
Smart Images

Figure CN115903515B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of model-independent optimal tracking control and fault-tolerant control technology in offshore oil production processes, and in particular to a fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning. Background Technology
[0002] Offshore oil and gas production generates large quantities of water and hydrocarbon mixtures. In mature offshore oil and gas fields, the ratio of produced water (PW) to well fluids can exceed 90%. PW must meet stringent oil-in-water (OiW) concentration standards, such as 30 mg / L in the North Sea, before it can be discharged or reinjected. Oil-water separation systems are crucial facilities for PW treatment in oil and gas production. Different installation schemes are required in different offshore areas, making model-based control system design challenging and necessitating a model-free solution. During system operation, actuator failures are common; valve rust, wear, and jamming are frequent problems. Actuator failures can severely degrade system performance, leading to excessively high OiW concentrations in emissions. Therefore, this issue needs to be addressed in the controller design. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a fault-tolerant H-infinite tracking control method for oil-water separation system based on reinforcement learning, which enables the system to learn the optimal control strategy autonomously from the operating data, improves the PDR tracking accuracy without relying on the model, and ensures that the system has good fault-tolerant performance while taking into account the optimal control performance of the system.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0005] A fault-tolerant H-infinite tracking control method for an oil-water separation system based on reinforcement learning includes the following steps:
[0006] S1. Establish an oil-water separation system model with interference and actuator failure. The interference is the flow rate of the oil-water mixture, and the actuator failure model is established based on the characteristics of the valve.
[0007] S2. The cooperative control problem of the oil-water separation system model is expressed as a cascaded fault-tolerant H∞ tracking control problem, which reduces the water level tracking accuracy while reducing the sensitivity of the pressure drop rate to inflow disturbances.
[0008] S3. The H∞ tracking control problem is transformed into a two-role zero-sum differential game problem. By establishing the game algebra Riccati equation, a model-based fault-tolerant control solution is obtained, and the optimal control solution is obtained.
[0009] S4. Construct a model-independent, fault-tolerant H∞ controller based on the optimal control solution;
[0010] S5. A non-strategy reinforcement learning algorithm is proposed to find the solution of the Riccati equation in a data-driven manner. The learned model-free solution is consistent with the model-based game algebra Riccati equation solution, and the optimal solution that restores the water level and pressure drop rate to normal is obtained, thus realizing the fault compensation of the actuator.
[0011] A further improvement to the technical solution of the present invention is that S1 includes the following steps:
[0012] S1.1 Establish an actuator fault model;
[0013] Based on the characteristics of the valve, assume the control input u i For i = 1, 2, there is partial loss of validity and bias fault:
[0014]
[0015] Among them, u Δi β represents the input uncertainty caused by actuator failure. i The efficiency loss coefficient, For bias;
[0016] S1.2 Establish mathematical models for the water level control subsystem and the PDR control subsystem, and simultaneously provide a system model with external disturbances and actuator failures: y = l;
[0017] Wherein, PDR represents the pressure drop rate;
[0018] The oil-water separation system can be viewed as a cascaded system consisting of a water level control subsystem and a PDR control subsystem.
[0019] In the water level control subsystem:
[0020] u = u1, uΔ = uΔ1;
[0021] Where l represents the liquid level, ΔP u =P i -P u and These are the pressure drop and the rate of change of pressure drop at the relief valve, respectively, P i It is the valve pressure. It is the valve pressure change rate, u1 is the control signal for opening the relief valve, u Δ =u Δ1 The input uncertainty is caused by actuator failure;
[0022] The state of the PDR control subsystem is: Where, ΔP o =P i -P oand These represent the pressure drop and the rate of change of pressure drop at the downstream valve, respectively. u2 is the control signal for opening the downstream valve. Δ =u Δ2 The input uncertainty is caused by an actuator malfunction, and the output value is y = ΔP. o ;
[0023] The water level control subsystem and PDR control subsystem models with external disturbances and actuator failures are both described by the following linear system:
[0024]
[0025] Where x is the system state, y is the system output, u is the control signal for valve opening, and w = F in It is external interference, ρ(tt) s ) is a step function, t s The time when the fault occurs satisfies t > t s At that time, ρ(tt) s ) = 1, otherwise ρ(tt) s ) = 0; u Δ The input uncertainty is caused by system failure, given by the failure model established in S1.1; system matrices A, B, C, and E are unknown system matrices.
[0026] A further improvement of the technical solution of the present invention is that: in S2, robust H∞ control is applied to the oil-water separation system model, and the tracking performance is relaxed while satisfying the H∞ performance, so that the change of the inflow rate of the hydrocyclone and the control input u2 are less sensitive to the fluctuation of disturbances.
[0027] The control objective is to make the output y track a linear reference trajectory r. d To achieve tracking, we define ξ = [x T ,r d T ] T When the system is fault-free, the augmented system is constructed based on the linear system model given in S1 as follows:
[0028]
[0029] in, e = yr d It is the tracking error, w = F in It's external interference.
[0030] A further improvement to the technical solution of the present invention lies in: in S3, according to the L2 gain condition... Define a performance function:
[0031]
[0032] Where α is a discount factor that satisfies α>0, Q and R are positive definite matrices, and γ>0 represents the amount of attenuation from the perturbation input w(t) to the defined performance output variable z(t);
[0033] Based on the performance function, the H∞ tracking control problem is treated as a two-player zero-sum game, and the Nash equilibrium solution (u) is found. * ,w * This makes the closed-loop system stable and robust. For any disturbance w, the closed-loop control system satisfies the L2 gain condition, that is, the tracking error e converges to the origin neighborhood whose boundary depends on γ.
[0034] When the system is fault-free, the Riccati equation can be derived, and the optimal control strategy can be given; for model-based fault-tolerant control, actuator faults are addressed by adding an estimation term u to the controller. c To compensate.
[0035] A further improvement to the technical solution of this invention lies in: in S4, u is used. * Instead of the terms containing the system matrix B in the model-based fault-tolerant controller derived in S3, u * It is learned through reinforcement learning algorithms.
[0036] A further improvement to the technical solution of the present invention is that S5 includes the following steps:
[0037] S5.1 collects system data, including system state information and measured disturbance data, to update the Bellman equation in the non-policy reinforcement learning algorithm, and iteratively learns to obtain the system's control gain K. u * Interference gain K w * Thus, the optimal control strategy can be obtained;
[0038] The Bellman equation is updated iteratively to obtain K. u * ,K w * :
[0039]
[0040] Among them, V j It has the form of a quadratic form: V j =ξ T P j Substituting ξ into the above equation, we can further obtain P. j That is, the solution to the Riccati equation, which is consistent with the optimal solution of the model-based fault-tolerant control in S3;
[0041] S5.2 Based on the online parameter estimation information of the fault signal provided by the adaptive law, a fault-tolerant compensation controller is constructed on the basis of the learned optimal control strategy, and a model-free fault-tolerant compensation control scheme is given to compensate for the fault.
[0042] The technological advancements achieved by this invention due to the adoption of the above technical solutions are as follows:
[0043] 1. The H∞ control based on reinforcement learning proposed in this invention allows the system to autonomously learn the optimal control strategy from operating data. It has the advantages of H∞ control but does not rely on a model. PID control, which is commonly used in industry, provides accurate liquid level tracking and does not require a model, but it cannot achieve coordination between liquid level control and PDR control. PDR cannot be guaranteed to be between 1.5 and 3. H∞ control relaxes the liquid level tracking performance and improves PDR tracking accuracy; however, it relies on an accurate model, resulting in high maintenance costs for industrial systems.
[0044] 2. This invention takes into account both external disturbances and actuator failures during system operation. By introducing a fault compensation term in the controller, it can effectively recover the tracking performance degradation caused by system actuator failure, ensuring that the system has good fault tolerance while taking into account the optimal control performance of the system. Attached Figure Description
[0045] Figure 1 This is a control block diagram of the cascade structure of the oil-water separation system in this invention. Detailed Implementation
[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:
[0047] The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning includes the following steps:
[0048] S1. Establish an oil-water separation system model with interference and actuator failure. The interference is the flow rate of the oil-water mixture, and the actuator failure model is established based on the characteristics of the valve.
[0049] S1.1 Establish an actuator fault model;
[0050] During actual oil removal processes, valves are prone to malfunctions due to prolonged use, improper operation, or excessive operation. These malfunctions include insensitive switching and failure to achieve the required opening degree. Therefore, the identified fault types are valve ineffectiveness loss and offset failure. External interference refers to the flow rate of the oil-water mixture in the system.
[0051] Based on the characteristics of the valve, assume the control input u i For i = 1, 2, there is partial loss of validity and bias fault:
[0052]
[0053] Among them, u Δi β represents the input uncertainty caused by actuator failure. i The efficiency loss coefficient, For bias;
[0054] S1.2 Establish mathematical models for the water level control subsystem and the PDR control subsystem, and simultaneously provide a system model with external disturbances and actuator failures; y = l
[0055] Considering that the overflow rate at the overflow valve in the hydrocyclone is significantly less than the downstream rate at the downstream valve, the influence of the overflow rate on the water level control can be ignored. Therefore, the control inputs at the downstream valve and the overflow valve can be decoupled, corresponding to water level control and PDR control respectively. Adopting a cascaded control framework can reduce the nonlinear characteristics of the system and better reflect the actual operation of the system, making it easier to establish a mathematical model.
[0056] like Figure 1 As shown, the oil-water separation system can be viewed as a cascaded system consisting of a water level control subsystem and a PDR (pressure drop rate) control subsystem;
[0057] In the water level control subsystem:
[0058] u = u1, u Δ =u Δ1 ;
[0059] Where l represents the liquid level, ΔP u =P i -P u and These are the pressure drop and the rate of change of pressure drop at the relief valve, respectively, P i It is the valve pressure. It is the valve pressure change rate, u1 is the control signal for opening the relief valve, u Δ =u Δ1 The input uncertainty is caused by actuator failure; it is given in S1.1.
[0060] In the PDR control subsystem, system performance is related to ΔP u The changes are very sensitive. In practice, the required PDR is allowed to be within the desired range. Similar to the water level control subsystem, we can also use robust H-infinite control for the PDR control subsystem, described using a linear system model, with PDR... d ×ΔP u F is the reference value for tracking. in This is system interference;
[0061] The state of the PDR control subsystem is: Where, ΔPo =P i -P o and These represent the pressure drop and the rate of change of pressure drop at the downstream valve, respectively. u2 is the control signal for opening the downstream valve. Δ =u Δ2 The input uncertainty is caused by an actuator malfunction, and the output value is y = ΔP. o Therefore, the controller design for the water level and PDR control subsystems can be studied within the same system framework.
[0062] The water level control subsystem and PDR control subsystem models with external disturbances and actuator failures are both described by the following linear system:
[0063]
[0064] Where x is the system state, y is the system output, u is the control signal for valve opening, and w = F in It is external interference, ρ(tt) s ) is a step function, t s The time when the fault occurs satisfies t > t s At that time, ρ(tt) s ) = 1, otherwise ρ(tt) s ) = 0; u Δ The input uncertainty is caused by system failure, given by the failure model established in S1.1; system matrices A, B, C, and E are unknown system matrices.
[0065] S2. The cooperative control problem of the oil-water separation system model is expressed as a cascaded fault-tolerant H∞ tracking control problem, which reduces the sensitivity of PDR to inflow disturbances while reducing the water level tracking accuracy.
[0066] Robust H∞ control is applied to the oil-water separation system model. While satisfying the H∞ performance, the tracking performance is relaxed, making the changes in the inflow rate of the hydrocyclone and the control input u2 less sensitive to disturbance fluctuations.
[0067] The control objective is to make the output y track a linear reference trajectory r. d To achieve tracking, we define ξ = [x T ,r d T ] T When the system is fault-free, the augmented system is constructed based on the linear system model given in S1 as follows:
[0068]
[0069] in, e = yr dIt is the tracking error, w = F in It's external interference.
[0070] S3. The H∞ tracking control problem is transformed into a two-role zero-sum differential game problem. By establishing the game algebra Riccati equation (GARE), a model-based fault-tolerant control (FTC) solution is obtained, and the optimal control solution is obtained.
[0071] According to the L2 gain condition Define a performance function:
[0072]
[0073] Where α is a discount factor that satisfies α>0, Q and R are positive definite matrices, and γ>0 represents the amount of attenuation from the perturbation input w(t) to the defined performance output variable z(t);
[0074] Based on the performance function, the H∞ tracking control problem is treated as a two-player zero-sum game, and the Nash equilibrium solution (u) is found. * ,w * This makes the closed-loop system stable and robust. For any disturbance w, the closed-loop control system satisfies the L2 gain condition, that is, the tracking error e converges to the origin neighborhood whose boundary depends on γ.
[0075] When the system is fault-free, the GARE equations can be derived, and the optimal control strategy can be given; for model-based FTC, actuator faults can be addressed by adding an estimate term u to the controller. c To compensate.
[0076] S4. Construct a model-independent, fault-tolerant H∞ controller based on the optimal control solution;
[0077] Since the system dynamics are unknown, it is necessary to use u. * The term containing the system matrix B in the model-based FTC derived in S3 is replaced by... u * It is learned through reinforcement learning algorithms.
[0078] S5. A non-policy reinforcement learning algorithm is proposed to find the solution of GARE in a data-driven manner. The learned model-free solution is consistent with the model-based GARE solution, realizing actuator fault compensation and restoring the water level and PDR value to the near-normal optimal solution. That is, it restores the tracking performance degradation caused by the system actuator failure, ensuring that the system has good fault tolerance performance while taking into account the optimal control performance of the system.
[0079] Reinforcement learning methods can be divided into policy-based algorithms and non-policy-based algorithms. The former requires updating the old policy with the learned new policy in each iteration to generate new data. However, the interference policy w in the system of this invention... * It cannot be used for policy updates because it cannot control disturbances (i.e., inflows into Fin) to follow the desired policy. The latter uses data generated by a fixed control policy, such as a PID controller, and the policy iteration is independent of the system control loop. Therefore, non-policy reinforcement learning is more suitable for solving the tracking control problem of oil-water separation systems.
[0080] S5 specifically includes the following steps:
[0081] S5.1 collects system data, including system state information and measured disturbance data, to update the Bellman equation in the non-policy reinforcement learning algorithm and iteratively learns the system's control gain K. u * Interference gain K w * Thus, the optimal control strategy can be obtained;
[0082] Applying non-policy reinforcement learning algorithms to learn u * The core idea is to update the Bellman equation as follows to obtain K. u * ,K w * :
[0083]
[0084] Among them, V j It has the form of a quadratic form: V j =ξ T P j Substituting ξ into the above equation, we can further obtain P. j That is, the solution to the GARE equation, which is consistent with the model-based solution obtained in S3;
[0085] The adaptive update law of fault signals is applied to directly obtain the reinforcement learning-based FTC, and the optimal solution obtained by the non-policy reinforcement learning algorithm can converge to the model-based GARE solution, that is, the learned model-free solution is consistent with the model-based GARE solution.
[0086] S5.2 Based on the online parameter estimation information of the fault signal provided by the adaptive law, a fault-tolerant compensation controller is constructed on the basis of the learned optimal control strategy, and a model-free fault-tolerant compensation control scheme is given to compensate for the fault.
[0087] The designed fault-tolerant controller based on reinforcement learning can restore the water level and PDR value to near-normal optimal solutions, that is, restore the tracking performance degradation caused by the failure of the system actuator, and ensure that the system has good fault tolerance performance while taking into account the optimal control performance of the system.
[0088] Example
[0089] First, the model parameters of the oil-water separation system are obtained through system identification, including the water level control subsystem and the PDR control subsystem, such as the system matrix A, B, C, E.
[0090] The expected reference water level is l d =0.15m, the expected PDR value is PDR d =2.25, the safe range is [0.15, 3]. Using a model-based method, the Nash equilibrium K of the water level control subsystem and the PDR control subsystem can be obtained. u * ,K w * A fixed PID control strategy is applied to the water level control subsystem and the PDR control subsystem to generate system data. Using this system data, the K value is iteratively learned using the non-policy reinforcement learning algorithm proposed in S5.1. u j ,K w j They can all converge to K with small errors. u * ,K w * This demonstrates that the proposed RL algorithm can accurately learn the optimal control strategy.
[0091] Assume that at 800s, the system experiences simultaneous valve failure and biasing fault, with the valve opening between (0, 1) and both water level and PDR exceeding safe limits. Then, the proposed model-free FTC is used to compensate for the actuator fault, restoring the water level and PDR to near-normal optimal solutions. Based on the actuator faults and estimation results for the downstream and relief valves, it can be shown that the adaptive law is effective for estimating actuator faults. Therefore, it can be concluded that the proposed reinforcement learning-based FTC solution is effective.
[0092] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
[0093] In summary, this invention enables the system to autonomously learn the optimal control strategy from operational data, improving PDR tracking accuracy without relying on a model, and ensuring both good fault tolerance and optimal control performance.
Claims
1. A fault-tolerant H-infinite tracking control method for an oil-water separation system based on reinforcement learning, characterized in that: Includes the following steps: S1. Establish an oil-water separation system model with interference and actuator failure. The interference is the flow rate of the oil-water mixture, and the actuator failure model is established based on the characteristics of the valve. S2. The cooperative control problem of the oil-water separation system model is expressed as a cascaded fault-tolerant problem. The tracking control problem aims to reduce the water level tracking accuracy while simultaneously decreasing the sensitivity of the pressure drop rate to inflow disturbances; S3, will The tracking control problem is transformed into a two-role zero-sum differential game problem. By establishing the game algebra Riccati equation, a model-based fault-tolerant control solution is obtained, and the optimal control solution is obtained. S4. Construct a model-independent fault-tolerant system based on the optimal control solution. Controller; S5. A non-strategy reinforcement learning algorithm is proposed to find the solution of the Riccati equation in a data-driven manner. The learned model-free solution is consistent with the model-based game algebra Riccati equation solution, and the optimal solution that restores the water level and pressure drop rate to normal is obtained, thus realizing the fault compensation of the actuator.
2. The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning according to claim 1, characterized in that: S1 includes the following steps: S1.1 Establish an actuator fault model; Based on the characteristics of the valve, assume the control input There are partial loss of effectiveness and bias faults: in, Input uncertainty caused by actuator failure, The efficiency loss coefficient, For bias; S1.2 Establish mathematical models for the water level control subsystem and the PDR control subsystem, and also provide a system model with external disturbances and actuator failures: ; Wherein, PDR represents the pressure drop rate; The oil-water separation system can be viewed as a cascaded system consisting of a water level control subsystem and a PDR control subsystem. In the water level control subsystem: , , ; in, Indicates liquid level. and These are the pressure drop and the rate of change of pressure drop at the relief valve, respectively. It is the valve pressure. It is the rate of change of valve pressure. It is the control signal for opening the relief valve. The input uncertainty is caused by actuator failure; The state of the control subsystem is ,in, and These are the pressure drop and the rate of change of pressure drop at the downstream valve, respectively. It is the control signal for opening the downstream valve. The input uncertainty is caused by actuator failure, and the output value is ; The water level control subsystem is susceptible to external interference and actuator failure. The control subsystem models are all described using the following linear system description: in, It is the system status. It is system output. It is the control signal for opening the valve. It's external interference. It is a step function. It is the time when the fault occurs, which satisfies... hour, ,otherwise ; The input uncertainty is caused by system faults, given according to the fault model established in S1.1; system matrix. , , , It is an unknown system matrix.
3. The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning according to claim 1, characterized in that: In S2, robust applications are used in the oil-water separation system model. Control, in order to satisfy While maintaining performance, relax tracking performance to allow for changes in the inflow rate and control input of the hydrocyclone. It is less sensitive to fluctuations caused by disturbances; The control objective is to make the output Tracking a linear reference trajectory To achieve tracking, define When the system is fault-free, the augmented system is constructed based on the linear system model given in S1 as follows: in, , , , , It is tracking error. It's external interference.
4. The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning according to claim 1, characterized in that: In S3, according to Gain condition Define a performance function: in, To meet Discount factor, and It is a positive definite matrix. Indicates the input from the disturbance To the defined performance output variable The amount of attenuation; Based on performance functions The tracking and control problem can be viewed as a two-player zero-sum game; the goal is to find the Nash equilibrium solution. This makes the closed-loop system stable and robust against any disturbance. Closed-loop control systems all meet Gain condition, i.e., tracking error Convergence to boundary depends The origin neighborhood; When the system is fault-free, the Riccati equation can be derived, and the optimal control strategy can be given; for model-based fault-tolerant control, actuator faults are addressed by adding an estimation term to the controller. To compensate.
5. The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning according to claim 4, characterized in that: In S4, use Replace the terms in the model-based fault-tolerant controller derived in S3 that contain the system matrix B. It is learned through reinforcement learning algorithms.
6. The fault-tolerant H-infinite tracking control method for oil-water separation systems based on reinforcement learning according to claim 1, characterized in that: S5 includes the following steps: S5.1 collects system data, including system state information and measured disturbance data, to update the Bellman equation in the non-policy reinforcement learning algorithm and iteratively learn the system's control gain. Interference gain Thus, the optimal control strategy can be obtained; The Bellman equation is updated iteratively as follows to obtain : in, It has a quadratic form: Substituting into the above formula, we can further obtain... That is, the solution to the Riccati equation, which is consistent with the optimal solution of the model-based fault-tolerant control in S3; S5.2 Based on the online parameter estimation information of the fault signal provided by the adaptive law, a fault-tolerant compensation controller is constructed on the basis of the learned optimal control strategy, and a model-free fault-tolerant compensation control scheme is given to compensate for the fault.