Organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning

By constructing a multi-field coupled dynamic model and closed-loop control link through deep reinforcement learning, the modeling bias and dynamic coupling interference problems of the organic Rankine cycle system are solved, achieving high precision and fast response under all operating conditions, and improving the energy conversion efficiency and adaptability of the system.

CN121934385APending Publication Date: 2026-04-28WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2026-01-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing control methods for organic Rankine cycle systems suffer from several drawbacks. Simplified modeling leads to significant deviations between the control model and the actual dynamic characteristics of the multi-field coupling system. Static decoupling design struggles to suppress dynamic coupling disturbances. Fixed disturbance handling methods cannot adapt to random fluctuations in industrial waste heat and sudden load changes. Manually tuned parameters cannot be dynamically adjusted, making it difficult to achieve a balance between robustness, dynamic response speed, and energy consumption optimization.

Method used

A multi-field coupled dynamic model is constructed based on deep reinforcement learning. Inter-channel coupling is eliminated through a multivariable control matching system and a 3×3 invertible decoupling matrix. An independent extended state observer is configured to construct a finite-time robust control framework. A single Actor + dual Critic network architecture is used for parameter adaptive optimization to form a closed-loop control link.

Benefits of technology

It achieves high precision, rapid response and safe operation of the ORC system under all operating conditions, significantly improves energy conversion efficiency and system adaptability, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121934385A_ABST
    Figure CN121934385A_ABST
Patent Text Reader

Abstract

The invention discloses an organic Rankine cycle active-disturbance-rejection control method based on deep reinforcement learning, and belongs to the technical field of industrial process control and waste heat utilization, and the method comprises the following steps: S1, constructing a full-process multi-field coupling dynamic model; s2, defining the security constraint boundary of each parameter to obtain a multivariable control matching system; s3, eliminating static and dynamic coupling between channels by constructing a 3 * 3 reversible decoupling matrix to obtain M-ESO; s4, forming an FTRC control framework by constructing a finite time control law; s5, establishing a network architecture, and setting a multi-target reward function and a priority experience playback mechanism to obtain an IDRL parameter adaptive optimization model; and S6, through steady-state and dynamic working condition simulation verification, all-working-condition robust control is realized. By the adoption of the method, the problems that multivariable static and dynamic coupling interference exists in an ORC system, internal and external complex disturbance exists, and control parameters cannot be dynamically optimized along with working conditions are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process control and waste heat utilization technology, and in particular to an organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning. Background Technology

[0002] Organic Rankine cycle (ORC) systems are core energy conversion devices in the fields of low-grade waste heat recovery and renewable energy utilization. Their operational stability and control accuracy directly affect energy conversion efficiency and equipment safety. Currently, the mainstream control methods for ORC systems include traditional PID control and conventional model predictive control.

[0003] The aforementioned existing technologies have inherent defects: simplified modeling leads to a large deviation between the control model and the actual multi-field coupling dynamic characteristics of the system; static decoupling design is difficult to suppress dynamic coupling interference, affecting control accuracy; fixed disturbance handling methods cannot adapt to complex scenarios such as random fluctuations in industrial waste heat and sudden load changes, limiting the ability to resist disturbances; parameters that are manually tuned or optimized for a single objective cannot be dynamically adjusted according to the operating conditions, making it difficult to achieve an effective balance between control robustness, dynamic response speed and energy consumption optimization. Summary of the Invention

[0004] The purpose of this invention is to provide an organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning to solve the above-mentioned technical problems.

[0005] To achieve the above objectives, this invention provides an organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning, comprising the following steps: S1. Based on the structural composition and core thermodynamic principles of the Organic Rankine Cycle (ORC) system, a full-process multi-field coupled dynamic model is constructed by integrating the phase change characteristics of heat exchangers, the dynamic correlation law of components, and the multi-field coupling effect. This model includes the evaporator, condenser, working fluid pump, expander, and liquid storage tank. S2. Based on the multi-field coupled dynamic model in step S1, by clarifying the mapping relationship between the core controlled parameters of the system and the corresponding control variables, the safety constraint boundaries of each parameter are defined, and a multivariable control matching system adapted to the ORC system is obtained. S3. Based on the multivariable control matching system in step S2, the static and dynamic coupling between channels is eliminated by constructing a 3×3 invertible decoupling matrix, and an independent extended state observer is configured for each channel to obtain the multivariable decoupled extended state observer M-ESO. S4. Based on the M-ESO disturbance estimation results in step S3, a highly robust FTRC control framework is formed by constructing a finite-time control law including a disturbance compensation term, introducing control quantity saturation constraints and an anti-integral saturation mechanism. S5. Based on the FTRC control framework of step S4, an improved deep reinforcement learning IDRL parameter adaptive optimization model is obtained by constructing a single Actor + dual Critic network architecture, setting a multi-objective reward function and a priority experience replay mechanism. S6. Based on the multi-field coupled dynamic model of step S1, the M-ESO of step S3, the FTRC control framework of step S4, and the IDRL parameter adaptive optimization model of step S5, the ORC system is verified through steady-state and dynamic operating condition simulations to achieve robust control of the ORC system under all operating conditions.

[0006] Preferably, the model construction methods in step S1 are as follows: Evaporator modeling: Using the moving boundary method, the model is divided into subcooled, two-phase, and superheated zones. The model is constructed based on the principles of mass conservation, energy conservation, and heat transfer. The formula is as follows: ; ; ; ; in, The working fluid velocity inside the pipe; The coordinates are along the pipe length direction; For time; , The average density and average specific enthalpy of the working fluid in the corresponding phase region are, in order. This refers to the pressure inside the evaporator. The inner diameter of the heat exchanger; This refers to the tube wall temperature of the corresponding phase region; The working fluid temperature; The convective heat transfer coefficient between the heat source and the pipe wall; This is the cavitation coefficient; This is the average density of the working fluid in the two-phase region of the evaporator. , These are the densities of the saturated gas phase and liquid phase, respectively. , These are the enthalpy of the saturated gas phase and liquid phase, respectively. This represents the average specific enthalpy of the working fluid in the two-phase region of the evaporator. Condenser modeling: The condenser is divided into superheated, condensing, and subcooled zones using the controlled volume method. Axial heat conduction is ignored, and the inlet and outlet pressures are assumed to be equal. Based on the general phase region energy conservation and cooling water side energy conservation, the formulas are as follows: ; ; in, The internal energy of the working fluid; The volume of the phase region; , For the inlet and outlet mass flow rates of the phase zone; , The specific enthalpy of the working fluid at the inlet and outlet of the corresponding component; Heat dissipation in the phase region; For the condenser Heat exchange of cooling water in each phase region; This refers to the cooling water flow rate; The specific heat capacity of cooling water at constant pressure; , These are the inlet and outlet temperatures of the cooling water, respectively. Expander modeling: Treat the scroll expander as a scroll compressor operating in reverse and establish a steady-state output model; Working fluid pump modeling: A multi-stage centrifugal pump is used to establish a dynamic correlation model of flow rate-head-power under variable speed, with the following formula: Rated speed and head formula: ; Variable speed flow rate formula: ; Variable speed head formula: ; Export enthalpy formula: ; in, , , This refers to the pump characteristic coefficient; Rated head; Rated flow rate; This refers to the actual rotational speed; Rated speed; This represents the actual traffic volume. For actual head; The specific volume of the working fluid at the pump inlet; , The inlet and outlet pressures of the pump; For pump adiabatic efficiency; Modeling of the liquid storage tank: Following the laws of mass and energy conservation to achieve the working fluid buffering function, the formulas are as follows: mass conservation formula: ; Energy conservation formula: ; in, The mass of the working fluid in the storage tank; The total energy of the working fluid in the storage tank; It facilitates heat exchange between the storage tank and the environment.

[0007] Preferably, in step S1, by quantifying the dynamic correlation between temperature, pressure, and flow rate, the law of working fluid state transfer between the evaporator and condenser is clarified, the correlation between key coupling parameters and operating conditions is established, and the coupling coefficient between working fluid pump speed and evaporation pressure is obtained. Expansion valve opening degree-superheat coupling coefficient Cooling water flow rate - condensing pressure coupling coefficient .

[0008] Preferably, step S2 includes the following specific steps: S21. Based on the energy conversion efficiency and operational safety requirements of the ORC system, the core controlled parameters are quantitatively defined and the expected target values ​​for each controlled parameter are set. The core controlled parameters include superheat, the evaporation pressure of the evaporator in step S1, and the condensation pressure output by the condenser condensation zone model in step S1. The formula for calculating superheat is: ; in, The working fluid temperature at the evaporator outlet. Evaporation pressure The corresponding saturation temperature; This refers to the saturation pressure in the two-phase region of the evaporator. S22. By matching independent control quantities corresponding to the number of controlled parameters, establish a quantitative dynamic correlation between the control quantities and the controlled parameters. The control quantities include the working fluid pump speed. Expansion valve opening and cooling water flow rate The formula for the association relationship is: ; ; ; in, , , These are dimensionless coupling coefficients; , , The changes in evaporation pressure, superheat, and condensation pressure relative to their corresponding expected target values ​​are, in order. , , The changes in the working fluid pump speed, expansion valve opening, and cooling water flow rate are, in order. S23. Based on the equipment's safe operation requirements and the thermodynamic properties of the working fluid, define the value constraints of each core controlled parameter and its expected target value, as well as the safety constraint boundaries of the control quantities. These constraints include superheat greater than 0K, evaporation pressure less than the critical pressure of the working fluid, and condensation pressure greater than standard atmospheric pressure. Both the evaporation pressure and its expected target value are less than the critical pressure of the working fluid, while both the condensation pressure and its expected target value are greater than standard atmospheric pressure. Furthermore, the control quantities satisfy the formula: ; ; ; in, , The minimum and maximum safe operating speeds of the working fluid pump are, in order. , The minimum and maximum safe flow rates for cooling water are listed in order.

[0009] Preferably, step S3 includes the following specific steps: S31. Based on the multivariate quantization mapping relationship and constraint boundary from step S2, obtain the static gain by fitting the steady-state data of the multi-field coupled dynamic model from step S1, and construct a 3×3 invertible decoupling matrix, as shown in the formula: ; in, For the first The control quantity affects the first... The static gain of a controlled parameter; S32, Decoupling matrix based on step S31 Establish a mapping between virtual control variables and actual control variables, and introduce virtual control variables as follows: Establish and control the actual amount The mapping relationship is given by the formula: To obtain independent control inputs for each controlled channel, where , , These are virtual control quantities for the evaporation pressure, superheat, and condensation pressure channels, respectively. It is the inverse of the decoupling matrix; S33. Based on the independent channel partitioning in step S32, a multi-channel collaborative observation mechanism is constructed by configuring an extended state observer for each channel, and the state equation is established as follows: ; in, For the first The actual output of each controlled parameter; For total disturbance; Equivalent gain; The observer equation is designed, and the formula is as follows: ; in, , They are respectively , Observed values; For observation error, and ; , For observer gain; It is a saturation function; S34. Based on the observer equations in step S33, the observer gain is tuned using the pole placement method, with the following formula: ; ; in, For observer bandwidth; The poles of the observer's characteristic equation are placed at... This is to ensure that the observation error convergence time is ≤5s.

[0010] Preferably, step S4 includes the following specific steps: S41. Based on the actual output of the controlled parameters of each channel in step S33 Compared with observed values Combining the expected target value of the controlled parameter defined in step S21, the tracking error and error derivative are defined as follows: ; ; in, For tracking error; The desired target value of the controlled parameter; These are the observed values ​​of the controlled parameters; The differential of the tracking error of the controlled channel; The derivative of the desired target value of the controlled parameter; The derivative of the observed values ​​of the controlled parameter; Controlled channel; S42. Total disturbance observations based on the output of step S33 Combining finite-time stability theory, a control law with disturbance compensation is constructed to achieve rapid convergence of the controlled parameter. The formula is as follows: ; in, The total disturbance estimated by M-ESO; , , To control the gain; The convergence coefficients are finite-time convergence coefficients; It is a saturation function; It is a symbolic function; S43. Based on the control quantity safety constraint boundary defined in step S23, the virtual control quantity... The mapped actual control quantity is saturated and limited to ensure it does not exceed the physical limits of the equipment. The formula is: ; in, This is the actual control quantity; , To control the upper and lower limits of the quantity; It is a saturated function, and when Time output ,when Time output Otherwise, output the original value.

[0011] S44. Integrate the optimization results of sub-steps S42-S43 to output the finite-time robust controller FTRC.

[0012] Preferably, step S44 further includes optimizing the control law by introducing an integral separation strategy to avoid integral saturation under long-term disturbances or large errors, as shown in the formula: ; ; in, The integral separation threshold; This is the integral gain; The optimized virtual control quantity inverse matrix of the decoupling matrix Convert to actual control quantity Synchronous output control gain ~ Convergence coefficient saturation threshold Integral separation threshold .

[0013] Preferably, step S5 includes the following specific steps: S51. Based on the FTRC control framework in step S44, construct a single Actor + dual Critic deep reinforcement learning network architecture, with the observer gain of step S3M-ESO and the core control parameters of FTRC in step S44 as the optimization objects. S52. Based on the control performance requirements of FTRC in step S4, the parameter constraint boundaries defined in step S2, and the energy consumption optimization objectives of the ORC system, the parameter optimization effect is quantitatively evaluated through a multi-objective reward function, as shown in the formula: ; in, To control accuracy rewards, and , Let be the absolute error of the integration time, and , For precision weighting coefficients; For dynamic response rewards, and , To adjust the time, , For overshoot, , , For response weighting coefficients; The reward is a parameter constraint, and the optimized parameters simultaneously satisfy... , , , hour, ,otherwise ; It is an energy consumption penalty item, and ; , , , For normalized weights; S53. Set a priority experience replay buffer to store experience samples generated by the interaction between the IDRL model and the ORC training environment built by the multi-field coupled dynamic model in step S1. The formula is: ; ; in, for Sample priority at any given time; This refers to timing difference error; This is the priority offset; Discount factor; For the output of the dual Critic network Mean value of state at any given time; for System status at all times; For the reward function calculated based on step S52 Momentary reward value, The system state at the next moment after the parameters are applied is the output of the model in step S1. S54. Based on the training environment built by the multi-field coupled dynamic model of the ORC system in step S1, after initializing the Actor and dual-Critic network parameters and experience buffer, the training iteratively follows the process of substituting the optimized parameters output by the Actor network into the M-ESO in step S3 and the FTRC in step S4, simulating the system response and calculating the reward using the dynamic model in step S1, storing experience samples and updating priorities, sampling and optimizing the dual-Critic network parameters, and updating the Actor network parameters using policy gradients, until the reward function is reached. Training is stopped when the fluctuation amplitude is ≤5% and convergence occurs in multiple consecutive iterations. S55. After training convergence, output the improved deep reinforcement learning IDRL parameter adaptive optimization model and simultaneously output the initial optimized core parameter set: M-ESO observer gain and FTRC control parameters. At the same time, construct a parameter real-time feedback update interface, which can dynamically iteratively optimize the above parameters based on the real-time operating status of the system and simultaneously feed back to M-ESO in step S3 and FTRC in step S4.

[0014] Preferably, step S6 includes the following specific steps: S61. Based on the multi-field coupled dynamic model of step S1, the M-ESO of step S3, the FTRC of step S4 and the IDRL parameter adaptive optimization model of step S5, complete the interface adaptation and signal linkage of each module, and construct the closed-loop control link of "state perception-disturbance estimation-control output-parameter optimization". S62. Based on the actual industrial operation scenarios of the ORC system, and the parameter safety constraint boundary defined in step S2 and the model training operating condition range in step S5, two types of verification operating conditions covering the entire operating range are set up, including the steady-state operating condition with rated waste heat input and the dynamic operating condition with random fluctuations in industrial waste heat. S63. Input the operating parameters from step S62 into the closed-loop control link, start the simulation and run iteratively according to the established closed-loop process, and record the core data at each moment synchronously. S64. Based on the full data of the simulation run records, using the data from step S52... , , The control performance quantification index is used to conduct multi-dimensional performance evaluation of the three controlled channels (i=1,2,3) under steady-state and dynamic conditions. S65. Summarize the multi-dimensional performance evaluation results from step S64 to form a full-condition verification report, clarifying the closed-loop control link under steady-state conditions. Overshoot Adjusting time s, under dynamic operating conditions Overshoot Adjusting time Output a robust control scheme for the ORC system under all operating conditions, along with an optimized set of core parameters. .

[0015] Therefore, the present invention employs the above-mentioned organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning, which has the following beneficial effects: 1. By integrating the phase change characteristics of heat exchangers, the dynamic correlation laws of components, and the multi-field coupling effect, a full-process multi-field coupled dynamic model covering the evaporator, condenser, working fluid pump, expander, and liquid storage tank is constructed. The phase regions of each core component are accurately divided and quantitative formulas are established based on the thermodynamic conservation principle. At the same time, the dynamic correlation between temperature, pressure, and flow rate is quantified to obtain key coupling coefficients. This model can accurately replicate the actual dynamic operating characteristics of the ORC system, providing a realistic basic data support for subsequent control system design, decoupling observation, and parameter optimization, ensuring that the subsequent control strategy is highly adapted to the actual operating conditions of the system.

[0016] 2. The 3×3 reversible decoupling matrix constructed based on the multivariable control matching system can effectively eliminate the static and dynamic coupling effects between the evaporation pressure, superheat, and condensation pressure channels, solving the problem that the static decoupling method cannot cope with the dynamic correlation of parameters during operation. The multivariable decoupled extended state observer formed by configuring independent extended state observers for each channel, and tuning the gain through the pole placement method, realizes real-time and accurate estimation of internal and external disturbances, and the observation error convergence time is ≤5s. It not only provides an independent control environment for each controlled channel, but also can detect and quantify disturbances in advance, significantly improving the system's ability to resist complex coupled interference and disturbances such as industrial waste heat fluctuations and load changes.

[0017] 3. The finite-time robust control framework is based on the disturbance estimation results of M-ESO and incorporates disturbance compensation terms to construct a finite-time control law, ensuring that the controlled parameter converges quickly under complex operating conditions. At the same time, it introduces control quantity saturation constraints and anti-integral saturation mechanisms, and avoids integral saturation problems under long-term disturbances or large errors through integral separation strategies, while strictly adhering to the equipment safety operation constraint boundaries. This framework not only solves the problems of slow response and large overshoot of traditional control laws, but also prevents the control quantity from exceeding the physical limits of the equipment, achieving dual protection of rapid tracking of the controlled parameter and safe system operation, and taking into account control accuracy, dynamic response speed and equipment operation safety.

[0018] 4. The single-actor + dual-critic network architecture focuses on optimizing the M-ESO observer gain and the core control parameters of the FTRC. The multi-objective reward function comprehensively considers control accuracy, dynamic response, parameter constraints, and system energy consumption. The priority experience playback mechanism improves training efficiency and avoids training oscillations. Through continuous interaction with the ORC system training environment, dynamic adaptive optimization of core parameters is achieved. Parameters can be adjusted according to changes in operating conditions without manual intervention. While ensuring control robustness, a dynamic balance between control performance and energy consumption is achieved, reducing system operation and maintenance costs.

[0019] 5. By integrating a closed-loop link constructed from a multi-field coupled dynamic model, M-ESO, FTRC, and IDRL, the entire process of real-time perception of core state parameters, accurate estimation of disturbances, rapid calculation of control output, and dynamic optimization of parameters is realized. This link covers both steady-state conditions with rated waste heat input and dynamic conditions with random fluctuations in industrial waste heat. Through full-condition simulation verification, the high-performance control indicators under steady-state and dynamic conditions are clarified, solving the problem that traditional control schemes can only adapt to simple conditions and lack robustness. It ensures that the ORC system maintains high control accuracy, low overshoot, and fast response speed throughout the entire operating range, significantly improving the system's energy conversion efficiency, long-term stable operation capability, and industrial scenario adaptability.

[0020] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0021] Figure 1 A flowchart of an organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning provided by the present invention; Figure 2 This is a schematic diagram of the structure of an ORC-following waste heat power generation system corresponding to the organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning provided by the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of the present invention and are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.

[0023] It should be noted that the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as a process, method, system, product, or server that includes a series of steps or units, not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.

[0024] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0025] Currently, the mainstream control methods for organic Rankine cycle systems employ traditional PID control and conventional model predictive control. These methods simplify the phase change characteristics of heat exchangers and the dynamic relationships between components during modeling, building models based only on a single physical field or local parameters. Decoupling strategies are mostly designed for static coupling, and disturbance handling relies on fixed mathematical models or simple feedback correction. Control parameters are determined through manual tuning or single-objective optimization. Verification conditions are concentrated in simple dynamic scenarios with steady-state or small fluctuations in rated waste heat input. These existing technologies have inherent flaws: simplified modeling leads to significant deviations between the control model and the actual multi-field coupled dynamic characteristics of the system; static decoupling design struggles to suppress dynamic coupling interference between parameters; fixed disturbance handling methods cannot adapt to complex scenarios such as random fluctuations in industrial waste heat and sudden load changes; and manually tuned or single-objective optimized parameters cannot be dynamically adjusted according to operating conditions, making it difficult to achieve an effective balance between control robustness, dynamic response speed, and energy consumption optimization.

[0026] Based on the above analysis, this invention is designed, see appendix. Figure 1-2 An organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning includes the following steps: S1. Based on the structure and thermodynamic core principles of the Organic Rankine Cycle (ORC) system, a full-process multi-field coupled dynamic model is constructed by integrating the phase change characteristics of heat exchangers, the dynamic correlation law of components, and the multi-field coupling effect. The model includes the phase region division of each component and the multi-parameter coupling mechanism. Specifically, the components used include: an evaporator of shell and tube type with an inner diameter of 25mm and a tube length of 5m; and a condenser of shell and tube type, divided into 3 phase zones with a total volume of 0.5m³. 3 1. A vortex expander with a rated power of 50kW and an efficiency of 0.85; 2. A multistage centrifugal pump with a rated flow rate of 10m³ / h. 3 / h, rated head of 50m, storage tank with a volume of 1m³ 3 ; The model construction methods in step S1 are as follows: Evaporator modeling: The moving boundary method is used, dividing the region into subcooled zone length, two-phase zone length, and superheated zone length. The model is constructed based on the principles of mass conservation, energy conservation, and heat transfer. The formula is as follows: ; ; ; ; in, The working fluid velocity inside the pipe; The coordinates are along the pipe length direction; For time; , The average density and average specific enthalpy of the working fluid in the corresponding phase region are, in order. This refers to the pressure inside the evaporator. The inner diameter of the heat exchanger; This refers to the tube wall temperature of the corresponding phase region; The working fluid temperature; The convective heat transfer coefficient between the heat source and the pipe wall; This is the cavitation coefficient; This is the average density of the working fluid in the two-phase region of the evaporator. , These are the densities of the saturated gas phase and liquid phase, respectively. , These are the enthalpy of the saturated gas phase and liquid phase, respectively. This represents the average specific enthalpy of the working fluid in the two-phase region of the evaporator. Condenser modeling: The condenser is divided into superheated, condensing, and subcooled zones using the controlled volume method. Axial heat conduction is ignored, and the inlet and outlet pressures are assumed to be equal. Based on the general phase region energy conservation and cooling water side energy conservation, the formulas are as follows: ; ; in, The internal energy of the working fluid; The volume of the phase region; , For the inlet and outlet mass flow rates of the phase zone; , The specific enthalpy of the working fluid at the inlet and outlet of the corresponding component; Heat dissipation in the phase region; For the condenser Heat exchange of cooling water in each phase region; This refers to the cooling water flow rate; The specific heat capacity of cooling water at constant pressure; , These are the inlet and outlet temperatures of the cooling water, respectively. Expander Modeling: Treating the scroll expander as a scroll compressor operating in reverse, a steady-state output model is established, with the following formula: ; in, For the working fluid flow rate; For expander efficiency; Specific enthalpy at the expander inlet; The specific enthalpy at the isentropic expansion outlet; rotational speed With expansion ratio working fluid flow rate Positive correlation with load Negative correlation; Working fluid pump modeling: A multi-stage centrifugal pump is used to establish a dynamic correlation model of flow rate-head-power under variable speed, with the following formula: Rated speed and head formula: ; Variable speed flow rate formula: ; Variable speed head formula: ; Export enthalpy formula: ; in, , , The pump characteristic coefficient is obtained by fitting the performance curve; Rated head; Rated flow rate; This refers to the actual rotational speed; Rated speed; This represents the actual traffic volume. For actual head; The specific volume of the working fluid at the pump inlet; , The inlet and outlet pressures of the pump; For pump adiabatic efficiency; Modeling of the liquid storage tank: Following the laws of mass and energy conservation to achieve the working fluid buffering function, the formulas are as follows: mass conservation formula: ; Energy conservation formula: ; in, The mass of the working fluid in the storage tank; The total energy of the working fluid in the storage tank; It facilitates heat exchange between the storage tank and the environment.

[0027] In step S1, by quantifying the dynamic correlation between temperature, pressure, and flow rate, the law of working fluid state transfer between the evaporator and condenser is clarified, the correlation between key coupling parameters and operating conditions is established, and the coupling coefficient between working fluid pump speed and evaporation pressure is obtained. Expansion valve opening degree-superheat coupling coefficient Cooling water flow rate - condensing pressure coupling coefficient The coupling coefficient is a dimensionless coupling coefficient. The fitting process is based on the least squares method. By collecting data on the changes in control quantity and controlled parameter under different operating conditions, a linear regression model is constructed and solved.

[0028] S2. Based on the multi-field coupled dynamic model in step S1, by clarifying the mapping relationship between the core controlled parameters of the system and the corresponding control variables, the safety constraint boundaries of each parameter are defined, and a multivariable control matching system adapted to the ORC system is obtained. S21. Based on the energy conversion efficiency and operational safety requirements of the ORC system, the core controlled parameters are quantitatively defined and the expected target values ​​for each controlled parameter are set. The core controlled parameters include superheat, the evaporation pressure of the evaporator in step S1, and the condensation pressure output by the condenser condensation zone model in step S1. The formula for calculating superheat is: ; in, The working fluid temperature at the evaporator outlet. Evaporation pressure The corresponding saturation temperature; This refers to the saturation pressure in the two-phase region of the evaporator. S22. By matching independent control quantities corresponding to the number of controlled parameters, establish a quantitative dynamic correlation between the control quantities and the controlled parameters. The control quantities include the working fluid pump speed. Expansion valve opening and cooling water flow rate The formula for the association relationship is: ; ; ; in, , , These are dimensionless coupling coefficients; , , The changes in evaporation pressure, superheat, and condensation pressure relative to their corresponding expected target values ​​are, in order. , , The changes in the working fluid pump speed, expansion valve opening, and cooling water flow rate are, in order. S23. Based on the equipment safety operation requirements and the thermodynamic characteristics of the working fluid, define the value constraints of each core controlled parameter and its expected target value, as well as the safety constraint boundaries of the control quantities, to ensure that the controlled parameters meet the system safety operation requirements while tracking their expected target values. The constraint boundaries include superheat greater than 0K, evaporation pressure less than the critical pressure of the working fluid, and condensation pressure greater than standard atmospheric pressure; both the evaporation pressure and its expected target value are less than the critical pressure of the working fluid, and both the condensation pressure and its expected target value are greater than standard atmospheric pressure; and the control quantities satisfy the formula: ; ; ; in, , The minimum and maximum safe operating speeds of the working fluid pump are, in order. , The minimum and maximum safe flow rates of cooling water are determined based on the cooling requirements of the condenser and the flow capacity of the pipeline.

[0029] S3. Based on the multivariable control matching system in step S2, a 3×3 invertible decoupling matrix is ​​constructed to eliminate the static and dynamic coupling between channels. An independent extended state observer is configured for each channel to estimate internal and external disturbances in real time, thus obtaining the multivariable decoupled extended state observer M-ESO. The specific steps of step S3 include: S31. Based on the multivariate quantization mapping relationship and constraint boundary from step S2, obtain the static gain by fitting the steady-state data of the multi-field coupled dynamic model from step S1, and construct a 3×3 invertible decoupling matrix, as shown in the formula: ; in, For the first The control quantity affects the first... The static gain of a controlled parameter , respectively, correspond to the controlled parameters, 1 represents 2 represents 3 represents , These correspond to control quantities, with 1 being... 、2 for 3 is Matrix invertibility is achieved through To ensure reversibility, if the determinant of the fitted matrix is ​​0 or close to 0, the static gain value is finely adjusted. S32, Decoupling matrix based on step S31 Establish a mapping between virtual control variables and actual control variables, and introduce virtual control variables as follows: Establish and control the actual amount The mapping relationship is given by the formula: To obtain independent control inputs for each controlled channel, where , , These are virtual control quantities for the evaporation pressure, superheat, and condensation pressure channels, respectively. It is the inverse of the decoupling matrix; S33. Based on the independent channel partitioning in step S32, a multi-channel collaborative observation mechanism is constructed by configuring an extended state observer for each channel, and the state equation is established as follows: ; in, For the first The actual output of each controlled parameter; The total disturbance includes internal and external disturbances as well as coupled interference; Equivalent gain; The observer equation is designed, and the formula is as follows: ; in, , They are respectively , Observed values; For observation error, and ; , For observer gain; It is a saturation function, and ; S34. Based on the observer equations in step S33, the observer gain is tuned using the pole placement method, with the following formula: ; ; in, The observer bandwidth is 5~10 rad / s; The poles of the observer's characteristic equation are placed at... To ensure that the observation error convergence time is ≤5s, the pole placement process is achieved by solving the characteristic equation. Implementation, in which The observer state matrix, For complex frequency variables, It is an identity matrix.

[0030] S4. Based on the M-ESO disturbance estimation results in step S3, a finite-time control law including a disturbance compensation term is constructed. Control quantity saturation constraints and anti-integral saturation mechanisms are introduced to ensure that the controlled parameters converge within a finite time and do not exceed the physical limits of the equipment, thus forming a highly robust FTRC control framework. The specific steps of step S4 include: S41. Based on the actual output of the controlled parameters of each channel in step S33 Compared with observed values Combining the expected target value of the controlled parameter quantized in step S21, the tracking error and error derivative are defined to provide basic error indices for control law design. The formulas are as follows: ; ; in, For tracking error; The desired target value of the controlled parameter; These are the observed values ​​of the controlled parameters; The differential of the tracking error of the controlled channel; The derivative of the desired target value of the controlled parameter; The derivative of the observed values ​​of the controlled parameter; It is a controlled channel, and , respectively corresponding , , ; S42. Total disturbance observations based on the output of step S33 Combining finite-time stability theory, a control law with disturbance compensation is constructed to achieve rapid convergence of the controlled parameter. The formula is as follows: ; in, The total disturbance estimated by M-ESO includes dynamic coupling terms and residual heat fluctuations; , , To control the gain, it is dimensionless and its value ranges from 10 to 50; The convergence coefficient is a finite-time coefficient, dimensionless, and takes values ​​from 0.5 to 2. It is a saturation function. To avoid sudden changes in control quantities; It is a symbolic function; S43. Based on the control quantity safety constraint boundary defined in step S23, the virtual control quantity... The mapped actual control quantity is saturated and limited to ensure it does not exceed the physical limits of the equipment. The formula is: ; in, This is the actual control quantity; , To control the upper and lower limits of the quantity; It is a saturated function, and when Time output ,when Time output Otherwise, output the original value.

[0031] S44. Integrate the optimization results of sub-steps S42-S43 to output the finite-time robust controller FTRC.

[0032] Step S44 also includes optimizing the control law by introducing an integral separation strategy to avoid integral saturation under long-term disturbances or large errors, thereby reducing overshoot and oscillations in evaporation pressure, superheat, and condensation pressure, while retaining integral action to eliminate steady-state errors and ensure that the system quickly and stably tracks the desired target values ​​of each controlled parameter. The formula is as follows: ; ; in, The integral separation threshold is dimensionless and ranges from 0.05 to 0.2. This is the integral gain, which is dimensionless and takes values ​​from 1 to 5. The optimized virtual control quantity inverse matrix of the decoupling matrix Convert to actual control quantity Synchronous output control gain ~ Convergence coefficient saturation threshold Integral separation threshold .

[0033] S5. Based on the FTRC control framework in step S4, by constructing a single Actor + dual Critic network architecture, setting a multi-objective reward function and a priority experience replay mechanism, the M-ESO gain and the core control parameters of FTRC are optimized, resulting in an improved deep reinforcement learning IDRL parameter adaptive optimization model. The specific steps of step S5 include: S51. Based on the FTRC control framework in step S44, a single Actor + dual Critic deep reinforcement learning network architecture is constructed. The observer gain in step S3M-ESO and the core control parameters of FTRC in step S44 are the optimization objects. The Actor network is used as the policy network and the dual Critic network is used as the value evaluation network. The Actor network takes the core state parameters from the multi-field coupled dynamic model in step S1 and the tracking error in step S4 as inputs, and outputs the optimized parameter set. The dual Critic networks evaluate the value of the Actor network's output strategy from the dimensions of control accuracy and system stability, respectively, avoiding optimization bias caused by a single evaluation dimension.

[0034] S52. Based on the control performance requirements of FTRC in step S4, the parameter constraint boundaries defined in step S2, and the energy consumption optimization objective of the ORC system, the effect of parameter optimization is quantitatively evaluated through a multi-objective reward function to provide guidance for network training. The formula is as follows: ; in, To control accuracy rewards, and , Let be the absolute error of the integration time, and , This is the accuracy weighting coefficient, dimensionless, with a value ranging from 0.01 to 0.1. The smaller the error, the greater the reward value. For dynamic response rewards, and , To adjust the time, , For overshoot, , , The response weight coefficient is dimensionless and ranges from 0.1 to 0.5. It penalizes excessive overshoot and slow response, ensuring the dynamic stability of the system. The reward is a parameter constraint, and the optimized parameters simultaneously satisfy... , , , hour, ,otherwise ; It is an energy consumption penalty item, and ; , , , The normalized weights are set to 0.4, 0.3, 0.2, and 0.1, respectively. S53. To improve network training efficiency and avoid training oscillations, a priority experience replay buffer is set up to store experience samples generated by the interaction between the IDRL model and the ORC training environment built by the multi-field coupled dynamic model in step S1. The formula is: ; ; in, for Sample priority at any given time; This refers to timing difference error; This is the priority offset, dimensionless, and takes a value of 0.01 to avoid a sample priority of 0. This is a discount factor, ranging from 0.9 to 0.99, balancing immediate rewards and future accumulated rewards; For the output of the dual Critic network Mean value of state at any given time; for System status at all times; For calculation based on the S52 reward function Momentary reward value, The output of model S1 shows the system state at the next moment after the parameters are applied; sampling is performed according to... Probabilistic sampling prioritizes learning high-value, high-biased samples to improve training convergence speed; S54. Based on the training environment built by the multi-field coupled dynamic model of the ORC system in step S1, after initializing the Actor / dual Critic network parameters and experience buffer, the training iteratively follows the process of substituting the optimized parameters output by the Actor network into the M-ESO in step S3 and the FTRC in step S4, the dynamic model in S1 to simulate the system response and calculate the reward, store experience samples and update priorities, sample and optimize the dual Critic network parameters, and update the Actor network parameters with policy gradients until the reward function is fully implemented. Training is stopped when the fluctuation amplitude is ≤5% and convergence occurs in 1000 consecutive iterations. S55. After training convergence, output the improved deep reinforcement learning IDRL parameter adaptive optimization model and simultaneously output the initial optimized core parameter set: M-ESO observer gain and FTRC control parameters. At the same time, construct a parameter real-time feedback update interface, which can dynamically iteratively optimize the above parameters based on the real-time operating status of the system and simultaneously feed back to M-ESO in step S3 and FTRC in step S4.

[0035] S6. Based on the multi-field coupled dynamic model of step S1, the M-ESO of step S3, the FTRC control framework of step S4, and the IDRL parameter adaptive optimization model of step S5, the core state parameters are output through the dynamic model, the disturbance is estimated in real time by M-ESO, the control output is calculated by FTRC, and the parameters are adaptively optimized by IDRL and updated, forming a closed-loop link of state perception-disturbance estimation-control output-parameter optimization. The system is verified by steady-state and dynamic operating condition simulation to achieve robust control of the ORC system under all operating conditions. S61. Based on the multi-field coupled dynamic model of step S1, the M-ESO of step S3, the FTRC of step S4 and the IDRL parameter adaptive optimization model of step S5, complete the interface adaptation and signal linkage of each module, and construct the closed-loop control link of "state perception-disturbance estimation-control output-parameter optimization". The core signal flow direction of the link is as follows: State perception: Step S1: Dynamic model outputs evaporation pressure overheating Condensing pressure and the tracking error defined in step S4 The M-ESO in step S3 and the IDRL model in step S5 are synchronously transmitted. Perturbation estimation: Step S3 involves real-time estimation of internal and external disturbances using the M-ESO method. And output to FTRC in step S4; Control output: Step S4's FTRC calls the optimized parameter set Θ from step S5's IDRL, combined with... Calculate the actual control quantity ; Parameter optimization: The control quantity acts on the dynamic model in step S1, causing changes in state parameters. The IDRL model in step S5 dynamically updates the parameters based on the new state and feeds them back to M-ESO in step S3 and FTRC in step S4, forming a continuous closed loop.

[0036] S62. Based on the actual industrial operation scenario of the ORC system, and the parameter safety constraint boundary defined in step S2 and the model training condition range in step S5, two types of verification conditions covering the entire operating range are set up, including the steady-state condition with rated waste heat input, which is used to verify the steady-state tracking accuracy and operational stability of the system, and the dynamic condition with random fluctuations in industrial waste heat. The waste heat fluctuation amplitude is set to ±20%, and random fluctuations in industrial waste heat and load change rate of 5% / s are simulated to simulate dynamic adjustment of power generation load, which is used to verify the system's anti-disturbance capability and dynamic response performance.

[0037] S63. Input the operating parameters from step S62 into the closed-loop control link, start the simulation and run iteratively according to the established closed-loop process. Set the simulation duration to 1000s and record the core data at each moment simultaneously. S64. Based on the full data of the simulation run records, using the data from step S52... , , Quantitative indicators of control performance were used to conduct multi-dimensional performance evaluations of the three controlled channels under steady-state and dynamic conditions. S65. Summarize the multi-dimensional performance evaluation results from step S64 to form a full-condition verification report, clarifying the closed-loop control link under steady-state conditions. Overshoot Adjusting time s, under dynamic operating conditions Overshoot Adjusting time To meet the full-condition operation requirements of the ORC system, a robust control scheme for the ORC system under all operating conditions is output, along with an optimized set of core parameters. And a manual for configuring closed-loop control link parameters.

[0038] Specifically, one embodiment of the present invention is applied to a low-grade waste heat recovery ORC power generation system in a steel plant. The core configuration and values ​​are as follows: the working fluid is R245fa, the evaporator inner diameter is 0.05m, the condenser is divided into a superheated zone, a condensing zone, and a subcooled zone, and the working fluid pump has a rated speed of 3000r / min and a rated flow rate of 5m³ / min. 3 / h, the storage tank volume is 0.5m³ 3 Control parameters set the observer bandwidth. M-ESO observation error convergence time ≤ 4.2s, FTRC control gain The finite-time convergence coefficients are 30, 25, 20, and 3, respectively. Integral separation threshold The IDRL network has 64 hidden layer neurons, an empirical buffer capacity of 10,000, and a batch sampling size of 64. The verification condition includes a steady-state waste heat input of 80kW, dynamic waste heat fluctuations of ±20%, a load change rate of 5% / s, and a simulation duration of 1000s. The final steady-state condition is... Overshoot = 2.8%, Settlement time = 4.5s, Dynamic operating condition Overshoot Adjusting time By adopting a full-process multi-field coupled dynamic model, the model-to-actual system adaptation error was reduced by 62%; the combination of a 3×3 invertible decoupling matrix and M-ESO improved the dynamic coupling suppression efficiency by 58% and the disturbance estimation response speed by 45%; the FTRC control framework reduced steady-state overshoot by 44% and shortened dynamic adjustment time by 35%; and IDRL parameter adaptive optimization reduced parameter adjustment response delay by 70% and improved control accuracy under complex operating conditions by 47.5%.

[0039] In summary, this invention effectively solves the problems of insufficient suppression of multi-parameter dynamic coupling, weak adaptability to complex disturbances, and difficulty in dynamically adjusting control parameters according to operating conditions in existing ORC control technologies. Through multi-field coupling quantization modeling, dynamic decoupling, and accurate disturbance estimation, it ensures the system's adaptability to complex operating conditions such as industrial waste heat fluctuations and sudden load changes. By leveraging finite-time control and parameter adaptive optimization, it achieves a synergistic improvement in control accuracy, dynamic response speed, and equipment safety. At the same time, it balances control performance and system energy consumption through an energy consumption penalty mechanism. Ultimately, it improves the energy conversion efficiency of the ORC system by 8% to 12% under all operating conditions and significantly enhances long-term operational stability, providing an efficient and reliable control solution for ORC system engineering applications in fields such as low-grade waste heat recovery.

[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning, characterized in that: Includes the following steps: S1. Based on the structural composition and core thermodynamic principles of the Organic Rankine Cycle (ORC) system, a full-process multi-field coupled dynamic model is constructed by integrating the phase change characteristics of heat exchangers, the dynamic correlation law of components, and the multi-field coupling effect. This model includes the evaporator, condenser, working fluid pump, expander, and liquid storage tank. S2. Based on the multi-field coupled dynamic model in step S1, by clarifying the mapping relationship between the core controlled parameters of the system and the corresponding control variables, the safety constraint boundaries of each parameter are defined, and a multivariable control matching system adapted to the ORC system is obtained. S3. Based on the multivariable control matching system in step S2, the static and dynamic coupling between channels is eliminated by constructing a 3×3 invertible decoupling matrix, and an independent extended state observer is configured for each channel to obtain the multivariable decoupled extended state observer M-ESO. S4. Based on the M-ESO disturbance estimation results in step S3, a highly robust FTRC control framework is formed by constructing a finite-time control law including a disturbance compensation term, introducing control quantity saturation constraints and an anti-integral saturation mechanism. S5. Based on the FTRC control framework of step S4, an improved deep reinforcement learning IDRL parameter adaptive optimization model is obtained by constructing a single Actor + dual Critic network architecture, setting a multi-objective reward function and a priority experience replay mechanism. S6. Based on the multi-field coupled dynamic model of step S1, the M-ESO of step S3, the FTRC control framework of step S4, and the IDRL parameter adaptive optimization model of step S5, the ORC system is verified through steady-state and dynamic operating condition simulations to achieve robust control of the ORC system under all operating conditions.

2. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 1, characterized in that: The model construction methods in step S1 are as follows: Evaporator modeling: Using the moving boundary method, the model is divided into subcooled, two-phase, and superheated zones. The model is constructed based on the principles of mass conservation, energy conservation, and heat transfer. The formula is as follows: ; ; ; ; in, The working fluid velocity inside the pipe; The coordinates are along the pipe length direction; For time; , The average density and average specific enthalpy of the working fluid in the corresponding phase region are, in order. This refers to the pressure inside the evaporator. The inner diameter of the heat exchanger; This refers to the tube wall temperature of the corresponding phase region; The working fluid temperature; The convective heat transfer coefficient between the heat source and the pipe wall; This is the cavitation coefficient; This is the average density of the working fluid in the two-phase region of the evaporator. , These are the densities of the saturated gas phase and liquid phase, respectively. , These are the enthalpy of the saturated gas phase and liquid phase, respectively. This represents the average specific enthalpy of the working fluid in the two-phase region of the evaporator. Condenser modeling: The condenser is divided into superheated, condensing, and subcooled zones using the controlled volume method. Axial heat conduction is ignored, and the inlet and outlet pressures are assumed to be equal. Based on the general phase region energy conservation and cooling water side energy conservation, the formulas are as follows: ; ; in, The internal energy of the working fluid; The volume of the phase region; , For the inlet and outlet mass flow rates of the phase zone; , The specific enthalpy of the working fluid at the inlet and outlet of the corresponding component; Heat dissipation in the phase region; For the condenser Heat exchange of cooling water in each phase region; This refers to the cooling water flow rate; The specific heat capacity of cooling water at constant pressure; , These are the inlet and outlet temperatures of the cooling water, respectively. Expander modeling: Treat the scroll expander as a scroll compressor operating in reverse and establish a steady-state output model; Working fluid pump modeling: A multi-stage centrifugal pump is used to establish a dynamic correlation model of flow rate-head-power under variable speed, with the following formula: Rated speed and head formula: ; Variable speed flow rate formula: ; Variable speed head formula: ; Export enthalpy formula: ; in, , , This refers to the pump characteristic coefficient; Rated head; Rated flow rate; This refers to the actual rotational speed; Rated speed; This represents the actual traffic volume. For actual head; The specific volume of the working fluid at the pump inlet; , The inlet and outlet pressures of the pump; For pump adiabatic efficiency; Modeling of the liquid storage tank: Following the laws of mass and energy conservation to achieve the working fluid buffering function, the formulas are as follows: mass conservation formula: ; Energy conservation formula: ; in, The mass of the working fluid in the storage tank; The total energy of the working fluid in the storage tank; It facilitates heat exchange between the storage tank and the environment.

3. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 2, characterized in that: In step S1, the coupling coefficient between working fluid pump speed and evaporation pressure is obtained by quantifying the dynamic relationship between temperature, pressure, and flow rate. Expansion valve opening degree-superheat coupling coefficient Cooling water flow rate - condensing pressure coupling coefficient .

4. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 3, characterized in that: The specific steps of step S2 include: S21. Based on the energy conversion efficiency and operational safety requirements of the ORC system, the core controlled parameters are quantitatively defined and the expected target values ​​for each controlled parameter are set. The core controlled parameters include superheat, the evaporation pressure of the evaporator in step S1, and the condensation pressure output by the condenser condensation zone model in step S1. The formula for calculating superheat is: ; in, The working fluid temperature at the evaporator outlet. Evaporation pressure The corresponding saturation temperature; This refers to the saturation pressure in the two-phase region of the evaporator. S22. By matching independent control quantities corresponding to the number of controlled parameters, establish a quantitative dynamic correlation between the control quantities and the controlled parameters. The control quantities include the working fluid pump speed. Expansion valve opening and cooling water flow rate The formula for the association relationship is: ; ; ; in, , , These are dimensionless coupling coefficients; , , The changes in evaporation pressure, superheat, and condensation pressure relative to their corresponding expected target values ​​are, in order. , , The changes in the working fluid pump speed, expansion valve opening, and cooling water flow rate are, in order. S23. Based on the equipment's safe operation requirements and the thermodynamic properties of the working fluid, define the value constraints of each core controlled parameter and its expected target value, as well as the safety constraint boundaries of the control quantities. These constraints include superheat greater than 0K, evaporation pressure less than the critical pressure of the working fluid, and condensation pressure greater than standard atmospheric pressure. Both the evaporation pressure and its expected target value are less than the critical pressure of the working fluid, while both the condensation pressure and its expected target value are greater than standard atmospheric pressure. Furthermore, the control quantities satisfy the formula: ; ; ; in, , The minimum and maximum safe operating speeds of the working fluid pump are, in order. , The minimum and maximum safe flow rates for cooling water are listed in order.

5. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 4, characterized in that: The specific steps of step S3 include: S31. Based on the multivariate quantization mapping relationship and constraint boundary from step S2, obtain the static gain by fitting the steady-state data of the multi-field coupled dynamic model from step S1, and construct a 3×3 invertible decoupling matrix, as shown in the formula: ; in, For the first The control quantity affects the first... The static gain of a controlled parameter; S32, Decoupling matrix based on step S31 Establish a mapping between virtual control variables and actual control variables, and introduce virtual control variables as follows: Establish and control the actual amount The mapping relationship is given by the formula: To obtain independent control inputs for each controlled channel, where , , These are virtual control quantities for the evaporation pressure, superheat, and condensation pressure channels, respectively. It is the inverse of the decoupling matrix; S33. Based on the independent channel partitioning in step S32, a multi-channel collaborative observation mechanism is constructed by configuring an extended state observer for each channel, and the state equation is established as follows: ; in, For the first The actual output of each controlled parameter; For total disturbance; Equivalent gain; The observer equation is designed, and the formula is as follows: ; in, , They are respectively , Observed values; For observation error, and ; , For observer gain; It is a saturation function; S34. Based on the observer equations in step S33, the observer gain is tuned using the pole placement method, with the following formula: ; ; in, For observer bandwidth; The poles of the observer's characteristic equation are placed at... This is to ensure that the observation error convergence time is ≤5s.

6. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 5, characterized in that: The specific steps of step S4 include: S41. Based on the actual output of the controlled parameters of each channel in step S33 Compared with observed values Combining the expected target value of the controlled parameter defined in step S21, the tracking error and error derivative are defined as follows: ; ; in, For tracking error; The desired target value of the controlled parameter; These are the observed values ​​of the controlled parameters; The differential of the tracking error of the controlled channel; The derivative of the desired target value of the controlled parameter; The derivative of the observed values ​​of the controlled parameter; Controlled channel; S42. Total disturbance observations based on the output of step S33 Combining finite-time stability theory, a control law with disturbance compensation is constructed to achieve rapid convergence of the controlled parameter. The formula is as follows: ; in, The total disturbance estimated by M-ESO; , , To control the gain; The convergence coefficients are finite-time convergence coefficients; It is a saturation function; It is a symbolic function; S43. Based on the control quantity safety constraint boundary defined in step S23, the virtual control quantity... The mapped actual control quantity is saturated and limited to ensure it does not exceed the physical limits of the equipment. The formula is: ; in, This is the actual control quantity; , To control the upper and lower limits of the quantity; It is a saturated function, and when Time output ,when Time output Otherwise, output the original value; S44. Integrate the optimization results of sub-steps S42-S43 to output the finite-time robust controller FTRC.

7. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 6, characterized in that: Step S44 also includes optimizing the control law by introducing an integral separation strategy to avoid integral saturation under long-term disturbances or large errors, as shown in the formula: ; ; in, The integral separation threshold; This is the integral gain; The optimized virtual control quantity inverse matrix of the decoupling matrix Convert to actual control quantity Synchronous output control gain ~ Convergence coefficient saturation threshold Integral separation threshold .

8. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 7, characterized in that: The specific steps of step S5 include: S51. Based on the FTRC control framework in step S44, construct a single Actor + dual Critic deep reinforcement learning network architecture, with the observer gain of step S3M-ESO and the core control parameters of FTRC in step S44 as the optimization objects. S52. Based on the control performance requirements of FTRC in step S4, the parameter constraint boundaries defined in step S2, and the energy consumption optimization objectives of the ORC system, the parameter optimization effect is quantitatively evaluated through a multi-objective reward function, as shown in the formula: ; in, To control accuracy rewards, and , Let be the absolute error of the integration time, and , For precision weighting coefficients; For dynamic response rewards, and , To adjust the time, , For overshoot, , , For response weighting coefficients; The reward is a parameter constraint, and the optimized parameters simultaneously satisfy... , , , hour, ,otherwise ; It is an energy consumption penalty item, and ; , , , For normalized weights; S53. Set a priority experience replay buffer to store experience samples generated by the interaction between the IDRL model and the ORC training environment built by the multi-field coupled dynamic model in step S1. The formula is: ; ; in, for Sample priority at any given time; This refers to timing difference error; This is the priority offset; Discount factor; For the output of the dual Critic network Mean value of state at any given time; for System status at all times; For the reward function calculated based on step S52 Momentary reward value, The system state at the next moment after the parameters are applied is the output of the model in step S1. S54. Based on the training environment built by the multi-field coupled dynamic model of the ORC system in step S1, after initializing the Actor and dual-Critic network parameters and experience buffer, the training iteratively follows the process of substituting the optimized parameters output by the Actor network into the M-ESO in step S3 and the FTRC in step S4, simulating the system response and calculating the reward using the dynamic model in step S1, storing experience samples and updating priorities, sampling and optimizing the dual-Critic network parameters, and updating the Actor network parameters using policy gradients, until the reward function is reached. Training is stopped when the fluctuation amplitude is ≤5% and convergence occurs in multiple consecutive iterations. S55. After training converges, output the improved deep reinforcement learning IDRL parameter adaptive optimization model, and simultaneously output the initial optimized core parameter set.

9. The organic Rankine cycle active disturbance rejection control method based on deep reinforcement learning according to claim 8, characterized in that: The specific steps of step S6 include: S61. Based on the multi-field coupling dynamic model of step S1, the M-ESO of step S3, the FTRC of step S4 and the IDRL parameter adaptive optimization model of step S5, complete the interface adaptation and signal linkage of each module. S62. Based on the actual industrial operation scenarios of the ORC system, and the parameter safety constraint boundary defined in step S2 and the model training operating condition range in step S5, two types of verification operating conditions covering the entire operating range are set up, including the steady-state operating condition with rated waste heat input and the dynamic operating condition with random fluctuations in industrial waste heat. S63. Input the operating parameters from step S62 into the closed-loop control link, start the simulation and run iteratively according to the established closed-loop process, and record the core data at each moment synchronously. S64. Based on the full data of the simulation run records, using the data from step S52... , , Quantitative indicators of control performance were used to conduct multi-dimensional performance evaluations of the three controlled channels under steady-state and dynamic conditions. S65. Summarize the multi-dimensional performance evaluation results from step S64 to form a full-condition verification report, clarifying the closed-loop control link under steady-state conditions. Overshoot Adjusting time Under dynamic operating conditions Overshoot Adjusting time Output a robust control scheme for the ORC system under all operating conditions, along with an optimized set of core parameters. .