A methanol reforming hydrogen production process control method and system based on an intelligent algorithm

By constructing an upstream-downstream coupled prediction model using a control method based on intelligent algorithms and utilizing a reinforcement learning agent to dynamically adjust the action space, the system instability problem caused by nonlinear correlation in the methanol reforming hydrogen production process was solved. This achieved efficient CO concentration control and energy consumption optimization, and improved the system stability and hydrogen recovery rate.

CN121149314BActive Publication Date: 2026-02-13JIANGSU IS ENERGY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511695607.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

In existing methanol reforming hydrogen production processes, there is a nonlinear relationship between upstream feed flow rate, water-to-carbon ratio, reaction temperature, bypass valve position and downstream purification pressure difference, which leads to increased CO concentration in the downstream purification section, decreased system stability, and traditional control methods are unable to cope with operating disturbances, resulting in problems such as regulation lag, amplified fluctuations, decreased hydrogen recovery rate and increased energy consumption.

Method used

A control method based on intelligent algorithms is adopted. By collecting real-time data from the upstream reforming unit and the downstream purification unit, a coupled prediction model is constructed, purification congestion indicators and their thresholds are defined, and the action space is dynamically adjusted using a reinforcement learning agent. Combined with sensitivity identification and minimal disturbance optimization models, the synergistic optimization of CO concentration and purification energy consumption is achieved, and an adaptive control strategy is constructed to quickly respond to changes in operating conditions.

Benefits of technology

It improves control precision and energy efficiency, avoids adjustment lag, enhances the robustness of process operation, shortens system stabilization time, improves control efficiency under load fluctuation scenarios, ensures CO concentration within a safe range, and increases hydrogen recovery rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121149314B_ABST
    Figure CN121149314B_ABST
Patent Text Reader

Abstract

The application provides a methanol reforming hydrogen production process control method and system based on an intelligent algorithm, relates to the technical field of fuel cell methanol hydrogen production process control, and comprises the following steps: collecting real-time operation data of an upstream reforming unit and a downstream purification unit, constructing an upstream and downstream coupling prediction model, and setting a purification congestion index and a threshold value; when the purification congestion index reaches the threshold value, triggering an enhanced learning intelligent agent to enter an active control state, dynamically adjusting an exploration rate and a utilization rate based on a reward function constraint of a CO concentration upper limit and a hydrogen recovery rate lower limit; collecting response data through disturbance and performing reverse sensitivity identification to construct a sensitivity and congestion coupling matrix; and adopting a strategy gradient and pruning joint search under the constraints of disturbance amplitude, energy consumption and safety. The application can realize the collaborative optimization of CO concentration stable control and purification energy consumption under complex working conditions close to the purification capacity limit, and improve the operation efficiency and stability of the hydrogen production process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of fuel cell methanol hydrogen production process control, and particularly relates to a methanol reforming hydrogen production process control method and system based on an intelligent algorithm. BACKGROUND

[0002] A fuel cell is a chemical device that directly converts chemical energy of fuel into electrical energy, and methanol reforming hydrogen production is a process of producing hydrogen by reforming reaction of methanol and water as raw materials under the action of a catalyst. The process has the advantages of compact equipment, high hydrogen purity, moderate reaction temperature, etc. Methanol reforming hydrogen production is widely used in the field of fuel cells, and its process generally includes an upstream methanol reforming unit and a downstream purification unit. The upstream reaction section is responsible for reforming methanol and water to generate hydrogen-containing mixed gas, and the downstream purification section removes impurities such as CO and CO2 from the gas to obtain high-purity hydrogen.

[0003] However, the existing methanol reforming hydrogen production has the following shortcomings:

[0004] Firstly, there is a nonlinear correlation between the upstream feed flow, the water-carbon ratio, the reaction temperature, the bypass valve position, and the downstream purification pressure difference, temperature, and beat; when the upstream load or reaction state is disturbed, it will directly lead to the downstream purification section approaching the capacity limit, the CO concentration rising or even exceeding the standard, and the system stability decreasing;

[0005] Secondly, the existing traditional PID control or simple logic threshold control method cannot cope with the nonlinear response under the working condition disturbance, and often appears the problems of regulation lag, fluctuation amplification, hydrogen recovery rate reduction, and energy consumption increase; therefore, we propose a methanol reforming hydrogen production process control method and system based on an intelligent algorithm.

[0006] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0007] The purpose of the present application is to overcome the shortcomings of the prior art, and to provide a methanol reforming hydrogen production process control method and system based on an intelligent algorithm, which solves the technical problems mentioned in the background.

[0008] To achieve the above-mentioned purpose, the present application provides the following technical scheme:

[0009] A methanol reforming hydrogen production process control method based on an intelligent algorithm, comprising the following steps:

[0010] S1. Collect real-time operating data of the upstream reforming unit and downstream purification unit of the methanol reforming system, construct an upstream and downstream coupled prediction model based on the operating data, and define purification congestion index and its threshold to characterize the margin status of the purification section.

[0011] S2. When the purification congestion index is less than the threshold, maintain normal process control; when the purification congestion index reaches or exceeds the threshold, trigger the reinforcement learning agent to enter the active control state, and dynamically adjust the exploration rate and utilization rate of the action space based on the hard constraints of the reward function on the upper limit of CO concentration and the lower limit of hydrogen recovery rate.

[0012] S3. In the balancing mode, apply a small disturbance to the upstream control quantity and collect the system response. Based on the disturbance and response data, perform reverse sensitivity identification to form a sensitivity vector. Adjust the action selection weight through sensitivity and congestion coupling matrix to prioritize the control of sensitive variables affecting CO concentration.

[0013] S4. Based on the sensitivity identification results and energy consumption target, construct a minimum perturbation balancing optimization model. Under the conditions of perturbation amplitude limit, hydrogen recovery rate lower limit, purification control limit, change rate limit and system temperature constraint, solve the optimal balancing action sequence through strategy gradient and pruning search iteratively, update the strategy network to achieve synergistic optimization of CO concentration and purification energy consumption.

[0014] S5. Based on the balance instruction sequence output by the strategy, and combined with sensitivity and action cost priority, the balance actions are executed in sequence. The residual deviation between CO concentration and hydrogen recovery rate is calculated based on the rolling prediction window. If the deviation exceeds the threshold, the synchronous correction of the strategy network and the coupled model is triggered.

[0015] S6. When the residual and sensitivity convergence criteria meet the set conditions, freeze the strategy parameters, sensitivity mapping and operating condition threshold, generate a stable strategy snapshot and store it in the experience pool; when similar conditions occur again, quickly load the historical snapshot through similarity retrieval to realize self-evolutionary control and rapid response from cold start to hot recovery.

[0016] S1 specifically includes: collecting real-time operating data of feed flow rate, water-to-carbon ratio, section temperature, bypass valve position, oxygen supply, purification pressure difference, purification temperature, and cycle time of the upstream reforming unit and the downstream purification unit; constructing an upstream and downstream coupled prediction model based on the operating data, and setting purification congestion indicators and their thresholds to characterize the capacity margin state of the purification section; initializing the state space and action space of the reinforcement learning agent, associating the agent with the coupled prediction model, and establishing an initial experience pool through small perturbation sampling to provide basic data for subsequent control strategy optimization.

[0017] S2 specifically includes: calculating the purification congestion index in real time and comparing it with a preset threshold; when the purification congestion index is lower than the threshold, maintaining process control based on conventional adjustment logic; when the purification congestion index reaches or exceeds the threshold, triggering the reinforcement learning agent to enter the active control state; the reinforcement learning agent constructs a reward function in the action space based on the upper limit of CO concentration and the lower limit of hydrogen recovery rate, and adjusts the action selection process with the reward function as a constraint; by dynamically adjusting the action exploration rate and utilization rate, the rapid triggering and stable control of the balancing mode are achieved, so that the system maintains a safety margin and hydrogen production efficiency when approaching the upper limit of purification capacity.

[0018] S3 specifically includes: after triggering the balancing mode, applying a small disturbance to the upstream control quantity and collecting real-time response data of CO outlet concentration and hydrogen recovery rate; performing reverse sensitivity identification based on the disturbance and response data to obtain a sensitivity vector; coupling the sensitivity vector with the purification congestion index to construct a sensitivity-congestion coupling matrix, and using this matrix as a basis to perform weighted adjustment of the action space, so that the control actions corresponding to high-sensitivity variables have higher selection priority, thereby achieving efficient regulation of CO concentration changes.

[0019] S4 specifically includes: constructing a minimum perturbation balancing optimization model based on sensitivity identification results and energy consumption targets, and setting constraints such as perturbation amplitude limits, hydrogen recovery rate lower limits, purification control limits, action change rates, and system temperature ranges; using a joint search method of policy gradient and pruning to iteratively optimize the balancing action sequence, by eliminating actions that contribute less to CO concentration improvement and rearranging action priorities, thereby reducing the search space and improving convergence efficiency; updating the policy network, and outputting the optimal balancing control command based on convergence criteria to achieve synergistic optimization of CO concentration suppression and purification energy consumption.

[0020] S5 specifically includes: after executing the balancing control action, based on the set rolling prediction window, collecting measured data of CO outlet concentration, hydrogen recovery rate and purification congestion index, and comparing them with the output of the coupled prediction model to calculate the residual deviation; when the residual deviation exceeds the threshold, triggering the policy correction of the reinforcement learning agent and the parameter correction of the coupled prediction model to realize the synchronous update of the control policy and the prediction model; when the residual meets the convergence condition, maintaining the execution of the existing policy, thereby forming a closed-loop adaptive control mechanism of rolling prediction, residual correction and policy update.

[0021] S6 specifically includes: when the residual and sensitivity convergence criteria meet the set conditions, triggering the control strategy's exit mechanism, freezing the policy parameters, sensitivity mapping, and operating condition thresholds of the reinforcement learning agent, generating a stable policy snapshot and binding it with operating condition feature information; when subsequent changes in operating conditions cause deviations to exceed limits, the system starts a relearning process, dynamically adjusts the threshold and reward weights, and re-triggers sensitivity identification and policy optimization; when the similarity to historical operating conditions is detected to reach a preset threshold, loading the corresponding stable policy snapshot to achieve rapid hot start and self-evolutionary recovery of the control strategy.

[0022] A control system for methanol reforming to hydrogen production based on intelligent algorithms, comprising:

[0023] The data acquisition module is used to collect real-time operating data of feed flow rate, water-to-carbon ratio, section temperature, bypass valve position, oxygen supply, purification pressure difference, purification temperature and circulation cycle of the upstream reforming unit and the downstream purification unit.

[0024] The threshold determination and triggering module is used to determine whether to trigger the reinforcement learning agent based on the comparison result between the congestion cleanup index and the preset threshold.

[0025] The sensitivity identification and weight adjustment module is used to apply a small disturbance to the upstream control quantity under active control, collect the CO outlet concentration and hydrogen recovery rate response, and perform reverse sensitivity identification.

[0026] The minimal perturbation optimization module is used to construct a minimal perturbation balancing optimization model based on sensitivity identification results and energy consumption targets under the constraints of perturbation amplitude, lower limit of hydrogen recovery rate, purification control limit, rate of change and system temperature.

[0027] The rolling correction module is used to collect the measured values ​​of CO outlet concentration and hydrogen recovery rate based on the rolling prediction window and compare them with the output of the coupled prediction model to calculate the residual deviation.

[0028] The snapshot and warm-start module is used to freeze the strategy parameters and sensitivity mapping when the residual and sensitivity convergence criteria meet the set conditions, generate a stable strategy snapshot and bind it to the operating condition characteristics.

[0029] The beneficial effects of this invention are as follows:

[0030] This invention collects multi-source operational data from upstream and downstream sources, constructs an upstream-downstream coupled prediction model, and introduces a purification congestion index. This model can reflect the state of the purification section's capacity approaching its limit in real time, identify potential fluctuation risks in advance, and has higher sensitivity and response speed compared to static threshold control. Through the purification congestion index and threshold judgment mechanism, the control system can intelligently switch between conventional adjustment and reinforcement learning control, avoiding the adjustment lag that traditional methods exhibit under large disturbance conditions, and enhancing the robustness of process operation. Through small disturbance triggering-response acquisition-reverse sensitivity identification and sensitivity-congestion coupling matrix construction, the selection of control actions has a quantitative basis, prioritizing the adjustment of variables sensitive to CO concentration, significantly improving control accuracy and energy efficiency.

[0031] This invention combines constraints such as disturbance amplitude, hydrogen recovery rate lower limit, purification control limit, action rate, and system temperature, employing a joint search method of policy gradient and pruning to improve the search efficiency and convergence speed of control actions, reduce energy consumption, and maintain CO concentration within a safe range. By calculating residual deviations through a rolling prediction window and simultaneously correcting the policy network and prediction model, continuous adaptive optimization of the control process is achieved, avoiding policy drift and improving the stability of the control system under load fluctuation scenarios. Once the system meets the convergence criteria, a stable policy snapshot bound to the operating condition characteristics is generated. When similar operating conditions occur, the control policy can be quickly loaded and restored, achieving a rapid "cold start → hot start" response, significantly shortening the system stabilization time and improving control efficiency. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of a methanol reforming hydrogen production process control method based on intelligent algorithms according to the present invention.

[0033] Figure 2 This is a schematic diagram of the control system framework for a methanol reforming hydrogen production process based on intelligent algorithms according to the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Example 1: As Figure 1 As shown, this embodiment provides a control method for methanol reforming hydrogen production process based on intelligent algorithms, including the following steps:

[0036] S1. Collect real-time operating data of the upstream reforming unit and downstream purification unit of the methanol reforming system, construct an upstream and downstream coupled prediction model based on the operating data, and define purification congestion index and its threshold to characterize the margin status of the purification section.

[0037] S2. When the purification congestion index is less than the threshold, maintain normal process control; when the purification congestion index reaches or exceeds the threshold, trigger the reinforcement learning agent to enter the active control state, and dynamically adjust the exploration rate and utilization rate of the action space based on the hard constraints of the reward function on the upper limit of CO concentration and the lower limit of hydrogen recovery rate.

[0038] S3. In the balancing mode, apply a small disturbance to the upstream control quantity and collect the system response. Based on the disturbance and response data, perform reverse sensitivity identification to form a sensitivity vector. Adjust the action selection weight through sensitivity and congestion coupling matrix to prioritize the control of sensitive variables affecting CO concentration.

[0039] S4. Based on the sensitivity identification results and energy consumption target, construct a minimum perturbation balancing optimization model. Under the conditions of perturbation amplitude limit, hydrogen recovery rate lower limit, purification control limit, change rate limit and system temperature constraint, solve the optimal balancing action sequence through strategy gradient and pruning search iteratively, update the strategy network to achieve synergistic optimization of CO concentration and purification energy consumption.

[0040] S5. Based on the balance instruction sequence output by the strategy, and combined with sensitivity and action cost priority, the balance actions are executed in sequence. The residual deviation between CO concentration and hydrogen recovery rate is calculated based on the rolling prediction window. If the deviation exceeds the threshold, the synchronous correction of the strategy network and the coupled model is triggered.

[0041] S6. When the residual and sensitivity convergence criteria meet the set conditions, freeze the strategy parameters, sensitivity mapping and operating condition threshold, generate a stable strategy snapshot and store it in the experience pool; when similar conditions occur again, quickly load the historical snapshot through similarity retrieval to realize self-evolutionary control and rapid response from cold start to hot recovery.

[0042] S1 specifically includes the following sub-steps:

[0043] S110. Collect real-time operating data of the upstream reforming unit and the downstream purification unit. The upstream operating data includes feed flow rate, water-to-carbon ratio (S / C), temperature of each reaction zone, bypass valve position, and micro-oxygen supply. The downstream operating data includes purification section pressure difference, purification temperature, and purification cycle time. The data acquisition frequency is set to 1Hz to 5Hz and can be dynamically adjusted according to the fluctuation characteristics of the operating conditions. The acquisition period covers stable operating conditions and load disturbance operating conditions.

[0044] To ensure data acquisition accuracy, the sampling period is matched with the thermal inertia time constant of the reaction section. When the thermal inertia time constant of the reaction section is... At that time, the sampling interval Δt is controlled within The range ensures that the dominant frequency characteristics of changes in CO concentration and hydrogen recovery rate can be captured.

[0045] S120. Based on the collected upstream and downstream operating data, an integrated upstream-downstream coupled prediction model is constructed, mapping the upstream and downstream control variables to the CO outlet concentration and hydrogen recovery rate outputs. The coupled prediction model adopts a multi-input multi-output nonlinear mapping structure and supports real-time online correction, ensuring a prediction accuracy of no less than 95%.

[0046] The structure of the coupled prediction model can adopt a feedforward neural network or a Gaussian process regression structure, and the input vector is denoted as:

[0047]

[0048] in It is the feed flow rate of the upstream methanol reforming unit, which directly determines the reaction load and gas production rate; S / C is the water-to-carbon ratio, that is, the feed ratio of water vapor to methanol, which is an important parameter for controlling hydrogen production and CO content. , , The temperature of each reaction section (e.g., preheating zone, reaction zone, outlet section) is a key variable controlling the reforming reaction rate and equilibrium state. The opening degree of the bypass valve determines the gas diversion ratio, which is used to achieve dynamic balance of upstream and downstream flow. It refers to the amount of oxygen supplementation, used for micro-oxygen supply regulation, which helps maintain the reaction heat balance and match the downstream purification load.

[0049] The output vector is denoted as

[0050]

[0051] in This indicates the CO concentration at the outlet of the purification section, used to reflect the system's CO slippage risk and purification efficiency; The hydrogen recovery rate is used to measure hydrogen production efficiency and energy consumption. The control system predicts the changing trend of the output target vector Y by real-time sensing and modeling of the input state vector X, and drives the reinforcement learning control strategy accordingly to achieve the synergistic goal of stable CO concentration control and maximizing hydrogen recovery rate. The weights are adjusted online through a recursive correction algorithm to adapt to dynamic load changes.

[0052] S130. Based on the operating status of the purification section, define a purification congestion index h to characterize the margin of the purification section from its capacity limit; simultaneously, set a purification congestion threshold. CO export concentration limit and lower limit of hydrogen recovery rate This is used to trigger subsequent balancing control and reinforcement learning policy constraints;

[0053] The purification congestion index is determined by the pressure differential margin of the purification section. Purification temperature margin margin factor with CO concentration Weighted calculation yields:

[0054]

[0055] in , The value of h is a weighting factor; the value of h ranges from 0 to 1, and the larger the value, the closer the purification section is to the capacity limit.

[0056] S140. Initialize the state space and action space of the reinforcement learning agent. The state space includes the feed flow rate F. in Water-to-carbon ratio (S / C), section temperature (T1~T3), bypass valve position Micro-oxygen supplementation The purification section pressure difference and the purification congestion index h are a total of 7 continuous variables;

[0057] The action space contains a set of perturbation actions corresponding to the upstream control variables, and defines the perturbation step size range; wherein the perturbation amplitude Process stability tolerances of each variable hook up:

[0058]

[0059] This represents the magnitude of the disturbance of the i-th control variable; This represents the stability tolerance or standard deviation of the i-th control variable, used to measure the allowable fluctuation range of the system; to avoid CO concentration spikes caused by disturbances; the initial value of the reward function of the reinforcement learning agent is set to 0, and the reward function weights are... , The initial values ​​were set in the range of 0.4-0.6 to facilitate balanced constraints on CO concentration and hydrogen recovery rate during the initial training phase.

[0060] S150. Associate the reinforcement learning agent with the coupled prediction model, and in the initial stage, form an initial experience pool by randomly sampling the action space with small perturbations. This is used for strategy initialization in subsequent reverse sensitivity balancing and action selection.

[0061] The initial sampling strategy combines Latin hypercube (LHS) sampling with uniform perturbation to ensure that the empirical pool covers at least 80% of the 7-dimensional state space; each sample stored in the empirical pool contains a state vector. Action vectors ,award and system response This provides a sufficient training foundation for subsequent reinforcement learning convergence.

[0062] S2 specifically includes the following sub-steps:

[0063] S210. Calculate the purification congestion index h in real time (formula shown in S130). The purification congestion index is determined by the pressure difference margin of the purification section. Purification temperature margin margin factor with CO concentration The weighted calculation was used to obtain the margins, where each margin was acquired through real-time sensor data and compared with the process baseline value.

[0064] The margins are defined as follows:

[0065]

[0066]

[0067]

[0068] in This is the upper limit of the differential pressure of the purification system, representing the maximum allowable differential pressure for the system to operate safely. This is the current real-time pressure difference; This is the base pressure, which is usually taken as the stable operating value under normal operating conditions. , , These represent the upper limit, current value, and baseline value of the purification section temperature, respectively. , , These represent the upper limit, current value, and baseline value of CO concentration, respectively. , , These represent the margin factors for pressure difference, temperature, and CO concentration, respectively; the value of h ranges from 0 to 1, and when h is close to 1, it indicates that the purification section is approaching its capacity limit. The system sampling period is matched with the reaction characteristic frequency to ensure real-time performance of no less than 1Hz.

[0069] S220, Combine the congestion index h with the preset congestion threshold. When comparing, At this time, the control system maintains the conventional control mode and relies solely on conventional process control logic for upstream and downstream coordination; the reinforcement learning agent maintains a "passive monitoring state" at this time, only updating the state space without issuing action commands.

[0070] in The typical setting range is 0.6-0.8, which can be adjusted online according to process load characteristics. To ensure control robustness, dynamic smoothing will be applied when the h fluctuation frequency exceeds 0.2Hz.

[0071]

[0072] in This represents the filtered signal value; This represents the raw signal value collected at the current moment (e.g., CO concentration, pressure, or temperature). This represents the filtered signal value at the previous moment; This represents the filter weight coefficient (range 0-1). The larger the value, the stronger the influence of the current signal and the more sensitive the response; The smaller the value, the stronger the smoothing effect, but the slower the response speed.

[0073] S230, When cleaning up congestion indicators At this time, the control system enters the balancing mode; the reinforcement learning agent switches to the "active control state", the action space is unfrozen, and the triggering of perturbation actions is allowed; at this time, the upstream-downstream coupling model enters the constrained optimization operation stage.

[0074] To prevent frequent switching, when h is within the threshold range (in When the recommended value is 0.02-0.05, the control system maintains the current mode state, forming a hysteresis band to avoid oscillation.

[0075] S240. In the balancing mode, the CO outlet concentration is kept within the upper limit and the hydrogen recovery rate is kept above the lower limit as hard constraints on the reward function of the reinforcement learning agent.

[0076] The reward function is defined as:

[0077]

[0078] in This represents the real-time CO outlet concentration. This is the upper limit threshold for CO concentration. For hydrogen recovery rate, This is the lower limit threshold for hydrogen recovery rate. and These are the reward weighting coefficients, with an initial suggested value of 0.5 ± 0.1, which can be dynamically adjusted online.

[0079] like and If the deviation is positive, then R=0; if a deviation occurs, the reward value is negative, and the larger the deviation, the stronger the penalty; this hard constraint ensures that the strategy action will not exceed the process safety boundary.

[0080] S250: The reinforcement learning agent dynamically adjusts the exploration rate of its strategy based on feedback from the congestion metric h and the reward function. and utilization rate :

[0081]

[0082] in It is a dynamic exploration rate based on the congestion cleanup index; It is the maximum exploration rate (when the congestion is low, the system can withstand more action disturbances). It is the minimum exploration rate (when the cleanup congestion is high, the system should minimize action disturbances); h is the cleanup congestion index, the larger the value, the closer it is to the capacity limit.

[0083] Recommended value: 0.8-1.0 Recommended value: 0.05-0.2. When... (Approaching capacity limit) The exploration rate decreases and the utilization rate increases to strengthen conservative control actions; when h approaches the lower limit of the threshold, Adjust appropriately upwards to enhance the system's sensitivity to potential disturbances.

[0084] In addition, the agent introduces a "sliding time window" to determine the stability of the strategy, and the window length... Values ​​range from 20 to 60 seconds; when Var(R) is continuously below the preset fluctuation threshold... (For example When the strategy enters a stable range, the exploration rate is further reduced to suppress unnecessary disturbances, ensuring that the balancing mode converges quickly under high load conditions, while avoiding frequent disturbances that could cause system oscillations.

[0085] S3 specifically includes the following sub-steps:

[0086] S310. After entering the balancing mode, the reinforcement learning agent triggers a small perturbation according to the action space definition, and performs a small adjustment to one or more of the upstream control variables.

[0087] Let the upstream manipulation vector be:

[0088]

[0089] The disturbance vector is:

[0090]

[0091] in It is the perturbation value of the i-th control variable at a certain time step, each term The disturbance amplitude is limited to 3% to 5% of the corresponding variable's rated value to ensure that the disturbance is small enough not to disrupt the thermal and material balance of the methanol reforming-purification system.

[0092] The perturbation action is triggered using a hybrid approach combining an ε-greedy strategy and sensitivity-priority weighting:

[0093]

[0094] in A is the perturbation action selected at time t, and A is the action space set. It is the action-value function, representing the state. Next action The expected returns that can be obtained; It is the sensitivity coefficient corresponding to the controlled quantity, representing the strength of the variable's response to CO concentration and hydrogen recovery rate; This is a sensitivity weighting adjustment factor used to balance the influence of Q value and sensitivity priority.

[0095] During the perturbation process, the response signals of CO concentration and hydrogen recovery rate are synchronously collected by sensors at a sampling frequency of no less than 1 Hz, forming a time-series data pair of perturbation input and response output, and constructing a sensitivity identification dataset. .

[0096] This dataset is used to calculate the coupling relationship between perturbation input and output response, providing a foundation for subsequent sensitivity matrix updates, action weight adjustments, and control strategy optimization. Through the above design, the system can achieve rapid sensitivity identification without compromising operating stability, and perform efficient perturbation action selection based on the dual weights of Q value and sensitivity, providing support for the efficient convergence and policy evolution of reinforcement learning agents.

[0097] S320. Dynamically set the disturbance amplitude threshold based on the purification congestion index h. When h approaches the capacity limit, the disturbance amplitude automatically decreases to prevent triggering peak fluctuations in CO concentration; when h is in the low to medium range, the disturbance amplitude is appropriately widened to enhance identification resolution.

[0098] The disturbance amplitude threshold is adopted in the form of a piecewise function:

[0099]

[0100] in This serves as the baseline disturbance amplitude, used for standard disturbances under low-risk operating conditions. is the minimum permissible disturbance under high-risk operating conditions; h is the congestion mitigation index. To clean up congestion thresholds; This is the width of the buffer. This is a scaling factor used to control the slope of the disturbance amplitude decrease.

[0101] When the cleanup congestion index is below the threshold ( At that time, the system has a large margin, and the agent adopts... The amplitude of the disturbance is adjusted to improve exploration efficiency;

[0102] When the congestion indicator enters the transition zone When h increases, the disturbance amplitude decreases linearly, achieving a smooth transition of the disturbance strategy;

[0103] When the congestion metric exceeds the upper limit of the threshold When the system approaches its capacity limit, the disturbance amplitude automatically shrinks to... This is to ensure the safety and stability of the control process.

[0104] S330. Using the input perturbation-output response data pair, perform reverse sensitivity identification on the mapping relationship between CO outlet concentration and upstream control quantity.

[0105] Let the change in CO outlet concentration at time t be:

[0106]

[0107] in This represents the change in CO concentration at time t; This represents the CO concentration at the outlet of the purification section at time t; This indicates the CO concentration value at the outlet of the purification section at the previous moment.

[0108] The disturbance vector is:

[0109]

[0110] in This represents the change (disturbance vector) of the upstream control variable at time t. Represents the upstream manipulation vector at time t; This represents the upstream control vector at the previous moment.

[0111] Sensitivity vector The update uses the recursive least squares (RLS) method:

[0112]

[0113] Gain vector The calculation is as follows:

[0114]

[0115] covariance matrix The updated formula is:

[0116]

[0117] in represents the sensitivity estimation vector at time t, and represents the sensitivity coefficients of each upstream control quantity to the CO outlet concentration; This represents the RLS gain vector, which determines the extent to which this update affects the estimation results; This represents the covariance matrix, used to describe the uncertainty of parameter estimation; It is a forgetting factor (0.95-0.99) used to balance the weights of historical data and newly sampled data; the method can converge quickly with minimal disturbance amplitude, update the sensitivity mapping relationship in real time, and is suitable for highly dynamic operating conditions.

[0118] S340, The reinforcement learning agent uses the updated inverse sensitivity vector It serves as a key input to the policy network, and the action selection weights are dynamically adjusted accordingly.

[0119] Reweight the action probability distribution:

[0120]

[0121] in Indicates the state Next, select the i-th action. The probability of; This represents the sensitivity coefficient corresponding to the i-th control variable in the sensitivity vector at time t; This represents a temperature coefficient or adjustment factor used to control the degree of "discretion / smoothing" in motion distribution; The larger the value, the sharper the probability distribution, indicating a bias towards highly sensitive actions; The smaller the value, the more uniform the probability distribution, and the greater the exploratory potential.

[0122] denominator This represents the normalization factor, ensuring that the sum of the probabilities of all actions is 1; "7" corresponds to the dimension of the action space, which here equals 7 upstream control variables (such as feed flow rate, water-to-carbon ratio, temperature). (Bypass valve, oxygen supply); actions corresponding to high-sensitivity variables have a higher selection probability, while the action probability of low-sensitivity variables is compressed, thereby achieving more efficient control of CO concentration under the same perturbation cost.

[0123] S350. Jointly determine the sensitivity mapping and congestion mitigation indicators to form a sensitivity-congestion coupling matrix:

[0124]

[0125] in The congestion weighting function can be set as follows: ( (Recommended range: 0.2-0.5); the rows of matrix M correspond to upstream control variables, and the columns correspond to the operating conditions weights under different congestion levels.

[0126] As h increases, the matrix-driven policy network shrinks its action range, reduces the exploration rate, and decreases the perturbation amplitude, achieving continuous online tracking and stable response to the upstream-downstream coupling relationship. To ensure the convergence of the action policy, a sensitivity convergence criterion is set:

[0127]

[0128] in Recommended setting: 10 -3 ~10 -4 When the condition is met N times consecutively (N=5~10 is recommended), the sensitivity identification is considered stable, and the sensitivity weights of the policy network are frozen.

[0129] S4 specifically includes the following sub-steps:

[0130] S410. After completing the online identification of reverse sensitivity, construct the objective function for minimum perturbation balancing optimization. The objective is to minimize the weighted sum of CO deviation and purification section energy consumption while ensuring that the CO outlet concentration is below the upper limit threshold COmax.

[0131]

[0132] in Real-time CO outlet concentration; This is the upper limit threshold for CO concentration; Energy consumption per unit time in the purification section; , These are the weighting coefficients, This is a reference value for energy consumption (e.g., rated operating energy consumption).

[0133] To enhance numerical stability, the objective function is standardized: ,in and These are the minimum and maximum target values ​​observed within the rolling time window.

[0134] S420. Set optimization constraints to ensure that the balancing action is completed within the process safety boundary:

[0135]

[0136] in This indicates the amplitude of the upstream control variable disturbance. This is the upper limit of the disturbance amplitude, used to avoid sudden increases in CO concentration or fluctuations in reactor pressure caused by large movements; For hydrogen recovery rate, This is the lower limit for hydrogen recovery rate, ensuring that hydrogen production efficiency meets process requirements. This indicates the permissible disturbance level of the purification system. To ensure a safe disturbance threshold and prevent penetration or valve impact; Indicates the rate of change of the manipulated variable. To maximize the rate of change and prevent excessively rapid movements that could cause control oscillations; For the predicted time The system temperature, This is within the allowable temperature range of the reaction system to ensure catalyst activity and reaction selectivity.

[0137] The above constraint is represented as matrix: Au≤b; where u is the action vector, A is the constraint coefficient matrix, and b is the corresponding process limit threshold vector, ensuring that the optimization action does not cause process thermal shock or purification instability.

[0138] The above constraints are equivalent to drawing boundaries for the intelligent control algorithm, ensuring that the methanol reforming hydrogen production system always meets the requirements of safety, energy efficiency and stability when performing dynamic optimization and action exploration.

[0139] S430, Reinforcement Learning Agent Based on Objective Function With constraint matrix Guided by inverse sensitivity mapping, the upstream balancing action sequence is analyzed. Perform iterative search;

[0140] The search strategy employs a combined approach of policy gradient and greedy pruning: Policy gradient update:

[0141]

[0142] in It is the objective function For strategy parameters The gradient; Indicates the state Next, strategy Select Action The probability distribution; determined by parameters Decide; Taking the logarithm makes it easier to calculate the gradient, resulting in a more stable gradient update. Indicates the execution of an action The actual reward value after the transaction. This represents the baseline value, used to reduce gradient variance and preferentially select the average reward. Indicates the current strategy The expectation of the state-action distribution.

[0143] Greedy pruning: After each iteration, remove branches that contribute less than a threshold to improving CO bias. Actions are implemented to narrow the search space and accelerate convergence; action priorities are reordered: combined with the sensitivity-congestion matrix M, actions are sorted according to their contribution.

[0144]

[0145] in Indicates the i-th action Priority value; Indicates action The absolute value of the sensitivity to CO concentration reflects the degree of impact of the action on the system response; h is the purification congestion index, reflecting the load status of the purification system. It is an adjustment factor used to control the weight of congestion indicators in priority calculation; it prioritizes actions with high contribution and compresses the probability of actions with low impact.

[0146] S440, The reinforcement learning agent uses an Actor-Critic structure for policy updates: the Actor network outputs the action distribution; the Critic network evaluates the CO concentration deviation and purification energy consumption changes caused by the balancing actions.

[0147] Update rules:

[0148]

[0149]

[0150] in Represents the parameters of the policy network (Actor); Parameters representing the value network (Critic); , The learning rates are for the Actor and the Critic, respectively. Indicates the state Next, strategy selection action The probability of; The advantage function represents the action. Compared to the performance of the benchmark strategy, ;in For actual returns, The state-value function predicted by Critic; Indicates the parameter Gradient calculation.

[0151] Convergence Criterion:

[0152]

[0153] in This represents the Actor network parameters after the k-th iteration; This represents the parameters of the Actor network after the (k+1)th iteration; This represents the policy convergence threshold (usually set to 10). -4 -10 -6 After N consecutive convergences (N=5~10 recommended), the policy network is considered to have converged and enters the action execution phase.

[0154] S450. Based on the strategy iteration results, output the final upstream balancing instruction sequence and downstream coordination constraints; simultaneously, establish an action priority table: The priority table can be dynamically adjusted during strategy iteration, and the arrows indicate that the action priorities are sorted from high to low.

[0155] Final action command It is given by the following formula:

[0156]

[0157] in represents the optimal action (optimal upstream control vector); u represents the candidate action; U represents the action feasible region, i.e., the set of actions after constraint pruning; represents the normalized objective function value (e.g., CO deviation + energy consumption); A represents the constraint matrix, used to characterize the system operating condition constraints (e.g., disturbance amplitude, rate, temperature, etc.); b represents the constraint vector, corresponding to the boundary conditions for safe or process operation; argmin represents the selection of parameters that make the objective function... Take the minimum value of u; ensure that the CO outlet concentration is stably suppressed, the purification energy consumption is minimized, and a stable balanced closed loop is formed with the downstream system.

[0158] S5 specifically includes the following sub-steps:

[0159] S510. Based on the upstream balancing instruction sequence output by the reinforcement learning agent, and combined with the sensitivity-congestion coupling matrix M, dynamically generate an execution priority table.

[0160] Priority sorting rules:

[0161]

[0162] in This represents the absolute value of the sensitivity of the i-th control variable to the CO outlet concentration; This indicates the overall cost of the control action, including energy consumption, disturbance risk, and response lag.

[0163] Cost function It can be set as:

[0164]

[0165] in This indicates the energy cost of performing an action (such as heating energy consumption, gas supply energy consumption, electrical power, etc.). This indicates the potential risks and costs that the action may cause (such as system instability, excessive fluctuations in CO concentration, etc.). (This represents the cost of disturbance intensity), used to measure the magnitude of the instantaneous disturbance that an action causes to the entire system, such as the valve adjustment range or the sudden change in feed flow rate; , , The weighted coefficients represent the recommended values. , , Actions with higher priority are executed first, thereby achieving optimal CO concentration suppression and purification stability control within a limited number of actions.

[0166] S520: Perform balancing actions in order of priority, and trigger a short-term rolling prediction window after each action is performed.

[0167] Let the length of the rolling prediction window be... The predicted step size is The number of samples contained in the window ,recommend , .

[0168] In each window, the system collects: ;in This represents the measurement output dataset, indicating the time window. The collection of all key output variables acquired internally. It represents the CO outlet concentration of the system at time t and is one of the main monitoring indicators of the methanol reforming reaction process. represents the hydrogen recovery rate at time t, and represents the purified hydrogen yield or recovery efficiency. This represents the congestion index of the purification system at time t, used to measure whether there are problems such as blockage or saturation in the flow state of the purification section. Indicates the sampling start time. This indicates the sampling window length and the sampling duration.

[0169] And the corresponding predicted values ​​are calculated using a coupled prediction model: ;in This represents the model's predicted output, indicating the predicted values ​​of future states such as CO outlet concentration, hydrogen recovery rate, and congestion indicators obtained through model calculations. The process prediction model function can be: a mechanism-based mathematical model; a data-driven machine learning model (such as neural networks, random forests, etc.); or a coupled model of the two. This represents the control action vector currently being executed, such as feed flow rate, S / C ratio, reaction temperature, bypass valve opening, oxygenation rate, etc., forming a rolling prediction-actual measurement comparison data.

[0170] S530, Calculate the residual vector within the rolling prediction window. :

[0171]

[0172] This represents the deviation between the actual output of the system and the predicted output of the model at time t; and the mean square error (MSE) within the window is calculated:

[0173]

[0174] Simultaneously calculate the residual deviation between CO concentration and hydrogen recovery rate:

[0175]

[0176]

[0177] The CO concentration residual represents the deviation between the CO outlet concentration predicted by the model at time t and the measured value. The residual represents the hydrogen recovery rate, which is the deviation between the hydrogen recovery rate predicted by the model at time t and the measured value. This indicates the measured CO outlet concentration, obtained from an online CO concentration sensor; This indicates the predicted CO export concentration, derived from the model described above. calculate; This indicates the measured hydrogen recovery rate, derived from the flow rate and purity measurements of the purification unit; This represents the predicted hydrogen recovery rate, obtained from the model prediction; a residual threshold is set. , recommend: ; When the MSE or any residual exceeds the threshold, the policy correction mechanism is triggered.

[0178] S540: The reinforcement learning agent uses the bias residual as a policy correction signal to synchronously update the policy network and the coupled prediction model.

[0179] Strategy Adjustment:

[0180]

[0181] in These are the current parameters of the Actor (policy network); It is the new value of the Actor parameter after the update; It is the policy learning rate, which controls the step size of each update; The logarithmic probability gradient term represents the current action. In state The log probability with respect to the policy parameters The gradient; It is the action distribution output by the policy network; It is a residual signal;

[0182] Prediction model correction: Update the prediction model parameters using recursive least squares or Bayesian methods. Correction:

[0183]

[0184]

[0185]

[0186] in This represents the old values ​​of the model parameters (the parameter vector from the previous iteration). This represents the new values ​​of the model parameters after the update; The forgetting factor is 0.95–0.99. For gain vector (Kalman gain or RLS gain); It is the covariance matrix; It is the input feature vector at the current moment (the various control variables in the methanol reforming hydrogen production process). This dual correction ensures that the system strategy and the prediction model converge synchronously, and there will be no control prediction drift.

[0187] S550, forming a rolling closed-loop correction mechanism. When both MSE and residual are below the threshold:

[0188]

[0189] The system maintains the current policy execution, and the scrolling window enters "stable tracking mode";

[0190] If the residuals continue to exceed the limit, the reinforcement learning agent automatically: increases the exploration rate. Recalculate action priorities; re-trigger sensitivity identification; update strategy and model parameters; and start the next round of rolling prediction window.

[0191] This forms a closed-loop feedback chain: disturbance action → response acquisition → prediction-actual measurement comparison → residual correction → strategy-model synchronous update, ensuring that the system can still achieve adaptive control convergence under load fluctuation conditions.

[0192] S6 specifically includes the following sub-steps:

[0193] S610. After the rolling correction phase is completed, monitor the CO outlet concentration in real time. With hydrogen recovery rate And compare it with a set threshold.

[0194] The criteria for exiting the program upon meeting the standards are as follows:

[0195]

[0196] in Threshold for residuals in the rolling prediction window (recommended 10) -3 -10 -2 ), The CO concentration deviation threshold, This is the threshold for hydrogen recovery rate deviation. This indicates the sensitivity convergence threshold; when all four conditions are met, the system determines that the current control strategy has converged and stabilized under this operating condition, triggering the "achievement exit signal".

[0197] S620: After the target exit signal is triggered, the control system automatically switches from "balancing mode" to "normal control mode".

[0198] At this point, perform the following operation: Freeze the policy network parameters of the reinforcement learning agent. , Freeze sensitivity mapping vector Freeze and clean up congestion threshold and reward weight ; This will include the current state space, action space, and trim control commands. Store as a set of stable policy snapshots:

[0199]

[0200] in It is the optimal state, representing the current environmental state of the system, such as the state vector composed of feed flow rate, reaction temperature, S / C ratio, and purification congestion index; It is the optimal action, indicating that in the state... The next strategy selects a combination of control actions, such as adjusting the feed flow rate and the S / C ratio; This represents the optimal policy parameters of the Actor network, which are the converged parameters under the current conditions and are used for action decision-making. The optimal parameters of the Critic network are used to evaluate the effectiveness of the strategy (such as residual reduction and energy efficiency improvement). This represents the optimal sensitivity vector, indicating the direction of the maximum sensitivity response of each control variable to the CO outlet concentration at this moment; This represents the congestion threshold, indicating the maximum level of congestion that the cleanup segment can tolerate. This represents the CO concentration control weight, used to emphasize the penalty for exceeding the CO limit in the reward function; This represents the weight of the hydrogen recovery rate, used to emphasize hydrogen production efficiency in the reward function.

[0201] This snapshot is consistent with current load characteristics. (Including temperature range, purification pressure range, CO threshold level, and hydrogen recovery rate range) are bound together to form an "operating condition-strategy" mapping.

[0202] S630. When any of the CO outlet concentration, hydrogen recovery rate, or residual does not meet the above-mentioned compliance conditions, the reinforcement learning agent enters the adaptive relearning mode.

[0203] Its core mechanisms include: dynamically adjusting the congestion threshold for purification.

[0204]

[0205] in This is the congestion threshold of the purification system before adjustment; This is the adjusted congestion threshold; It is the adjustment coefficient (update step size), which controls the magnitude of the threshold adjustment; It is the CO outlet concentration residual: It is the residual of hydrogen recovery rate.

[0206] Dynamically reset the reward function weights:

[0207]

[0208] in , Is the CO control weight before adjustment and Control weights; , These are the adjusted weights; It is a weight adjustment coefficient that controls the update magnitude; it reactivates the action exploration mechanism with local micro-perturbations; and it triggers a new round of sensitivity identification and strategy optimization. This ensures that the system still has self-healing and rapid reconvergence capabilities under complex operating conditions.

[0209] S640, the reinforcement learning agent cyclically executes the control-optimization-feedback closed-loop process from S310 to S550, achieving adaptive rolling iterative updates:

[0210] The cyclical mechanism of disturbance action → response acquisition → sensitivity identification → balancing optimization → rolling correction → exit judgment / relearning requires no manual intervention and can achieve self-evolving convergence control in scenarios such as load disturbance, temperature fluctuation, and purification bottleneck.

[0211] S650, freeze a snapshot of the stable policy each time the target is met and exits. and the corresponding operating condition characteristic parameters (Including: load level L; temperature range) CO threshold Lower limit of hydrogen recovery rate Clean up congestion indicators Stored in the experience pool middle;

[0212] When the system detects a similar operating condition again When this happens, a similarity search is triggered:

[0213]

[0214] in The current state feature vector typically contains key features of the reaction control process, such as temperature, feed rate, S / C ratio, CO concentration, hydrogen recovery rate, and congestion index. The target state or historical best snapshot state vector represents the feature distribution when the optimal control effect is achieved under a certain operating condition; The dot product (inner product) of two vectors measures the consistency of their directions. , They are and The modulus (Euclidean norm); when (recommend When the value is 0.8, the corresponding snapshot is loaded directly. This enables a rapid response from "cold start" to "hot recovery," greatly shortening the system convergence time.

[0215] Example 2: Figure 2 As shown, this embodiment provides a methanol reforming hydrogen production process control system based on intelligent algorithms, including:

[0216] The data acquisition module is used to collect real-time operating data such as feed flow rate, water-to-carbon ratio, section temperature, bypass valve position, oxygen supply, purification pressure difference, purification temperature and cycle time of the upstream reforming unit and the downstream purification unit, and to build an upstream and downstream coupled prediction model and set purification congestion index and its threshold.

[0217] The threshold determination and triggering module is used to determine whether to trigger the reinforcement learning agent based on the comparison result between the congestion purification index and the preset threshold. When the congestion purification index is lower than the threshold, the normal process control is maintained. When the index reaches or exceeds the threshold, the reinforcement learning agent is triggered to enter the active control state.

[0218] The sensitivity identification and weight adjustment module is used to apply a small disturbance to the upstream control quantity under active control, collect the CO outlet concentration and hydrogen recovery rate response, perform reverse sensitivity identification, construct a sensitivity-congestion coupling matrix, and adjust the action space according to the matrix.

[0219] The minimal perturbation optimization module is used to construct a minimal perturbation balancing optimization model based on sensitivity identification results and energy consumption targets under the constraints of perturbation amplitude, lower limit of hydrogen recovery rate, purification control limit, rate of change and system temperature. It adopts a joint search method of policy gradient and pruning, and realizes the convergence of control policy and output of optimal balancing command through the Actor-Critic structure.

[0220] The rolling correction module is used to collect the measured values ​​of CO outlet concentration and hydrogen recovery rate based on the rolling prediction window and compare them with the output of the coupled prediction model to calculate the residual deviation. When the deviation exceeds the threshold, the synchronous correction of the strategy network and the model is triggered. When the deviation meets the convergence condition, the current strategy is maintained to achieve closed-loop adaptive control.

[0221] The snapshot and warm-start module is used to freeze the strategy parameters and sensitivity mapping when the residual and sensitivity convergence criteria meet the set conditions, generate a stable strategy snapshot and bind it to the operating condition characteristics, and load the corresponding snapshot when a similar operating condition is detected, so as to realize the rapid warm-start and self-evolution recovery of the control strategy.

[0222] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0223] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0224] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for controlling a hydrogen production process from methanol reforming based on intelligent algorithms, characterized by, Comprise the following steps: S1, collecting real-time operation data of the upstream reforming unit and the downstream purification unit of the methanol reforming system, building an upstream and downstream coupling prediction model based on the operation data, defining a purification congestion index and its threshold to characterize the capacity margin state of the purification section; S2, when the purification congestion index is less than the threshold, maintain the conventional process control; when the purification congestion index reaches or exceeds the threshold, trigger the enhanced learning agent to enter the active control state, dynamically adjust the exploration rate and utilization rate of the action space based on the reward function hard constraint CO concentration upper limit and hydrogen recovery rate lower limit; S3, in the trim mode, a small disturbance is applied to the upstream control quantity and the system response is collected, the disturbance and response data are used for reverse sensitivity identification, a sensitivity vector is formed, and the action selection weight is adjusted through the sensitivity and congestion coupling matrix, and the sensitive variables affecting CO concentration are controlled preferentially; S4, based on the sensitivity identification results and the energy consumption target, a minimum disturbance trimming optimization model is built, and the optimal trimming action sequence is solved through policy gradient and pruning search iteration under the constraints of disturbance amplitude limit, hydrogen recovery rate lower limit, purification control limit, change rate limit and system temperature, and the strategy network is updated to realize the cooperative optimization of CO concentration and purification energy consumption; S5, according to the trimming instruction sequence output by the strategy, combining the sensitivity and action cost priority, the trimming actions are executed in turn, the residual deviation of CO concentration and hydrogen recovery rate is calculated based on the rolling prediction window, and if the deviation exceeds the threshold, the strategy network and the coupling model are triggered to synchronize correction; S1 specifically includes: collecting real-time operation data of upstream reforming unit and downstream purification unit, including feed flow, water-carbon ratio, section temperature, bypass valve position, oxygen supply, purification pressure difference, purification temperature and cycle beat; based on the operation data, an upstream and downstream coupling prediction model is built, and a purification congestion index and its threshold are set to characterize the capacity margin state of the purification section; the state space and action space of the enhanced learning agent are initialized, the agent is associated with the coupling prediction model, and the initial experience pool is established by small disturbance sampling to provide basic data for subsequent control strategy optimization; The structure of the coupled prediction model adopts a feedforward neural network or a Gaussian process regression structure, and the input vector is denoted as: ; is the feed flow rate of the upstream methanol reforming unit; S / C is the steam-to-carbon ratio, that is, the feed ratio of steam to methanol; , , is the temperature of each reaction section; is the bypass valve opening degree, which determines the gas split ratio; is the supplemental oxygen amount, which is used for trace oxygen supply adjustment; The output vector is denoted by: ; represents the CO concentration at the outlet of the purification section; represents the hydrogen recovery. The purification congestion index is obtained by weighting the following factors: purification pressure difference margin purification temperature margin and CO concentration margin ; wherein h is a weight factor; the greater the value of h, the closer the purification section is to the capacity limit; the margins are defined as follows: ; ; ; is the upper limit value of the differential pressure of the purification system; is the current real-time differential pressure; is the reference differential pressure; , , respectively represent the upper limit value, the current value and the reference value of the temperature of the purification section; , , respectively represent the upper limit value, the current value and the reference value of the CO concentration.

2. The method of claim 1, wherein the control method is based on an intelligent algorithm. It also includes S6, when the residual error and sensitivity convergence criterion meet the set conditions, freeze the strategy parameters, sensitivity mapping and working condition threshold, generate a stable strategy snapshot and store it in the experience pool; when similar conditions occur again, quickly load the historical snapshot through similarity retrieval to realize self-evolution control and fast response from cold start to hot recovery.

3. The method of claim 1, wherein the control method is based on an intelligent algorithm. S2 specifically includes: Real-time calculation of the purification congestion index and comparison with the preset threshold, when the purification congestion index is lower than the threshold, maintain the process control based on the conventional regulation logic; When the purification congestion index reaches or exceeds the threshold, trigger the enhanced learning agent to enter the active control state; The enhanced learning agent adjusts the action selection process according to the reward function constructed based on the CO concentration upper limit and the hydrogen recovery rate lower limit in the action space; By dynamically adjusting the action exploration rate and utilization rate, the rapid triggering and stable control of the trimming mode are realized, and the system still maintains a safe margin and hydrogen production efficiency when approaching the upper limit of purification capacity.

4. The method of claim 1, wherein the control method is based on an intelligent algorithm. S3 specifically includes: After triggering the trim mode, a small disturbance is applied to the upstream manipulated variable, and real-time response data of CO outlet concentration and hydrogen recovery rate are collected; Based on the disturbance and response data, reverse sensitivity identification is performed to obtain a sensitivity vector; The sensitivity vector is coupled with the purification congestion index to construct a sensitivity and congestion coupling matrix, and the action space is weighted and adjusted based on the matrix, so that the control action corresponding to the high sensitivity variable has a higher selection priority, realizing efficient regulation and control of CO concentration changes.

5. The method of claim 1, wherein the method is characterized by: S4 specifically includes: Based on the sensitivity identification result and the energy consumption target, a minimum disturbance trimming optimization model is constructed, and disturbance amplitude limit, hydrogen recovery rate lower limit, purification control limit, action change rate and system temperature interval constraints are set; A strategy gradient and pruning combined search method is used to iteratively optimize the trimming action sequence, eliminate actions with low contribution to CO concentration improvement, rearrange action priorities, reduce search space and improve convergence efficiency; The strategy network is updated, and the optimal trimming control instruction is output based on the convergence criterion, realizing the collaborative optimization of CO concentration suppression and purification energy consumption.

6. The method of claim 1, wherein the method is characterized by: S5 specifically includes: After executing the trimming control action, based on the set rolling prediction window, the measured data of CO outlet concentration, hydrogen recovery rate and purification congestion index are collected, and compared with the output of the coupled prediction model to calculate the residual deviation; When the residual deviation exceeds the threshold, the strategy correction of the reinforcement learning agent and the parameter correction of the coupled prediction model are triggered to realize the synchronous update of the control strategy and the prediction model; When the residual meets the convergence condition, the existing strategy is maintained, thereby forming a closed-loop adaptive control mechanism of rolling prediction, residual correction and strategy update.

7. The method of claim 2, wherein the control method is based on an intelligent algorithm. S6 specifically includes: When the residual and sensitivity convergence criterion meet the set conditions, the standard exit mechanism of the control strategy is triggered, the strategy parameters of the reinforcement learning agent, the sensitivity mapping and the working condition threshold are frozen, a stable strategy snapshot is generated and bound with the working condition feature information; When the subsequent working condition changes cause the deviation to exceed the limit, the system starts the relearning process, dynamically adjusts the threshold and reward weight, and re-triggers the sensitivity identification and strategy optimization; When the similarity to the historical working condition reaches the preset threshold, the corresponding stable strategy snapshot is loaded to realize the rapid hot start and self-evolution recovery of the control strategy.

8. A control system for hydrogen production from methanol reforming based on intelligent algorithm, using the control method for hydrogen production from methanol reforming based on intelligent algorithm according to any one of claims 1-7, characterized in that, It includes: A data acquisition module is used to collect real-time operating data of the upstream reforming unit and the downstream purification unit, including feed flow, water-carbon ratio, section temperature, bypass valve position, oxygen supply, purification pressure difference, purification temperature and cycle beat; A threshold determination and triggering module is used to determine whether to trigger the reinforcement learning agent based on the comparison result of the purification congestion index and the preset threshold; A sensitivity identification and weight adjustment module is used to apply a small disturbance to the upstream manipulated variable in the active control state, collect the response of CO outlet concentration and hydrogen recovery rate, and perform reverse sensitivity identification; A minimum disturbance optimization module is used to construct a minimum disturbance trimming optimization model based on the sensitivity identification result and the energy consumption target under the constraints of disturbance amplitude, hydrogen recovery rate lower limit, purification control limit, change rate and system temperature; A rolling correction module is configured to collect the measured values of the CO outlet concentration and the hydrogen recovery rate based on a rolling prediction window, compare the measured values with the output of the coupled prediction model, and calculate residual deviations; A snapshot and warm start module is configured to freeze the strategy parameters and the sensitivity mapping when the residual and the sensitivity convergence criterion meet the set conditions, generate a stable strategy snapshot, and bind the stable strategy snapshot with the working condition characteristics.

Citation Information

Patent Citations

  • Water electrolysis efficiency dynamic scheduling method and system based on reinforcement learning

    CN120485871A