Event-triggered safety reinforcement learning based optimization method for crude oil scheduling in oil refineries

CN122840544APending Publication Date: 2026-09-29TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611033622.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明的目的就是为了克服上述现有技术存在的缺陷而提供一种基于事件触发安全强化学习的炼油厂原油调度优化方法,以解决或部分解决传统固定时间切片调度方法在系统稳态时仍频繁下发决策导致无效计算和设备磨损、统强化学习依赖软性罚函数约束无法在指令下发前绝对保证零溢罐、防抽空与防混油等刚性工业安全底线、以及传统混合整数规划或鲁棒优化方法求解耗时长、难以适应船期波动和原油性质扰动的问题

Benefits of technology

(1)本发明通过构建基于事件驱动的半马尔可夫决策模型,仅在预设的事件触发时启动调度决策,解决了传统固定时间切片调度方法在系统稳态时仍频繁下发决策导致无效计算和设备磨损的问题,实现了决策频率与实际物流节奏同步,显著减少无效计算并提升实时响应能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840544A_ABST
    Figure CN122840544A_ABST
Patent Text Reader

Abstract

The present application relates to a refinery crude scheduling optimization method based on event-triggered safety reinforcement learning, comprising the following steps: collecting the operating states of crude oil tankers, storage tanks, feed tanks and atmospheric-vacuum distillation units, and constructing an event-driven semi-Markov decision model; when events such as the arrival of a tanker, the approaching of inventory boundaries, the completion of a transportation task or the early warning of a unit running out of feed are triggered, using a hierarchical reinforcement learning model to output candidate topology decisions and candidate flow decisions; then, through anti-mixing oil interlocking rules, action inertia retention bias and quadratic programming constraint correction models, executable scheduling decisions that meet the hard physical constraints of the industrial site are generated; finally, based on the operation cost feedback closed loop, online policy updating is driven. Compared with the prior art, the present application has the advantages of considering strict physical constraints and eliminating invalid execution exploration, significantly reducing the switching cost and decision efficiency in long process scheduling, and having the advantages of real-time and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of refinery production scheduling, and in particular to a refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning. Background Technology

[0002] Crude oil dispatching is at the forefront of refinery production, serving as a crucial link connecting the terminal and atmospheric and vacuum distillation units for continuous processing. This process is subject to multiple physical constraints, including shipping schedule fluctuations, tank capacity limits, and blending procedures, making it a typical problem of coupled continuous logistics and discrete decision-making. Traditional mixed-integer programming or robust optimization methods are time-consuming, and existing methods often employ fixed-time slices for control, failing to reflect the inherent "event-driven" nature of industrial dispatching. This leads to frequent calculation and switching commands being issued even in steady-state conditions, unnecessarily increasing computational energy consumption and valve wear. Current reinforcement learning methods lack absolute hard safety constraints. Relying on soft penalty functions to constrain out-of-bounds behavior, they can only suppress violations statistically during training, failing to ensure that rigid industrial safety baselines such as zero tank overflow, anti-vacuuming, and anti-blending are met before commands are issued. Therefore, a new crude oil dispatching optimization method is urgently needed, requiring both rapid dynamic adaptive response under complex disturbances and absolute assurance that control commands conform to on-site physical constraints, while significantly reducing unnecessary equipment wear and switching costs.

[0003] Chinese invention patent CN118966431A discloses a collaborative optimization method for refinery production and maintenance based on reinforcement learning. This invention can balance the conflict between production decisions and equipment maintenance decisions, thereby improving production and maintenance efficiency. However, it still suffers from problems such as: traditional fixed-time slice scheduling methods frequently issuing decisions during system steady state, leading to invalid calculations and equipment wear; reinforcement learning relies on soft penalty function constraints, which cannot absolutely guarantee zero overflow before issuing instructions; rigid industrial safety bottom lines such as preventing cavitation and oil mixing; and traditional mixed-integer programming or robust optimization methods are time-consuming to solve and difficult to adapt to shipping schedule fluctuations and crude oil property disturbances.

[0004] In summary, there is currently a lack of a refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning to solve or partially solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a refinery crude oil scheduling optimization method based on event-triggered safety reinforcement learning. This method aims to solve or partially solve the problems of traditional fixed-time slice scheduling methods, which still issue decisions frequently in the steady state of the system, resulting in invalid calculations and equipment wear; reinforcement learning, which relies on soft penalty function constraints, cannot absolutely guarantee zero tank overflow before the instruction is issued; rigid industrial safety bottom lines such as preventing evacuation and mixing of oil; and traditional mixed integer programming or robust optimization methods, which are time-consuming to solve and difficult to adapt to shipping schedule fluctuations and crude oil property disturbances.

[0006] The objective of this invention can be achieved through the following technical solutions: This invention provides a refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning, the method comprising: S1. Obtain the status information of crude oil scheduling in the refinery, construct an event-driven semi-Markov decision model, and dynamically determine the next decision interval of the system according to the preset triggering events; S2. Based on the decision interval, the current decision state is input into the hierarchical reinforcement learning model. The discrete action output layer of the hierarchical reinforcement learning model outputs candidate discrete topology decisions, and the continuous action output layer outputs candidate continuous flow decisions. Anti-mixing interlock rule screening and action inertia maintenance rule screening are performed on the candidate discrete topology decisions. S3. Based on the discrete topology decision after rule screening, construct a quadratic programming model with the objective of minimizing the deviation between the candidate continuous flow decision and the corrected flow decision, and with the upper bound of dynamic flow and the minimum feed requirement as constraints. Solve the model to obtain the corrected continuous flow decision and generate executable scheduling control instructions. S4. Based on the execution result of the executable scheduling control instruction, calculate the total operating cost, and use the total operating cost as a feedback signal to drive the hierarchical reinforcement learning model to perform online policy updates, thereby realizing refinery crude oil scheduling optimization.

[0007] As a preferred technical solution, the status information includes: the current scheduling time, the remaining oil volume, crude oil properties and estimated arrival time of each crude oil tanker to be unloaded, the real-time inventory, minimum safety stock, maximum safety stock and crude oil properties of each storage tank, the real-time inventory, maximum allowable inventory and crude oil properties of each feed tank, the feed demand of each atmospheric and vacuum distillation unit, and the pipeline start / stop status or topology connection status of the previous decision cycle.

[0008] As a preferred technical solution, the next decision interval Represented as: in, This refers to the time required for the next crude oil tanker to arrive at the port. For storage tanks or filling tanks to reach their safety threshold, This indicates that the transfer operation has been completed.

[0009] As a preferred technical solution, the discrete action output layer is equipped with a heuristic inertial masking mechanism, which modifies the logit distribution of the discrete strategy header before sampling, and applies a bias based on the action of the previous moment and the volume of the current active tank, as follows: in, Candidate discrete actions after inertial mask correction The corresponding logit value, This is the raw logits output. For candidate discrete actions, For constant bias, This is an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise. The action in the previous moment. For the current active tank volume, Minimum volume threshold, For safety margin.

[0010] As a preferred technical solution, the hierarchical reinforcement learning model constructs a dual security mechanism, including a security constraint layer and a quadratic differentiable programming layer.

[0011] As a preferred technical solution, the safety constraint layer includes hard interlock rules and inertial rules; the hard interlock rules are used to block transport paths that are not allowed to be used simultaneously; the inertial rules are used to prioritize keeping the previous topology unchanged when the current active tank has sufficient inventory.

[0012] As a preferred technical solution, the quadratic differentiable programmable layer projects the output of the original continuous actions to obtain the executable action that is closest to the original action and satisfies all physical constraints, expressed as: in, The original continuous target traffic output by the policy network. For discrete topological actions, This is an upper bound for dynamic traffic related to the current state and topology actions. The executable continuous action is obtained after projection through a second-order differentiable programmable layer. For the continuous flow decision variable to be optimized, For the index of the feed tank, For indexing crude oil distillation units, A collection of crude oil distillation units. This is a collection of feeding tanks.

[0013] As a preferred technical solution, the reward function of the hierarchical reinforcement learning model is expressed as the sum of the negative of the total operating cost and the reward shaping term.

[0014] As a preferred technical solution, the total operating cost is a weighted sum of all costs by unit weight, including inventory holding cost, unloading operation cost, vessel waiting cost, and mode switching cost.

[0015] As a preferred technical solution, the reward shaping item Used to guide exploration, represented as: in, For the already processed crude oil, the coefficient is... Adopting a course-based learning strategy, For training rounds.

[0016] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) This invention constructs an event-driven semi-Markov decision model, which initiates scheduling decisions only when a preset event is triggered. This solves the problem that the traditional fixed-time slice scheduling method still issues decisions frequently when the system is in steady state, leading to invalid calculations and equipment wear. It realizes that the decision frequency is synchronized with the actual logistics rhythm, significantly reduces invalid calculations and improves real-time response capabilities.

[0017] (2) By constructing a dual security mechanism, the synergistic effect of the logical security layer and the quadratic differentiable programming layer, this invention strictly satisfies the constraints of mass conservation, tank capacity boundary, material supply requirements and non-simultaneous material inflow and outflow during the strategy execution phase. This solves the problem that traditional reinforcement learning, which relies on soft penalty function constraints, cannot absolutely guarantee the rigid industrial safety bottom line such as zero tank overflow, anti-vacuuming and anti-mixing before the instruction is issued. It achieves the technical effect of ensuring that the scheduling instruction meets the logical security constraints and has physical feasibility.

[0018] (3) The present invention adopts a hierarchical reinforcement learning structure. The upper-layer policy network outputs discrete topology actions and continuous target flow parameters, while the lower layer achieves dynamic adaptive response through online policy updates. This solves the problems of long solution time and difficulty in adapting to shipping schedule fluctuations and crude oil property disturbances in traditional mixed integer programming or robust optimization methods. It achieves short single-step decision time and significantly improves scheduling efficiency under complex disturbances. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the event-triggered security reinforcement learning framework of the present invention; Figure 2This is a schematic diagram of the crude oil dispatching process in a refinery according to the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0021] To address the problems existing in the prior art, this embodiment provides a refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning, such as... Figure 1 As shown, the proposed event-triggered security reinforcement learning framework consists of two layers. The upper-layer reinforcement learning agent is responsible for making decisions on discrete operating modes and reference flows, and a heuristic logic security layer is embedded in this layer to suppress high-frequency mode switching. The lower execution layer then transforms these abstract decisions into strictly feasible continuous flows through differentiable quadratic programming (QP) projection. The method specifically includes: S1. Obtain the status information of crude oil scheduling in the refinery, construct an event-driven semi-Markov decision model, and dynamically determine the next decision interval of the system based on preset trigger events.

[0022] This invention represents crude oil scheduling objects as a directed network. , where the set of nodes Including ship collection Storage tank assembly , Mixing tank collection and crude oil distillation unit assembly edge set Indicates the permitted crude oil transport routes.

[0023] At any scheduling time The volume of the tank satisfies the following mass conservation law: in, For storage tanks volume, For the upstream node Transport to node Traffic, For the node Transported to downstream nodes Traffic, For continuous scheduling time variables, It is the set of upstream adjacent nodes. For the set of downstream adjacent nodes, For upstream node indexing, For downstream node indexing.

[0024] For crude oil distillation units, the continuous feed requirement is met: in, For feeding tank Transported to crude oil distillation unit crude oil flow crude oil distillation unit The minimum continuous feed requirement, Index for crude oil distillation units, A collection of crude oil distillation units. For the feeding tank index, This is a collection of feeding tanks.

[0025] Status information includes: the current scheduling time, the remaining oil volume, crude oil properties and estimated arrival time of each crude oil tanker to be unloaded, the real-time inventory, minimum safety stock, maximum safety stock and crude oil properties of each storage tank, the real-time inventory, maximum allowable inventory and crude oil properties of each feed tank, the feed demand of each atmospheric and vacuum distillation unit, and the pipeline start / stop status or topology connection status of the previous decision cycle.

[0026] Triggering events include: 1. When a ship arrives at the port, berth allocation or unloading allocation needs to be carried out; 2. If the liquid level in the storage tank or mixing tank is close to the upper or lower limit, the transport route needs to be switched in time; 3. Once the oil unloading or transfer operation is completed, equipment resources need to be reallocated; 4. If the liquid level in the blending tank is too low, it may cause a feed interruption in the crude oil distillation unit.

[0027] state space Represented as: in, For the current time, For detailed information on all vessels, including volume, concentration, and arrival time, and The status of storage tanks and blending tanks, including liquid level and concentration. The value of the discrete variable is the value at the previous time step.

[0028] Will Inclusion of status can provide the necessary context for high-frequency switching penalties, thereby helping to maintain stable refinery operations.

[0029] Action space is defined as a hybrid space combining discrete logic and continuous parameters, denoted as .

[0030] Discrete subspace Topological connectivity control is represented by logic gates: in, Used to determine the connection between the tanker and the storage tank. Used to control the feed lines from the blending tank to the CDU. These binary variables are equivalent to logic switches, used to enable or disable specific flow paths.

[0031] Continuous subspace Indicates the target flow parameters: Each element Corresponding edge in the network graph The target traffic.

[0032] The goal is to maximize throughput while minimizing operating costs. To facilitate learning in complex environments, rewards will be... Expressed as the sum of the negative of total operating costs and incentive shaping items: in, Total operating cost, To reward body shaping.

[0033] Total operating costs Costs are categorized into four types based on their unit weight. Weighted composition: in, , , , These represent inventory holding costs, unloading operation costs, vessel waiting costs, and mode switching costs, respectively. , , , These represent the coefficients for inventory holdings, unloading operations, vessel waiting (demurrage), and mode switching penalties, respectively. Demurrage is weighted to prevent port operations from becoming deadlocked. It is assigned a large value to impose a strong penalty on delayed vessels.

[0034] To encourage exploration, further shaping rewards will be introduced. : in, This refers to the processing of crude oil. (Coefficient) A learning strategy is employed instead of constant weights: initial weights are higher to encourage early learning and improve throughput; subsequent weights are adjusted according to training rounds. Gradual decay allows the agent to shift its focus to minimizing costs.

[0035] S2. Based on the decision interval, the current decision state is input into the hierarchical reinforcement learning model. The discrete action output layer of the hierarchical reinforcement learning model outputs candidate discrete topology decisions, and the continuous action output layer outputs candidate continuous flow decisions. Anti-mixing interlock rules and action inertia maintenance rules are then applied to the candidate discrete topology decisions.

[0036] If an agent relies solely on immediate responses to short-term rewards, it may frequently switch active containers, resulting in chattering behavior that can damage actuators and incur excessively high switching costs. In traditional numerical optimization, chattering is typically addressed by adding an explicit switching penalty term to the objective function. To suppress, among which, For the current scheduling time Discrete topological actions, For the previous scheduling time The discrete topological action directly penalizes changes in the control action at adjacent time steps. However, in reinforcement learning frameworks, this penalized policy is difficult to implement effectively. If it is incorporated into reward shaping, it only exists as a soft constraint: when short-term gains are high enough, the agent may still tolerate frequent switching, especially in the early stages of training when the policy has not yet converged. Furthermore, switching penalty weights... The relative size between the reward and other rewards is difficult to adjust, and improper settings may also disrupt training stability.

[0037] This invention employs a hierarchical reinforcement learning strategy to output scheduling actions. The upper-level strategy outputs discrete logic actions and continuous target flow parameters. Discrete logic actions are used to determine the connection relationships between ships and storage tanks, storage tanks and blending tanks, and blending tanks and crude oil distillation units; continuous target flow parameters are used to provide the expected transport intensity for each side.

[0038] To avoid unsafe actions, this invention constructs a dual security mechanism between the strategy output and execution layers, including a logical security layer and a quadratic differentiable layer.

[0039] The logical safety layer includes hard-interlock rules and inertial rules. Hard-interlock rules are used to block transport paths that are not allowed to be used simultaneously. For example, when a blending tank is feeding a crude oil distillation unit, the blending tank is prohibited from receiving oil from a storage tank at the same time. Inertial rules are used to prioritize keeping the previous topology operation unchanged when the currently active tank still has sufficient inventory, in order to reduce frequent switching.

[0040] This logical safety layer enforces stability directly at the policy output level, rather than relying solely on reward signals for guidance. Specifically, it constructs a heuristic inertial masking mechanism to modify the logit distribution of the discrete policy header before sampling. Let... These are the raw logits output for a specific resource. Based on the actions from the previous time step. With the current active tank volume Apply the following conditional bias: in, It is a large constant bias (e.g.) ), It has a safety margin The minimum volume threshold, Candidate discrete actions after inertial mask correction The corresponding logit value, For candidate discrete actions, This is an indicator function that takes a value of 1 when the condition within the parentheses is true, and a value of 0 otherwise. This logic forms a soft lock: if the current active tank still has sufficient inventory, i.e. If the system fails to do so, the intelligent system will be strongly biased towards continuing to use the container. Essentially, it functions as a low-pass filter in decision dynamics, reducing high-frequency switching by over 95% in experiments while still allowing timely switching when the container is nearly empty.

[0041] S3. Based on the discrete topology decision after rule screening, construct a quadratic programming model with the objective of minimizing the deviation between the candidate continuous flow decision and the corrected flow decision, and with the upper bound of dynamic flow and the minimum feed requirement as constraints. Solve the model to obtain the corrected continuous flow decision and generate executable scheduling control instructions.

[0042] The original continuous quantity output by the policy network This represents the desired traffic, but may violate physical constraints or logical interlocks. To strictly guarantee feasibility while maintaining differentiability, a quadratic differentiable programming layer is introduced. This layer projects the original continuous actions output by the policy network to find the executable action that best approximates the original action and satisfies all physical constraints. in, The original continuous target traffic output by the policy network. For discrete topological actions, This is an upper bound for dynamic traffic related to the current state and topology actions. The executable continuous action is obtained after projection through a second-order differentiable programmable layer. For the continuous flow decision variable to be optimized, For the index of the feed tank, For indexing crude oil distillation units, A collection of crude oil distillation units. This is a collection of feeding tanks.

[0043] Furthermore, Constraints provide dynamic upper bounds It integrates topological logic, structural limits, and inventory availability. For each flow edge... Its upper boundary Defined as: in, It is a discrete action logic gate. This represents the maximum expected time interval between adjacent events in SMDP. Source tank Maximum safe volume threshold For storage tanks exist Real-time volume at any given moment. When activated, the corresponding flow rate is forced to zero; when activated, the flow rate limit is determined by the most stringent of the following three factors: the pipeline's physical rated limit. Source tank Maximum discharge volume and receiving tank Available empty capacity This effectively prevents cavitation, receiving tank overflow, and violation of mass conservation. More importantly, it ensures the proper functioning of the logic gate vectors. It also incorporates a hard-coded interlocking mechanism for the mixing tank. To prevent contamination, when the mixing tank... While feeding a CDU, it is not allowed to simultaneously receive oil from a storage tank. This rule is set by... This ensures strict separation of operating modes.

[0044] S4. Based on the execution results of executable scheduling control instructions, calculate the total operating cost, and use the total operating cost as a feedback signal to drive the hierarchical reinforcement learning model to update the strategy online, thereby optimizing crude oil scheduling in the refinery.

[0045] To verify the performance of the present invention, the following experiments were designed, and the experimental results are shown in Table 1.

[0046] Table 1 Performance Comparison of Different Methods A real-world refinery scheduling test platform was built based on a real-world case study to validate the proposed framework. This scenario simulates the operation of a coastal refinery's front-end over a continuous 15-day time domain, such as... Figure 2 As shown. The system configuration includes: a shipping terminal with 3 berths for handling crude oil imports; a tank farm containing 6 storage tanks and 4 blending tanks for intermediate buffering; and 3 crude oil distillation units (CDUs) with different demand curves and concentration constraints. To simulate real disturbances, we introduce stochastic uncertainty: ship arrival times follow a truncated normal distribution. ,in The number of days indicates significant fluctuations in maritime logistics, and this distribution is strictly truncated within [a specific timeframe]. Within a certain timeframe, to avoid physically unreasonable negative delays. In addition, crude oil quality fluctuations are mitigated by superimposing Gaussian noise. Modeling is based on sulfur content, where This represents the standard deviation of the concentration error, used in the hybrid control logic of the explicit challenge system.

[0047] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning, characterized in that, The method includes: S1. Obtain the status information of crude oil scheduling in the refinery, construct an event-driven semi-Markov decision model, and dynamically determine the next decision interval of the system according to the preset triggering events; S2. Based on the decision interval, the current decision state is input into the hierarchical reinforcement learning model. The discrete action output layer of the hierarchical reinforcement learning model outputs candidate discrete topology decisions, and the continuous action output layer outputs candidate continuous flow decisions. Anti-mixing interlock rule screening and action inertia maintenance rule screening are performed on the candidate discrete topology decisions. S3. Based on the discrete topology decision after rule screening, construct a quadratic programming model with the objective of minimizing the deviation between the candidate continuous flow decision and the corrected flow decision, and with the upper bound of dynamic flow and the minimum feed requirement as constraints. Solve the model to obtain the corrected continuous flow decision and generate executable scheduling control instructions. S4. Based on the execution result of the executable scheduling control instruction, calculate the total operating cost, and use the total operating cost as a feedback signal to drive the hierarchical reinforcement learning model to perform online policy updates, thereby realizing refinery crude oil scheduling optimization.

2. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 1, characterized in that, The status information includes: the current scheduling time, the remaining oil volume, crude oil properties and estimated arrival time of each crude oil tanker to be unloaded, the real-time inventory, minimum safety stock, maximum safety stock and crude oil properties of each storage tank, the real-time inventory, maximum allowable inventory and crude oil properties of each feed tank, the feed demand of each atmospheric and vacuum distillation unit, and the pipeline start / stop status or topology connection status of the previous decision cycle.

3. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 1, characterized in that, The next decision interval Represented as: in, This refers to the time required for the next crude oil tanker to arrive at the port. For storage tanks or filling tanks to reach their safety threshold, This indicates that the transfer operation has been completed.

4. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 1, characterized in that, The discrete action output layer employs a heuristic inertial masking mechanism, modifying the logit distribution of the discrete strategy header before sampling. A bias is applied based on the action from the previous time step and the volume of the current active tank, expressed as: in, Candidate discrete actions after inertial mask correction The corresponding logit value, This is the raw logits output. For candidate discrete actions, For constant bias, This is an indicator function that takes the value 1 when the condition within the parentheses is true, and 0 otherwise. The action in the previous moment. For the current active tank volume, Minimum volume threshold, For safety margin.

5. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 1, characterized in that, The hierarchical reinforcement learning model constructs a dual security mechanism, including a logical security layer and a quadratic differentiable layer.

6. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 5, characterized in that, The logical safety layer includes hard interlock rules and inertial rules; the hard interlock rules are used to block transport paths that are not allowed to be used simultaneously; the inertial rules are used to prioritize keeping the previous topology action unchanged when the current active tank has sufficient inventory.

7. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 5, characterized in that, The quadratic differentiable layer projects the output of the original continuous actions to obtain the executable action that is closest to the original action and satisfies all physical constraints, expressed as: in, The original continuous target traffic output by the policy network. For discrete topological actions, This is an upper bound for dynamic traffic related to the current state and topology actions. The executable continuous action is obtained after projection through a second-order differentiable programmable layer. For the continuous flow decision variable to be optimized, For the index of the feed tank, For indexing crude oil distillation units, A collection of crude oil distillation units. This is a collection of feeding tanks.

8. The refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 1, characterized in that, The reward function of the hierarchical reinforcement learning model is expressed as the sum of the negative of the total operating cost and the reward shaping term.

9. A refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 8, characterized in that, The total operating cost is a weighted sum of all costs by their respective unit weights, including inventory holding costs, unloading operation costs, vessel waiting costs, and mode switching costs.

10. A refinery crude oil scheduling optimization method based on event-triggered security reinforcement learning according to claim 8, characterized in that, The reward shaping item Used to guide exploration, represented as: in, For the already processed crude oil, the coefficient is... Adopting a course-based learning strategy, For training rounds.

Citation Information

Patent Citations

  • Oil refinery production and maintenance collaborative optimization method based on reinforcement learning

    CN118966431A