Optimization control method and equipment for refinery plant production device, medium and product
By combining deep reinforcement learning with mathematical programming models, a scheduling model and reinforcement learning environment for refinery production units are constructed. This solves the problem that traditional methods are difficult to adapt to in dynamic environments, and achieves efficient and feasible scheduling optimization of refinery production units, thereby improving operational efficiency and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA UNIV OF SCI & TECH
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional refinery production unit optimization methods are difficult to adapt flexibly to fluctuations in unit operating conditions and market demand, making it difficult for scheduling schemes to balance feasibility and economy in dynamic environments. Furthermore, scheduling results generated by purely data-driven or purely reinforcement learning methods under complex constraints are either infeasible or lack engineering interpretability.
By combining deep reinforcement learning and mathematical programming models, a scheduling model and reinforcement learning environment for refinery production units are constructed. The trained reinforcement learning model predicts product output parameters, and the feasibility is verified and calculated in the unit scheduling model. The scheduling parameters are adaptively adjusted to meet operational constraints and material balance.
This approach enables flexible adaptation to market and operating condition uncertainties while ensuring the feasibility of the scheduling plan. It significantly improves the economy, stability, and automation level of refinery production scheduling, and ensures that the scheduling plan strictly meets the requirements of unit capacity and material balance.
Smart Images

Figure CN121900352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial control technology, and in particular to an optimized control method, equipment, medium and product for a refinery production unit. Background Technology
[0002] The petrochemical industry is a crucial foundation of the national economy, and its production efficiency and stability are of strategic importance to economic security and resource conservation. Refinery production units are subject to the coupled influence of multiple constraints during operation, including processing capacity, material balance, storage tank capacity, and product delivery plans. Furthermore, in actual production, both unit operating conditions and market demand are uncertain, and the material flow network is complex, leading to significant challenges in refinery production scheduling.
[0003] Currently, traditional optimization methods typically rely on mathematical programming. These methods establish mathematical models to precisely characterize the operational constraints of the equipment, material balance relationships, and changes in tank inventory, thereby solving for the optimal equipment load and material flow under given parameters. However, the scheduling parameters of traditional methods (such as product output) usually need to be fixed in advance or set based on manual experience. This makes it difficult for fixed parameters to adapt flexibly in actual production, especially when equipment operating conditions or market demand fluctuate. Frequent manual adjustments lack a systematic approach, making it difficult for scheduling schemes to continuously balance feasibility and economy in dynamic environments.
[0004] Therefore, there is an urgent need for an optimized control method for refinery production units to improve their operating efficiency and resource utilization. Summary of the Invention
[0005] This invention provides an optimized control method, equipment, medium, and product for refinery production units, which improves the operating efficiency and resource utilization of refinery production units.
[0006] In a first aspect, this application provides an optimized control method for a refinery production unit, the method comprising: Based on the production data of the refinery's production units, a unit scheduling model and a reinforcement learning environment are constructed. The production data represents the production operation data related to the refinery's production units, storage tanks, and product output. The unit scheduling model is used to describe the operating constraints, material balance relationships, product flow relationships, and storage tank inventory changes of the refinery's production units. The state space of the reinforcement learning environment represents the overall operating status of the refinery's production units during the current scheduling period. Based on the trained reinforcement learning model and combined with the state information of the reinforcement learning environment, the product output parameters for the current scheduling period are predicted; the reinforcement learning model is obtained through offline training based on the historical production data of the refinery's production units. Based on the aforementioned equipment scheduling model, the feasibility of the product output parameters is verified and calculated to determine the scheduling decision results for the current scheduling period, so as to optimize the control of the operating load, material flow, and storage tank inventory status of the refinery production equipment.
[0007] Optionally, after determining the scheduling decision result for the current scheduling period, the method further includes: Based on the scheduling decision results, the status information of the refinery production unit is updated; Based on the scheduling decision results and preset device constraints, a decision evaluation is performed to determine the corresponding evaluation information; the device constraints are used to constrain the operation of the refinery production unit, material balance, and storage tank inventory, and the evaluation information characterizes the degree to which the scheduling decision results violate the device constraints. Based on the evaluation information, the scheduling decision result is corrected.
[0008] Optionally, the construction of the unit scheduling model based on the production base data of the refinery's production units includes: Based on the aforementioned production data, the decision variables, equipment constraints, and objective function of the equipment scheduling model are determined. Based on the decision variables, the device constraints, and the objective function, and combined with a preset discrete-time mathematical programming strategy, the device scheduling model is constructed. The decision variables represent the operating load of the refinery production unit in each scheduling period, the flow rate of each material stream in each scheduling period, and the inventory level of each storage tank in each scheduling period. The objective function is used to minimize at least one of the following during the scheduling cycle: inventory fluctuation of intermediate material storage tanks, discharge volume of intermediate material storage tanks, and the number of product storage tanks used.
[0009] Optionally, the construction of the reinforcement learning environment based on the production data of the refinery's production units includes: Based on the tank inventory information in the production base data, determine the tank inventory state components in the state space; Based on the product outgoing demand information in the production base data, determine the product remaining outgoing demand state components in the state space. Based on the time identifier information of the current scheduling period, the time state component in the state space and the action space of the reinforcement learning environment are determined; the action space represents the planned output of each product within the current scheduling period.
[0010] Optionally, the reinforcement learning model is trained based on the following steps: Based on the initial state information corresponding to the historical production base data, the initial scheduling parameters of the reinforcement learning model in the initial scheduling period are determined. The initial scheduling parameters are solved based on the device scheduling model to obtain the initial scheduling decision results and initial evaluation information; Based on the initial scheduling decision results, update the status information of the refinery production units in the next scheduling period; Based on the updated state information and the initial evaluation information, the reinforcement learning model is iteratively trained until the reinforcement learning model converges, thus obtaining the trained reinforcement learning model.
[0011] Optionally, the step of performing feasibility verification and calculation on the product output parameters based on the device scheduling model to determine the scheduling decision result for the current scheduling period includes: Based on the device constraints of the device scheduling model, the feasibility of the product output parameters is verified. When the product output parameters are determined to meet the device constraints, the device scheduling model is solved to obtain the scheduling decision result.
[0012] Optionally, the production base data includes at least the rated processing capacity of the refinery production units, the upper and lower limits of unit load, the upper and lower limits of storage tank inventory, the initial inventory information of storage tanks, the inflow and outflow relationship between the refinery production units and storage tanks, the discharge diversion relationship of the units, and the product outflow demand information within the scheduling cycle.
[0013] In a second aspect, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the refinery production equipment optimization control methods described in the first aspect above.
[0014] Thirdly, this application provides a computer storage medium storing computer program instructions, which are executed by a processor using any of the refinery production unit optimization control methods described in the first aspect above.
[0015] Fourthly, an embodiment of this application provides a computer program product including computer program instructions, which, when executed by a processor, implement the optimization control method for any of the refinery production units described in the first aspect above.
[0016] The beneficial effects of this invention are as follows: This application provides an optimized control method for a refinery production unit. This method constructs a unit scheduling model and a reinforcement learning environment based on the refinery's production data. Based on the trained reinforcement learning model and the state information of the reinforcement learning environment, it predicts the product output parameters for the current scheduling period. Then, based on the unit scheduling model, it performs feasibility verification and calculation on the product output parameters to determine the scheduling decision for the current scheduling period. This optimizes the control of the refinery's operating load, material flow, and storage tank inventory status, thereby improving the refinery's operating efficiency and resource utilization. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 A flowchart illustrating an optimized control method for a refinery production unit provided in this application embodiment; Figure 2 A schematic diagram of a basic production data table for a refinery system provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the relationship between total reward value and training steps, provided in an embodiment of this application; Figure 4 A schematic diagram of a device scheduling optimization framework based on reinforcement learning and mathematical programming provided for an embodiment of this application; Figure 5 A schematic diagram of another optimized control method for a refinery production unit provided in an embodiment of this application; Figure 6 A schematic diagram illustrating the daily output arrangement of each product within a scheduling cycle, provided as an embodiment of this application; Figure 7 A schematic diagram illustrating the operating load results of various production units within a scheduling cycle, provided as an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0020] The terms "first" and "second" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and this application does not impose limitations.
[0021] The term "and / or" in the embodiments of this application is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0022] It is understood that the following specific embodiments of this application involve data related to refinery production, etc. When the various embodiments of this application are applied to specific products or technologies, relevant licenses or consents are required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, relevant volunteers can be recruited and agreements can be signed to authorize their data, thereby enabling the implementation using the data of these volunteers; or, implementation can be carried out within an authorized organization, using data from members of the organization to implement the following implementation methods for data management; or, the relevant data used in the specific implementation may be simulated data, such as simulated data generated in a virtual scene.
[0023] The design concept of the embodiments of this application is briefly introduced below: The petrochemical industry is a crucial foundation of the national economy, and its production efficiency and stability are of strategic importance to economic security and resource conservation. Refinery production units are subject to the coupled influence of multiple constraints during operation, including processing capacity, material balance, storage tank capacity, and product delivery plans. Furthermore, in actual production, both unit operating conditions and market demand are uncertain, and the material flow network is complex, leading to significant challenges in refinery production scheduling.
[0024] Currently, traditional optimization methods typically rely on mathematical programming. These methods establish mathematical models to precisely characterize the operational constraints of the equipment, material balance relationships, and changes in tank inventory, thereby solving for the optimal equipment load and material flow under given parameters. However, the scheduling parameters of traditional methods (such as product output) usually need to be fixed in advance or set based on manual experience. This makes it difficult for fixed parameters to adapt flexibly in actual production, especially when equipment operating conditions or market demand fluctuate. Frequent manual adjustments lack a systematic approach, making it difficult for scheduling schemes to continuously balance feasibility and economy in dynamic environments.
[0025] Furthermore, traditional data-driven methods based on historical data typically make decisions by learning experiential patterns to automatically adjust scheduling parameters. However, when faced with the complex process constraints and stringent material balance requirements of refineries, the resulting scheduling may be infeasible or lack clear engineering interpretability, thus limiting its application in real-world industrial scenarios where safety and reliability are paramount. On the other hand, pure reinforcement learning methods have the potential to adaptively update decision-making strategies based on environmental conditions. However, in the high-dimensional, complex constraint decision-making problem of refinery production unit scheduling, relying solely on reinforcement learning is prone to convergence issues during training, and the learned strategies often fail to guarantee compliance with process constraints and material balance, thus failing to generate feasible scheduling solutions.
[0026] In view of the above problems, this application provides an optimized control method for refinery production units. This method constructs a unit scheduling model and a reinforcement learning environment based on the basic production data of the refinery production units. Based on the trained reinforcement learning model and the state information of the reinforcement learning environment, it predicts the product output parameters for the current scheduling period. Based on the unit scheduling model, it performs feasibility verification and calculation on the product output parameters to determine the scheduling decision result for the current scheduling period. This optimizes the control of the operating load, material flow, and storage tank inventory status of the refinery production units, thereby improving the operating efficiency and resource utilization of the refinery production units.
[0027] Furthermore, this application combines deep reinforcement learning with mathematical programming models. Under the premise of satisfying unit operation constraints and material balance relationships, it achieves adaptive adjustment of scheduling model parameters, thereby obtaining a stable, feasible, and highly engineering-applicable refinery production unit scheduling scheme. Moreover, this application creatively integrates data-driven reinforcement learning with mathematical programming models, utilizing the perception and decision-making capabilities of the reinforcement learning model to adaptively and dynamically adjust key parameters in scheduling (e.g., product output), thus flexibly adapting to market and operating condition uncertainties and overcoming the low flexibility caused by the fixed parameters of traditional mathematical programming methods. Simultaneously, this application inputs the parameters output by reinforcement learning into the unit scheduling mathematical programming model in real time for solving, ensuring that the final generated scheduling scheme strictly meets complex physical and process constraints such as unit capacity, material balance, and tank capacity, thereby effectively solving the defects of poor feasibility and insufficient engineering interpretability of pure data-driven or pure reinforcement learning methods. Thus, the method provided by this application can achieve coordinated optimization of output plans and production load while ensuring the absolute feasibility of the scheduling scheme, significantly improving the overall economy, stability, and automation level of refinery production scheduling.
[0028] Please refer to Figure 1 The following is a flowchart illustrating an optimized control method for a refinery production unit according to an embodiment of this application. The specific implementation process of this method is as follows: Step 101: Based on the production data of the refinery's production units, construct a unit scheduling model and a reinforcement learning environment.
[0029] In this application embodiment, the production basic data is the production operation data related to refinery production units, storage tanks, and product delivery. Based on the production basic data, this application will pre-build a unit scheduling model and a reinforcement learning environment.
[0030] Specifically, this application involves the structured collection and organization of basic production data required for refinery production unit scheduling optimization. This basic production data includes unit information data, tank information data, product output information data, and material distribution relationship data. In this application, the basic production data is a structured set of information describing the physical structure and operating rules of the refinery. This may include, but is not limited to, the rated processing capacity of each production unit within the refinery, the upper and lower limits of allowable operating loads, the material yield of each unit, the upper and lower limits of inventory and initial inventory of each tank, the connection relationships of all material flows between production units and tanks, the distribution destinations of materials produced by each unit, and the planned output demand of each product within the current scheduling cycle. This provides basic data support for subsequent unit scheduling model construction and reinforcement learning decision-making.
[0031] In one possible implementation, this application will collect and organize the basic production data required for the optimization of refinery production unit scheduling in a structured manner. For example, the basic production data in this application may include unit information data, storage tank information data, finished product information data, and material distribution relationship data.
[0032] Specifically, the equipment information data in this application can be used to describe the basic operating characteristics of each production unit in the refinery, including the identification information of each unit, equipment type information, rated processing capacity of the unit, upper and lower limit coefficients of the unit's operating load, the set of feed materials of the unit, the set of output materials of the unit, and the material yield information of different side lines of each unit, which are used to characterize the output structure characteristics of the unit.
[0033] Furthermore, this application can define a set of refinery processing units as follows:
[0034] Each of them This represents a refinery processing unit. For any given unit... Define the following device operating parameters: Rated processing capacity of the device:
[0035] Upper and lower limits of the unit's operating load:
[0036] The output material of the device Corresponding material yield:
[0037] For the device The set of feed materials and the set of output materials are defined as follows:
[0038] Used to describe the material connection relationships between the device and upstream or downstream devices or storage tanks.
[0039] Specifically, the tank information data in this application can be used to describe the inventory status and material storage constraints of the refinery tank system, including the identification information of each tank, tank type information, current tank inventory, upper and lower limits of tank inventory, as well as the corresponding feed stream and discharge stream information of the tank, which are used to characterize the inventory change boundary of the tank during the scheduling cycle and its material connection relationship with the equipment.
[0040] Furthermore, this application may define a set of refinery storage tanks as follows:
[0041] Each of them This represents a storage tank. For any given storage tank, the following parameters are defined: Initial inventory at the start of the scheduling cycle:
[0042] Upper and lower limits for tank inventory:
[0043] For storage tanks The set of feed streams and the set of material feed streams are defined as follows:
[0044] Used to describe the material interaction relationships between storage tanks and equipment.
[0045] Specifically, the product information data in this application may include the product's storage tank information and the planned product shipment volume within the scheduling cycle.
[0046] Specifically, this application may define the product storage tank set as:
[0047] Define the set of planned output quantities of products that need to be shipped as follows:
[0048] Specifically, the material diversion relationship data in this application is used to describe the distribution relationship of the material produced by the device in different flow directions, including the material diversion identification information corresponding to each device and the corresponding relationship of each stream after diversion, which is used to depict the flow structure of the material produced by the device in downstream devices or storage tanks.
[0049] Furthermore, this application may define the refinery material set as:
[0050] For a certain output material m of device u, its branch stream set is defined as:
[0051] After completing the collection and organization of the above-mentioned basic production data, based on the correspondence between equipment, materials and storage tanks, a material flow network structure of the refinery production equipment system is constructed, which provides basic data support for the variable definition, constraint construction and reinforcement learning environment state space construction of the subsequent equipment scheduling mathematical programming model.
[0052] In one possible implementation, the refining system in this application may include, but is not limited to, 14 main production units, 18 storage tanks, and over 140 material flow relationships. The refinery production units include, but are not limited to, catalytic cracking units, hydrotreating units, reforming units, and other processing units. Intermediate materials and products are transferred between these units via multiple material flow streams. Storage tanks include intermediate material storage tanks and product storage tanks, used to hold materials of different processing stages and properties. Material flow directions describe the relationships between unit feed, unit output, unit product diversion, and storage tank feeding. Therefore, the refining system in this application may involve nine types of finished products, with a scheduling cycle set to 10 days, and the scheduling cycle is divided using a discrete time approach.
[0053] For details, please refer to Figure 2 The diagram shown is a schematic diagram of a refinery system production basic data table provided in an embodiment of this application. Figure 2 This application demonstrates the collection of basic production data from the refining system, with a collection period of 100 days and a collection interval of once a day, used to construct the state space of the device scheduling mathematical programming model and the reinforcement learning environment.
[0054] In one possible implementation, this application embodiment will determine the decision variables, device constraints, and objective function of the device scheduling model based on production base data, and construct a device scheduling model for refinery production device scheduling optimization based on the decision variables, device constraints, and objective function, combined with a preset discrete-time mathematical programming strategy, to describe device operating constraints, material balance relationships, material flow relationships, and storage tank inventory update processes.
[0055] Specifically, in the embodiments of this application, the decision variables of the unit scheduling model represent the operating load of the refinery production unit in each scheduling period, the flow rate of each material stream in each scheduling period, and the inventory of each storage tank in each scheduling period. The constraints are set to include unit operating load constraints, material balance constraints, material diversion ratio constraints, and storage tank inventory balance and upper and lower limit constraints. The objective function is set to minimize at least one of the following during the scheduling cycle: inventory fluctuation of intermediate material storage tanks, discharge volume of intermediate material storage tanks, and number of product storage tanks used.
[0056] In one possible implementation, embodiments of this application will collect and organize basic production data to construct a mathematical programming model for optimizing refinery production unit scheduling. This mathematical programming model can be constructed based on discrete time and is used to characterize the dynamic changes of each production unit, material flow stream, and storage tank inventory within the scheduling cycle.
[0057] Specifically, the decision variables of the mathematical programming model for equipment scheduling in this application may include: the operating load of each production unit in each discrete time period, the material flow rate of each unit in and out of the equipment and the flow rate of the diversion stream in each discrete time period, and the inventory of each storage tank in each discrete time period, which are used to characterize the equipment operating status, material flow and inventory status.
[0058] Specifically, the constraints of the device scheduling mathematical programming model in this application may include: Unit operating load constraints are used to limit the processing load of each unit to meet the upper and lower limits of the unit load constraints. This indicates the load on a certain device.
[0059] Material balance constraints are used to describe the mass conservation relationships during the feeding, output, and material diversion processes of a unit: Feed balance of the device:
[0060] Calculation of output materials from the unit:
[0061] Material diversion constraints:
[0062] in, This indicates the flow rate of the stream.
[0063] Tank inventory balance constraints are used to describe the inventory change process of each tank in different time periods and limit the inventory level to a preset upper and lower limit range.
[0064] in, This represents the tank inventory within that discrete time period. This indicates the number of cans shipped out during this discrete time period.
[0065] And the supply and demand balance constraints of key utility materials such as hydrogen:
[0066] in, This indicates the amount of hydrogen required by the device in the current discrete time period.
[0067] Specifically, the objective function of the mathematical programming model for equipment scheduling in this application can be used to optimize the inventory status of intermediate materials and the operating status of the equipment, under the premise of satisfying the above constraints, so as to reduce inventory deviation and unnecessary intermediate material turnover. The objective function can be as follows:
[0068] in, This refers to a collection of intermediate material storage tanks. This indicates the number of product tanks used. The first item represents the standardized deviation between the ending and initial inventory of intermediate tanks, used to reduce fluctuations in intermediate tank inventory; the second item represents the total output from intermediate tanks, used to reduce tank supply and encourage direct supply between units; the third item represents the number of storage tanks used for product delivery, reducing the usage of product storage tanks. This is represented as weight.
[0069] In one possible implementation, the embodiments of this application will determine the tank inventory state component in the state space based on the tank inventory information in the production basic data, determine the product remaining output demand state component in the state space based on the product output demand information in the production basic data, and determine the time state component in the state space and the action space of the reinforcement learning environment based on the time identifier information of the current scheduling period.
[0070] Specifically, in this embodiment of the application, the reinforcement learning environment will provide a simulated interaction framework for subsequent intelligent decision-making. Its state space is defined as including the inventory status of each storage tank in the current scheduling period, the remaining unfulfilled factory demand of each product in the scheduling cycle, and the time identifier of the current scheduling period. Its action space is defined as the planned factory output of each product in the current scheduling period, which is a set of continuous values.
[0071] In one possible implementation, the reinforcement learning model in this application can adopt a single-agent structure to make unified decisions on the overall operating status of multiple production units and storage tanks. The state space of the reinforcement learning environment may include: the inventory status of each storage tank at the current scheduling moment, the product output demand information within the scheduling cycle, and the time state within the scheduling cycle, which are used to comprehensively characterize the overall production operation status of the refinery at the current scheduling moment; wherein, the time state is used to reflect the stage position of the current scheduling moment within the scheduling cycle, so as to assist the reinforcement learning strategy module in making temporal decisions.
[0072] Specifically, the state space in this application can be defined as:
[0073] in, This indicates the inventory status of each storage tank at the current scheduling moment; This indicates the remaining factory demand for each product at the current scheduling moment; This represents the time step index of the current scheduling moment within the scheduling cycle, used to reflect the phased characteristics of the scheduling process.
[0074] Specifically, the action space of the reinforcement learning environment in this application can be a continuous action space, including the product output of the storage tank over a time period. This action space can be defined as:
[0075] in, This indicates the quantity shipped for each product.
[0076] In one possible implementation, the reinforcement learning module in this application can output the product output quantity based on the current state space and transmit the output quantity as guiding information to the device scheduling mathematical programming model, so as to adjust the relevant decision variables or constraint parameters in the mathematical programming model, thereby guiding the solution direction of the scheduling model.
[0077] Specifically, in order to reduce the impact of dimensional differences between different storage tank inventory ranges and different product demand quantities on the stability of model training, this application will normalize the continuous variables in the state space and action space using the Max-Min method, that is, for any variable value... The normalization result is:
[0078] in, and These represent the minimum and maximum values of the variable within historical operating data or a preset physically feasible range, respectively. This method can eliminate the influence of different units and scales, ensuring that various numerical features are trained on the same scale.
[0079] Step 102: Based on the trained reinforcement learning model and combined with the state information of the reinforcement learning environment, predict the product output parameters for the current scheduling period.
[0080] In this embodiment, the reinforcement learning model is obtained by simulating the operation of the refinery using historical production data and undergoing extensive offline training. Thus, this application can use the trained reinforcement learning model to predict the output parameters of each product in the current time period based on the status information of the current scheduling period.
[0081] In one possible implementation, this application embodiment determines the initial scheduling parameters of the reinforcement learning model during the initial scheduling period based on the initial state information corresponding to historical production baseline data. The initial scheduling parameters are then solved using the unit scheduling model to obtain the initial scheduling decision results and initial evaluation information. Next, based on the initial scheduling decision results, the state information of the refinery production units in the next scheduling period is updated. Based on the updated state information and the initial evaluation information, the reinforcement learning model is iteratively trained until it converges, resulting in a well-trained reinforcement learning model.
[0082] Specifically, during model training, this application will continuously optimize its internal decision-making strategy by outputting product output parameters in the action space and receiving reward signals from the unit scheduling model after solving the problem, until the strategy converges and stabilizes. After completing offline training of the reinforcement learning model and obtaining a converged reinforcement learning model, the application will predict the unit operating status within the scheduling cycle based on the converged model. For example, in the online application phase of the model, the application will acquire the state information representing the current operating point of the refinery, i.e., the specific values in the state space of the reinforcement learning environment, including the actual inventory of each tank, the remaining planned demand of each product, and the current time step, etc. This state information will be input into the trained reinforcement learning model. Based on its learned strategy, the model can automatically infer and output optimized product output parameters for the current scheduling period, and generate a refinery production unit scheduling scheme that satisfies unit operating constraints and material balance constraints, i.e., an intelligent decision-making scheme for high-level scheduling instructions such as "when to ship and how much to ship".
[0083] In one possible implementation, the present application embodiment may use the SoftActor-Critic (SAC) algorithm to experiment on the reinforcement learning model. This algorithm is based on the maximum entropy reinforcement learning framework and improves the stability of the training process while ensuring policy exploration by simultaneously learning the policy network and the value network. It is suitable for industrial scheduling scenarios with continuous decision variables and complex constraints.
[0084] For details, please refer to Figure 3 The diagram shown illustrates the relationship between total reward value and training steps according to an embodiment of this application. Figure 3 This paper demonstrates the relationship between the total reward value and the number of training steps in the reinforcement learning training process of the model in this application. The learning rate can be set to 0.001, the maximum number of training steps is 300,000, the batch size is 64, and the discount factor is 0.99. As the number of training steps increases, the total reward of the model in the scheduling cycle shows a steady upward trend and tends to converge after about 150,000 steps.
[0085] Step 103: Based on the equipment scheduling model, perform feasibility verification and calculation on the product output parameters to determine the scheduling decision result for the current scheduling period.
[0086] In this embodiment, performing feasibility verification and solution calculation on the equipment scheduling model is crucial for transforming decision-making from parameter instructions into detailed feasible solutions. This application uses the product output parameters predicted by the reinforcement learning model as known input conditions, which are then input into the constructed equipment scheduling mathematical programming model. For example, the product output parameters are substituted into the corresponding parts of the equipment scheduling model, such as the tank inventory balance constraint, thereby solving the equipment scheduling model.
[0087] Specifically, after training the reinforcement learning module and obtaining a converged reinforcement learning model, this application uses the converged model for refinery production unit scheduling prediction. For example, based on the current or planned tank inventory status and product output demand information of the refinery, the state for the corresponding scheduling period is constructed, and the state is input into the converged reinforcement learning model to obtain the product output prediction results for each scheduling period. Then, the product output prediction results are used as input parameters to the unit scheduling mathematical programming model. The unit scheduling mathematical programming model solves the problem under the conditions of satisfying unit operation constraints, material balance constraints, and tank inventory constraints, obtaining the unit operating load, material diversion path, and tank inventory changes within the corresponding scheduling period, thereby generating a feasible refinery production unit scheduling scheme. In the process of solving the mathematical programming model of this application, its built-in constraints such as unit operating load constraints, material balance constraints, and upper and lower limits of tank inventory constraints will automatically take effect, ensuring that any solution obtained naturally satisfies all these physical and process constraints. Thus, by solving the model, this application can determine the scheduling decision results for the current scheduling period. The results are complete, specific and feasible production instructions, including the precise operating load that each production unit should adopt, the specific flow rate of each material stream, and the expected inventory of each storage tank at the end of the period, in order to meet all constraints and optimize the objective function under the current output parameters.
[0088] In one possible implementation, the embodiments of this application will perform a feasibility verification on the product output parameters based on the device constraints of the device scheduling model, and when it is determined that the product output parameters meet the device constraints, solve the device scheduling model to obtain the scheduling decision result.
[0089] Specifically, the reinforcement learning model in this application will be used in each scheduling period. Based on current state information The system outputs the output parameters of each product for the corresponding time period and uses these parameters as input to the unit scheduling mathematical programming model to guide the setting of product output demand within the current scheduling period. The status information includes the inventory status of each storage tank, the remaining planned output demand of each product within the scheduling cycle, and the current time period number. Upon receiving the output parameters, the unit scheduling mathematical programming model solves the problem under the conditions of satisfying unit operation constraints, material balance constraints, and storage tank inventory constraints, obtaining the scheduling decision results for the operating load, unit output, material flow, and storage tank inventory status of each production unit within the current scheduling period. Thus, the reinforcement learning model can be used to make decisions on the directional parameters of the scheduling decision, while the unit scheduling mathematical programming model can be used to ensure that the scheduling decision results meet the requirements of physical feasibility and process constraints, thereby forming a collaborative scheduling mechanism combining reinforcement learning and mathematical programming.
[0090] In one possible implementation, refer to Figure 4 The diagram shows a schematic of a device scheduling optimization framework based on reinforcement learning and mathematical programming provided in this application embodiment. This application acquires basic production data related to refinery production units, storage tanks, and product output through a refinery production data acquisition system, and constructs a reinforcement learning environment based on the above data. This environment generates state information including remaining output demand and storage tank inventory status. Then, it feeds back reward values based on the results of subsequent mathematical model solving. The reinforcement learning decision module needs to receive the state information from the reinforcement learning environment, output the product output parameters for the current scheduling period, and input them into the mathematical model solving module. This allows the module to combine the basic production data for rigorous feasibility verification and optimization, generating detailed and feasible scheduling decision results, thereby achieving optimized control of the operating load, material flow, and storage tank inventory status of the refinery production units.
[0091] In one possible implementation, refer to Figure 5 The diagram shown is a schematic representation of another optimized control method for a refinery production unit provided in an embodiment of this application. Figure 5 This application demonstrates the collaborative workflow of offline training and online application in its embodiments. First, basic production data is collected, and based on this, a mathematical programming model for plant scheduling and a reinforcement learning environment are constructed in parallel. During the offline training phase, the reinforcement learning decision outputs parameters for the plant's output, which are then solved by the mathematical programming model. Based on the solution results, state feedback and reward calculations are performed to iteratively update the strategy, and the model is repeatedly checked for convergence. If convergence fails, training continues; if convergence occurs, the application enters the online application phase. Based on the converged model, the scheduling predicts the output for future time periods, which is then solved again by the mathematical programming model. Finally, a feasible scheduling scheme is output to optimize the control of the refinery's operating load, material flow, and storage tank inventory status.
[0092] Step 104: Based on the scheduling decision results, optimize the control of the operating load, material flow and storage tank inventory status of the refinery production units.
[0093] In this embodiment, the obtained scheduling decision results will be applied to the actual production guidance of the refinery production unit, so as to achieve optimized control of the operating load, material flow and storage tank inventory status of the refinery production unit.
[0094] Specifically, this application can format and output the scheduling decision results to form a production scheduling scheme that can be directly issued. For example, the control instructions of this scheduling scheme may include: a set value for the operating load of each production unit, which is within the upper and lower limits of the unit's allowable load and is obtained through material balance and upstream and downstream unit collaborative optimization, thereby achieving precise and stable control of the unit's production intensity. Alternatively, it may provide guidance for the regulation of material flow, clarifying the amount of each material transported between units and storage tanks, and between storage tanks, thereby achieving control over the balance and efficient flow of complex material networks, as well as prediction and management objectives for the inventory status of each storage tank, ensuring that the storage tank inventory always operates within safe upper and lower limits and tends to be stable, thereby achieving proactive control over the material buffer and storage risks of the entire plant. In this way, by executing the above production scheduling scheme, the refinery production system can achieve full-process collaborative optimization operation under the guidance of an adaptively set factory output rhythm.
[0095] In one possible implementation, the embodiments of this application will update the status information of the refinery production unit based on the scheduling decision results, and perform decision evaluation based on the scheduling decision results and preset unit constraints to determine the corresponding evaluation information, and correct the scheduling decision results based on the evaluation information.
[0096] Specifically, in this embodiment, the storage tank inventory status and related operating status are updated based on the scheduling decision results, and the updated status is then fed back to the reinforcement learning strategy module as the status input for the next scheduling period. Simultaneously, evaluation information is calculated based on the scheduling decision results. When a constraint violation occurs, the decision strategy of the reinforcement learning strategy module is corrected by imposing a penalty term on the evaluation information.
[0097] Specifically, in this embodiment, the corresponding evaluation information will be calculated based on the scheduling decision result. This evaluation information may include: 1) Objective function value of the mathematical programming model for device scheduling:
[0098] in, This represents the objective function value of the mathematical programming model.
[0099] 2) Penalty for potential risks associated with product can inventory: Risk of deviating from the safe zone:
[0100] Extreme value risk:
[0101] 3) Rewards and penalties based on batch size:
[0102] in, This is the threshold for the batch size of the factory shipment.
[0103] 4) Plan completion reward: To prevent agents from becoming overly conservative in order to avoid inventory fluctuations, resulting in insufficient execution rates for planned tasks, a reward for completing planned tasks is introduced at the end of the scheduling cycle:
[0104] in, This indicates the actual quantity of products shipped. This indicates the planned output quantity of the product.
[0105] Thus, reinforcement learning in this application is used during the scheduling period. The total reward can be defined as:
[0106] in, This represents the value of a reward or punishment normalized to (-1, 1).
[0107] In summary, this application modifies the decision-making strategy of the reinforcement learning strategy module by introducing a penalty term into the evaluation information when the scheduling decision results in situations such as tank inventory exceeding limits or other constraints being violated. This guides the reinforcement learning strategy module to gradually avoid infeasible or unstable scheduling decisions. Thus, this application can train the model based on offline historical operating data through the aforementioned iterative process of state feedback and evaluation updates until the model converges, thereby forming a reinforcement learning model for plant scheduling optimization.
[0108] In one possible implementation, refer to Figure 6 The diagram shown is a schematic representation of the daily output arrangement of each product within a scheduling period according to an embodiment of this application. This application will apply the model to a scheduling case with a 10-day scheduling period after the model training converges. Figure 6 The results show the daily production volume scheduling of each product under a scheduling cycle of 10 days. Figure 6The "Planned" row represents the planned total output of each product throughout the entire scheduling cycle, while the rows below it represent the actual output allocation for the corresponding product on each scheduling day. The "Incomplete" row represents the small number of computational tasks that remain unfinished for each product at the end of the scheduling cycle. This shows that the output tasks for each product exhibit a non-uniform distribution within the scheduling cycle. The reinforcement learning model, by combining tank inventory status and downstream demand changes, dynamically adjusts the daily output, ensuring overall plan completion while avoiding inventory overflows or fluctuations in plant operation caused by concentrated output.
[0109] Next, refer to Figure 7 The figure shown is an example of the operating load results of various production units during the scheduling cycle provided in this application embodiment. Figure 7 Shown on Figure 6 The results of the operating load of each production unit during the scheduling cycle are shown under the constraints of the factory dispatching scheme. Figure 7 The rated load and corresponding actual operating load of each unit are given. It can be seen that, under the guidance of the factory decision output by reinforcement learning, the mathematical programming model for unit scheduling has coordinated and optimized the load of each unit. The operating load of all units is controlled within their rated load and allowable load range, and there is no obvious overload or underload operation. This shows that the generated scheduling scheme has good feasibility and stability of unit operation while meeting material balance and inventory constraints.
[0110] Please see Figure 8 As shown, based on the same technical concept, this application also provides a computer device 80. In one embodiment, the computer device can be a device specifically for the optimized control of refinery production units, or it can be a device for overall control of energy storage scheduling. The computer device is as follows... Figure 8 As shown, it includes a memory 801, a communication module 803, and one or more processors 802.
[0111] The memory 801 is used to store computer programs executed by the processor 802. The memory 801 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0112] Memory 801 may be volatile memory, such as random-access memory (RAM); memory 801 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 801 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 801 may be a combination of the above-mentioned memories.
[0113] The processor 802 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 802 is used to implement the aforementioned optimized control method for the refinery production unit when it calls the computer program stored in the memory 801.
[0114] The communication module 803 is used to communicate with the industrial control system.
[0115] This application embodiment does not limit the specific connection medium between the memory 801, communication module 803, and processor 802 described above. This application embodiment... Figure 8 The memory 801 and the processor 802 are connected via a bus 804, and the bus 804 is in Figure 8 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 804 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 8 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.
[0116] The memory 801 stores a computer storage medium, which stores computer-executable instructions. The computer-executable instructions are used to implement the optimization control method of the refinery production apparatus in the embodiments of this application. The processor 802 is used to execute the optimization control method of the refinery production apparatus in the above embodiments.
[0117] Based on the same inventive concept, embodiments of this application also provide a storage medium storing a computer program that, when run on a computer, causes the computer to execute the steps in the optimized control method for a refinery production unit according to various exemplary embodiments of this application described above.
[0118] In some possible implementations, various aspects of the optimization control method for refinery production equipment provided in this application can also be implemented in the form of a computer program product, which includes a computer program that, when run on a computer device, causes the computer device to perform the steps in the optimization control method for refinery production equipment according to various exemplary embodiments of this application as described above. For example, the computer device can perform the steps of each embodiment.
[0119] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0120] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computer device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and the computer program included therein may be used by or in conjunction with a command execution system, apparatus, or device.
[0121] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.
[0122] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0123] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages.
[0124] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0125] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0126] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0128] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An optimized control method for a refinery production unit, characterized in that, The method includes: Based on the production data of the refinery's production units, a unit scheduling model and a reinforcement learning environment are constructed. The production data represents the production operation data related to the refinery's production units, storage tanks, and product output. The unit scheduling model is used to describe the operating constraints, material balance relationships, product flow relationships, and storage tank inventory changes of the refinery's production units. The state space of the reinforcement learning environment represents the overall operating status of the refinery's production units during the current scheduling period. Based on the trained reinforcement learning model and combined with the state information of the reinforcement learning environment, the product output parameters for the current scheduling period are predicted; the reinforcement learning model is obtained through offline training based on the historical production data of the refinery's production units. Based on the aforementioned equipment scheduling model, the feasibility of the product output parameters is verified and calculated to determine the scheduling decision results for the current scheduling period, so as to optimize the control of the operating load, material flow, and storage tank inventory status of the refinery production equipment.
2. The method as described in claim 1, characterized in that, After determining the scheduling decision result for the current scheduling period, the method further includes: Based on the scheduling decision results, the status information of the refinery production unit is updated; Based on the scheduling decision results and preset device constraints, a decision evaluation is performed to determine the corresponding evaluation information; the device constraints are used to constrain the operation of the refinery production unit, material balance, and storage tank inventory, and the evaluation information characterizes the degree to which the scheduling decision results violate the device constraints. Based on the evaluation information, the scheduling decision result is corrected.
3. The method as described in claim 1, characterized in that, The construction of the unit scheduling model based on the production base data of the refinery's production units includes: Based on the aforementioned production data, the decision variables, equipment constraints, and objective function of the equipment scheduling model are determined. Based on the decision variables, the device constraints, and the objective function, and combined with a preset discrete-time mathematical programming strategy, the device scheduling model is constructed. The decision variables represent the operating load of the refinery production unit in each scheduling period, the flow rate of each material stream in each scheduling period, and the inventory level of each storage tank in each scheduling period. The objective function is used to minimize at least one of the following during the scheduling cycle: inventory fluctuation of intermediate material storage tanks, discharge volume of intermediate material storage tanks, and the number of product storage tanks used.
4. The method as described in claim 1, characterized in that, The reinforcement learning environment is constructed based on the production data of the refinery's production units, including: Based on the tank inventory information in the production base data, determine the tank inventory state components in the state space; Based on the product outgoing demand information in the production base data, determine the product remaining outgoing demand state components in the state space. Based on the time identifier information of the current scheduling period, the time state component in the state space and the action space of the reinforcement learning environment are determined; the action space represents the planned output of each product within the current scheduling period.
5. The method as described in claim 1, characterized in that, The reinforcement learning model is trained based on the following steps: Based on the initial state information corresponding to the historical production base data, the initial scheduling parameters of the reinforcement learning model in the initial scheduling period are determined. The initial scheduling parameters are solved based on the device scheduling model to obtain the initial scheduling decision results and initial evaluation information; Based on the initial scheduling decision results, update the status information of the refinery production units in the next scheduling period; Based on the updated state information and the initial evaluation information, the reinforcement learning model is iteratively trained until the reinforcement learning model converges, thus obtaining the trained reinforcement learning model.
6. The method as described in claim 1, characterized in that, The process of performing feasibility verification and calculation on the product output parameters based on the device scheduling model to determine the scheduling decision result for the current scheduling period includes: Based on the device constraints of the device scheduling model, the feasibility of the product output parameters is verified. When the product output parameters are determined to meet the device constraints, the device scheduling model is solved to obtain the scheduling decision result.
7. The method as described in claim 1, characterized in that, The production base data includes at least the rated processing capacity of the refinery's production units, the upper and lower limits of unit load, the upper and lower limits of storage tank inventory, the initial inventory information of storage tanks, the inflow and outflow relationship between the refinery's production units and storage tanks, the output diversion relationship of the units, and the product outflow demand information within the scheduling cycle.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer storage medium storing computer program instructions thereon, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 1 to 7.
10. A computer program product comprising computer program instructions, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 1 to 7.