Method, device, and system for determining information on process composed of multiple steps

Multi-agent reinforcement learning enhances process scheduling resilience and efficiency in naphtha cracking centers by creating and evaluating branch schedules to optimize profitability and constraint compliance.

WO2026111391A1PCT designated stage Publication Date: 2026-05-28LG MANAGEMENT DEV INST CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG MANAGEMENT DEV INST CO LTD
Filing Date
2025-11-19
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing reinforcement learning-based scheduling methods for multi-stage processes are prone to failure if a single incorrect choice by the agent invalidates all subsequent steps, posing a risk of process collapse in complex environments like Naphtha Cracking Centers.

Method used

A method utilizing multi-agent reinforcement learning to create and expand branch schedules, evaluate them using a fitness estimator network, and update pivot schedules to determine a final schedule that maximizes profitability and compliance with operational constraints.

Benefits of technology

Improves learning speed and final performance by ensuring resilience against single-agent errors, enabling optimal scheduling in complex processes like naphtha cracking centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025019155_28052026_PF_FP_ABST
    Figure KR2025019155_28052026_PF_FP_ABST
Patent Text Reader

Abstract

One embodiment of the present disclosure may provide a system comprising one or more memories collectively storing instructions which, when executed by one or more processors, instruct the system to perform operations, the operations comprising: an operation of generating a plurality of groups including a first group and a second group; an operation of replicating a pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups; an operation of, for each of the plurality of branch schedules, expanding a schedule by performing a macro operation in which an agent trained on the basis of an artificial intelligence model reflects an operation scenario corresponding to each of the plurality of groups; an operation of evaluating a branch schedule of the first group and another group branch schedule at the same or higher level of constraint as the first group at a first synchronization time point with a suitability estimator network, and updating a pivot schedule of the first group according to the evaluation result; and an operation of determining a final schedule on the basis of the updated pivot schedule of the first group.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, and system for determining information about a process consisting of multiple steps

[0001] The present disclosure relates to a method, apparatus, and system for determining information about a process comprising multiple steps. More specifically, the present disclosure relates to a method, apparatus, and system for determining information about a process comprising multiple steps by utilizing artificial intelligence.

[0002] In complex production environments such as factories, multiple stages of processes are intertwined like dominoes, so there is a risk that the entire schedule will collapse if the schedule is slightly off at any one stage. For example, in a Naphtha Cracking Center (NCC), the receiving stage, where naphtha is unloaded from a ship and stored in tanks; the blending stage, where naphtha from multiple tanks is mixed to a target composition ratio; and the cracking stage, where ethylene and other substances are produced in a high-temperature cracking furnace, are performed in a domino-like sequence. If the timing is slightly off at any one stage, there is a risk that the entire process will be halted, such as by tank overflow, quality degradation, or shutdown of the cracking furnace.

[0003] Recently, various studies have been conducted to solve this multi-stage process scheduling problem, one of which is a Reinforcement Learning (RL)-based approach. Reinforcement Learning is a machine learning technique that trains an agent to observe the state of the environment and select an action or sequence of actions that maximizes rewards, and it is utilized in various fields such as manufacturing, logistics, and energy management. However, existing reinforcement learning-based scheduling methods posed a risk of the entire process failing if a single incorrect choice by the agent invalidated all subsequent steps.

[0004] To solve these problems, the present disclosure proposes a new method for determining information regarding a process consisting of multiple steps. Throughout the entire disclosure, a method according to one embodiment is described in combination with reinforcement learning; however, the embodiments of the present disclosure are not limited to reinforcement learning and can be applied to digital workflows with clearly separated steps, such as machine translation and generative AI model training, using the same principles.

[0005] The present disclosure aims to provide a method, apparatus, and system for determining information about a process comprising a plurality of steps.

[0006] One embodiment of the present disclosure may provide a method, apparatus, and system for determining information about a process comprising a plurality of steps.

[0007] One embodiment of the present disclosure provides a system comprising: one or more processors; and one or more memories that collectively store instructions that cause the system to perform operations when executed by the one or more processors, wherein the operations include: an operation of creating a plurality of groups including a first group and a second group; an operation of replicating a pivot schedule for each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups; an operation of expanding the schedule for each of the plurality of branch schedules by having an agent trained based on an artificial intelligence model perform a macro operation that reflects an operating scenario corresponding to each of the plurality of groups; an operation of evaluating the branch schedule of the first group and the branch schedule of another group with the same or higher constraint level as the first group at a first synchronization time using a fitness estimator network, and updating the pivot schedule of the first group according to the evaluation result; and an operation of determining a final schedule based on the updated pivot schedule of the first group.

[0008] In one embodiment, the final schedule may include a schedule consisting of a plurality of steps.

[0009] In one embodiment, the operation of replicating the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups includes the operation of replicating the pivot schedule of the first group into a first branch schedule and a second branch schedule, and the operation of expanding the schedule by performing a macro operation reflecting an operation scenario for each of the plurality of branch schedules may include the operation of replacing the first branch schedule with the second branch schedule when the first branch schedule of the first group fails.

[0010] In one embodiment, the second branch schedule may be determined by the one or more processors as the schedule having the highest suitability among the plurality of branch schedules of the first group.

[0011] In one embodiment, the operation of replicating a pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups includes: a operation of replicating the pivot schedule of the first group into a first branch schedule and a second branch schedule; and a operation of replicating the pivot schedule of the second group into a third branch schedule and a fourth branch schedule, and the operation of expanding the schedule by performing a macro operation reflecting an operation scenario for each of the plurality of branch schedules may include: a operation of determining that all of the plurality of branch schedules of the first group, including the first branch schedule and the second branch schedule, are failures; and a operation of determining one of the branch schedules of the second group as the pivot schedule of the first group at the time when all of the plurality of branch schedules of the first group have failed.

[0012] In one embodiment, any one of the branch schedules of the second group determined by the pivot schedule of the first group may be determined by the one or more processors as the schedule having the highest degree of suitability among the plurality of branch schedules of the second group.

[0013] In one embodiment, if all schedules of the plurality of groups fail, the operation may include performing a search based on the pivot schedule of each of the plurality of groups at the time of the previous synchronization.

[0014] In one embodiment, the operation of determining the final schedule may include the operation of selecting one or more final schedules by applying post-processing evaluation criteria to the completed pivot schedules after the predetermined planning period has ended.

[0015] In one embodiment, the operation scenario corresponding to each of the plurality of groups may include operation objectives and constraints.

[0016] In one embodiment, the operational objective includes at least one of increased profitability, increased process stability, increased energy efficiency, compliance with quality standards, reduced computing costs, reduced training time, reduced latency, reduced memory usage, and increased model accuracy, and the constraint may include conditions for at least one of equipment operating limits, storage capacity, material property specifications, quality specifications, safety standards, memory capacity, number of available server instances, network bandwidth, and target inference latency.

[0017] In one embodiment, the final schedule includes the operation schedule of the naphtha cracking center, and the first synchronization point may include at least one of the time of arrival of the vessel, the time of completion of the process step, and the time of elapsed simulation time.

[0018] In one embodiment, the operation of expanding the schedule by reflecting the operation scenario for each of the plurality of branch schedules may include the operation of expanding the schedule by having at least one agent perform at least one of a naphtha receiving operation, a blending operation, and a cracking operation based on a pre-learned reinforcement learning policy.

[0019] In one embodiment, the operation of extending the schedule by performing a macro operation reflecting an operation scenario for each of the plurality of branch schedules may include the operation of extending the schedule by performing at least one of a data preprocessing operation, a distributed model learning operation, a fine-tuning operation, and an inference operation.

[0020] One embodiment of the present disclosure provides a method performed by one or more processors, comprising: generating a plurality of groups including a first group and a second group; replicating a pivot schedule for each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups; expanding the schedule for each of the plurality of branch schedules by having an agent trained based on an artificial intelligence model perform a macro operation reflecting an operation scenario corresponding to each of the plurality of groups; evaluating the branch schedule of the first group and the branch schedule of another group with the same or higher constraint level as the first group at a first synchronization time using a fitness estimator network based on the expanded schedule, and updating the pivot schedule of the first group according to the evaluation result; and determining a final schedule based on the pivot schedule of the first group.

[0021] One embodiment of the present disclosure includes a program stored on a recording medium to execute a method according to one embodiment of the present disclosure on a computer.

[0022] One embodiment of the present disclosure includes a computer-readable recording medium having a program for executing a method according to one embodiment of the present disclosure on a computer.

[0023] One embodiment of the present disclosure includes a computer-readable recording medium that records a database used in one embodiment of the present disclosure.

[0024] According to one embodiment of the present disclosure, in determining information about a process consisting of multiple steps using an artificial intelligence algorithm, the learning speed can be improved and the final performance can be improved.

[0025] FIG. 1 is a schematic diagram of a product production process of a naphtha cracking center according to one embodiment of the present disclosure.

[0026] FIG. 2 is a diagram illustrating a reinforcement learning method using multiple agents according to one embodiment of the present disclosure.

[0027] FIGS. 3a and 3b are drawings illustrating a reinforcement learning method using multiple agents according to one embodiment of the present disclosure.

[0028] FIG. 4 is a diagram showing the schedule of a naphtha cracking center according to one embodiment of the present disclosure.

[0029] FIG. 5 is a diagram illustrating a method for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0030] FIG. 6 is a flowchart illustrating a method for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0031] FIG. 7 is a block diagram of a system for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0032] To clarify the technical concept of the present disclosure, embodiments of the present disclosure will be described in detail with reference to the attached drawings. In describing the present disclosure, detailed descriptions of related known functions or components will be omitted if it is determined that such detailed descriptions would unnecessarily obscure the essence of the present disclosure. Components having substantially the same functional configuration among the drawings have been assigned the same reference numerals and symbols as much as possible, even if they are shown in different drawings. For convenience of explanation, devices and methods will be described together where necessary. Each operation of the present disclosure does not necessarily have to be performed in the order described and may be performed in parallel, selectively, or individually.

[0033] The terms used in the embodiments of this disclosure have been selected to be as widely used as possible, taking into account the functions of this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, terms used in this specification should be defined not merely by their names, but based on their meanings and the overall content of this disclosure.

[0034] Throughout this disclosure, singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms such as “comprising” or “having” are intended to specify the presence of features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. That is, throughout this disclosure, when a part is described as “comprising” a certain component, it means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0035] Expressions such as "at least one" modify the entire list of components and do not modify the components of the list individually. For example, "at least one of A, B, and C" and "at least one of A, B, or C" refer to only A, only B, only C, both A and B, both B and C, both A and C, all of A, B, and C, or any combination thereof.

[0036] Additionally, terms such as “...part,” “...module,” etc., as described in this disclosure refer to a unit that processes at least one function or operation, and may be implemented in hardware or software, or a combination of hardware and software.

[0037] Throughout the entire disclosure, when a part is described as being “connected” to another part, this includes not only cases where they are “directly connected” but also cases where they are “electrically connected” with other elements interposed between them. Furthermore, when a part is described as “comprising” a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0038] As used throughout this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware. Instead, in some situations, the expression “system configured to” may mean that the system is “capable of” together with other devices or components. For example, the phrase “a processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing said operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in memory.

[0039] One embodiment of the present disclosure aims to provide a method for determining information regarding a process consisting of multiple steps using artificial intelligence. Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0040] Artificial intelligence (AI) is a field of computer science and information technology that studies methods to enable computers to perform thinking, learning, and self-development—tasks achievable by human intelligence—and refers to the ability of computers to mimic intelligent human behavior. Furthermore, AI does not exist in isolation but is closely related, directly or indirectly, to many other fields of computer science. Particularly in the modern era, there are very active attempts to introduce AI elements into various sectors of information technology and utilize them to solve problems within those fields.

[0041] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0042] Machine learning is a field of artificial intelligence that enables computers to learn without explicit programming. Specifically, machine learning can be defined as a technology that studies and builds systems and algorithms capable of learning, making predictions, and improving their own performance based on empirical data. Rather than executing strictly defined static program commands, machine learning algorithms adopt an approach of constructing specific models to derive predictions or decisions based on input data. The term 'machine learning' may be used interchangeably with 'machine learning'.

[0043] Many machine learning algorithms have been developed to address how to classify data in machine learning. Representative examples include Decision Trees, Bayesian Networks, Support Vector Machines (SVMs), and Artificial Neural Networks (ANNs). A Decision Tree is an analytical method that performs classification and prediction by plotting decision rules in a tree structure. A Bayesian Network is a model that represents the probabilistic relationships (conditional independence) between multiple variables in a graph structure. Bayesian Networks are suitable for data mining through unsupervised learning. Support Vector Machines are supervised learning models for pattern recognition and data analysis, primarily used for classification and regression analysis. Artificial Neural Networks model the operating principles of biological neurons and the relationships between them; they are information processing systems in which multiple neurons, referred to as nodes or processing elements, are connected in a layered structure.

[0044] Artificial neural networks are models used in machine learning, serving as statistical learning algorithms inspired by biological neural networks (specifically the brain within the animal central nervous system). Specifically, an artificial neural network can refer to a model in which artificial neurons (nodes), forming a network through synaptic connections, change the strength of these connections through learning to possess problem-solving capabilities. The term artificial neural network may be used interchangeably with neural network.

[0045] An artificial neural network may include multiple layers, and each layer may include multiple neurons. Additionally, the artificial neural network may include synapses connecting the neurons. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights can be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as, for example, Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above. In this specification, the term 'layer' may be used interchangeably with the term 'layer'.

[0046] Artificial neural networks are classified into single-layer neural networks and multi-layer neural networks depending on the number of layers. A typical single-layer neural network consists of an input layer and an output layer. Additionally, a typical multi-layer neural network consists of an input layer, one or more hidden layers, and an output layer.

[0047] The input layer is a layer that receives external data, and the number of neurons in the input layer is equal to the number of input variables. The hidden layer is located between the input layer and the output layer, receives signals from the input layer, extracts features, and transmits them to the output layer. The output layer receives signals from the hidden layer and outputs an output value based on the received signals. Input signals between neurons are multiplied by their respective connection strengths (weights) and then summed; if this sum is greater than the neuron's threshold, the neuron is activated and outputs the value obtained through the activation function.

[0048] Meanwhile, a deep neural network containing multiple hidden layers between the input layer and the output layer can be a representative artificial neural network that implements deep learning, a type of machine learning technique. Meanwhile, the term 'deep learning' may be used interchangeably with the term 'deep learning'.

[0049] The machine learning workflow consists of a series of processes involving collecting data for learning and validation, modeling, and training the model, and may include the processes of collecting training data, checking and exploring data, data preprocessing and cleaning, modeling, and training.

[0050] In describing one embodiment of the present disclosure, for convenience of explanation, a method for determining the schedule of a naphtha cracking center (NCC) will be described as an example. However, the embodiments of the present disclosure are not limited to a method for determining the schedule of a naphtha cracking center, and can, of course, be applied to a method for determining the schedule of other processes or to a method for determining information other than a schedule. One embodiment of the present disclosure can be applied to a method for determining information about a process in multiple steps.

[0051] FIG. 1 is a schematic diagram of a product production process of a naphtha cracking center according to one embodiment of the present disclosure.

[0052] Referring to Fig. 1, at a Naphtha Cracking Center (NCC), naphtha, which is a gasoline fraction obtained from an atmospheric distillation unit of crude oil, is thermally cracked in a high-temperature cracking furnace, and then through processes such as rapid cooling, compression, and refining, ethylene, propylene, butylene, and BTX (benzene, toluene, xylene), which are basic raw materials for petrochemical products, can be produced.

[0053] In other words, naphtha can be converted into substances with high industrial utility, such as ethylene, propylene, benzene, toluene, and xylene, through steam cracking or thermal cracking. For example, ethylene serves as a raw material for making polyethylene and polystyrene, propylene serves as a raw material for making polypropylene, and butane or butylene can be used to make synthetic rubber. These substances serve as raw materials for the plastics processing, textile, rubber, paint, and detergent industries, and these raw materials can become final products such as daily necessities, adhesives, dyes, pesticides, pharmaceuticals, industrial products, and interior materials.

[0054] A naphtha cracking center is a core facility that produces petrochemical raw materials through a complex process and consists of a receiving stage for unloading naphtha, a mixing stage for blending naphtha, and a cracking stage for producing marketable products. More specifically, naphtha is initially transported from various geographically distributed refineries via vessels and unloaded into receiving tanks; various types of naphtha from these receiving tanks are then supplied to mixing tanks; and the naphtha blended in the mixing tanks is heated in a cracking furnace to produce marketable products of the desired quality. That is, the product production process of the naphtha cracking center may include a receiving process of storing naphtha supplied from one or more vessels (110) or companies (e.g., other oil companies) in one or more receiving tanks (120), a mixing process of transferring the naphtha from the receiving tanks (120) to a mixing tank (130) for a naphtha cracking process, and a cracking process of thermally cracking the naphtha supplied from the mixing tanks (130) at high temperature in a furnace (140). Here, the mixing tanks (130) may also be referred to as blending tanks or feed tanks.

[0055] In one embodiment, the product production process of the naphtha cracking center may further include a process of measuring the paraffin content of naphtha supplied from a ship (110) or company, and a process of measuring the paraffin content of naphtha stored in an receiving tank (120), a mixing tank (120), etc.

[0056] In one embodiment, the constraints may include a range for the paraffin content for each tank. For example, the paraffin content of the mixing tank (130) may be limited to a range of about 80 to 83% based on the total weight of the naphtha. Since naphtha has different properties depending on the country of origin or company, the naphtha stored in the receiving tank also has different properties, and the receiving process and the mixing process must be performed so that the paraffin content of the mixing tank (130) satisfies the range of the constraints.

[0057] Considering these constraints, determining the optimal schedule for the naphtha cracking center is crucial for profitability and efficiency. Generally, experts decide based on their experience and know-how which incoming tank to store naphtha, at what ratio to mix it into the mixing tank, and to what extent to heat it using which cracking furnace. However, relying on human experience to determine the schedule has limitations in predicting complex chemical reactions and actual results; results vary significantly depending on the level of the expert's experience and know-how; it is difficult to verify whether all constraints have been satisfied; and it is difficult to respond to sudden changes in circumstances.

[0058] Accordingly, the present disclosure aims to provide a method for determining optimal information (e.g., a schedule) using artificial intelligence. For example, multi-agent reinforcement learning may be utilized. A naphtha cracking center can be operated autonomously using multi-agent reinforcement learning, in which real-world constraints are overcome and each agent takes responsibility for a step and cooperates to achieve a common goal.

[0059] FIG. 2 is a diagram illustrating a reinforcement learning method using multiple agents according to one embodiment of the present disclosure.

[0060] Referring to FIG. 2, a naphtha cracking center can be operated according to an optimal schedule using scheduling information determined through reinforcement learning using multiple agents (210, 220, 230). A simulator using reinforcement learning can take actions (240, 250, 260) from the agents (210, 220, 230) and provide the next observation and reward (280) based on the current action. In one embodiment, each agent is responsible for a specific process and can cooperate with each other to achieve goals such as profit maximization while complying with real-world constraints. For example, there may be realistic constraints such as performing the transfer process from the receiving tank to the mixing tank for at least 8 hours, within a range that does not exceed the minimum naphtha storage capacity of the receiving tank and the maximum naphtha storage capacity of the mixing tank.

[0061] In one embodiment, each agent can determine the information necessary to generate scheduling information for naphtha cracking centers for a predetermined future period based on current information and various constraints, such as the inventory status of each tank, ship arrival plans, naphtha supply plans from other companies, and prices of naphtha and marketable products.

[0062] In one embodiment, each of the multiple agents may generate different results (e.g., durations) at different times. For example, the first agent (210) is an agent managing incoming shipments and may determine actions such as selecting an incoming tank to store naphtha when ships arrive irregularly and determining the amount to store in that tank, and the second agent (220) is an agent for mixing naphtha and may determine actions such as determining an incoming tank to bring naphtha to a mixing tank and determining the amount to bring from that tank when the level of a certain incoming tank reaches a threshold (e.g., 90% of the tank capacity). Additionally, the third agent (230) is an agent managing a cracking furnace and may determine actions such as receiving naphtha from a mixing tank and determining variables to operate the cracking furnace when the product inventory is below a predetermined amount. A virtual naphtha operating environment (270) can be created with actions (240, 250, 260) determined at different times. The simulation device can determine expected profits in a virtual naphtha operating environment (270) and determine a reward (280) based thereon. This reward can be delivered to multiple agents (210, 220, 230) and used by the agents to perform reinforcement learning. That is, multiple agents (210, 220, 230) can be trained using the same reward during reinforcement learning. However, it is also possible for multiple agents to be trained using different rewards.

[0063] In one embodiment, the reward (280) may be determined based on total revenue, facility operating costs, naphtha purchase costs, costs according to constraints, etc. For example, the reward may be determined by the following [Equation 1].

[0064] [Mathematical Formula 1]

[0065]

[0066] In [Equation 1], Constraints are constraint conditions, and w c is the weight per constraint, and Cost c can refer to the cost incurred per constraint. Accordingly, the more constraints are violated, the higher the Cost c The value can increase. For example, if the constraint includes the stability of the paraffin component—that is, the condition that the paraffin component must be maintained to a certain extent—the change in the paraffin component stored in the mixing tank can be used as variable c.

[0067] In addition, in [Equation 1], the profit can be calculated by taking into account facility operating costs (e.g., energy usage costs) and naphtha purchasing costs, and subtracting the estimated production costs of marketable products from the estimated revenue generated from the sale of naphtha. For example, the profit can be determined according to [Equation 2], which subtracts facility operating costs and naphtha purchasing costs from total revenue as follows.

[0068] [Mathematical Formula 2]

[0069]

[0070] In one embodiment, Revenue can be calculated as, for example, "CH4 production volume * CH4 product price + PSA OFF GAS production volume * PSA OFF GAS price + RC2 production volume * ethane product price + C3 LPG production volume * propane product price + ethylene production volume * C2 product price + propylene production volume * C3 product price + H2 (99%) production volume * 99% H2 product price + HRPG production volume * HRPG product price + PFO production volume * PFO product price + Raw C5 production volume * Raw C5 product price + (Mixed C4 production volume * Mixed C4 product price) + (RPG production volume * Raw C5 production volume) * RPG product price".

[0071] In one embodiment, energy usage can be calculated as, for example, "[(Naphtha input amount + C3 LPG input amount + C4 LPG input amount) * A + Mixed C4 production amount * B + (RPG production amount Raw C5 production amount) * C] * C3 LPG price / C3 LPG calorific value / 1000", where A is the average energy unit value of a naphtha cracking center plant, B is the average energy unit value of a BD plant, and C is the average energy unit value of a BTX plant (for example, a plant that produces aromatic products using pyrolysis gasoline produced from an ethylene plant).

[0072] In one embodiment, hh can be calculated as "Total naphtha feed input * naphtha price + C3 LPG input * C3 LPG price + C4 LPG input * C4 LPG price + RC2 input * ethane product price".

[0073] In one embodiment, when an optimal scheduling is determined through such reinforcement learning, the naphtha cracking center can be operated according to the generated optimal scheduling. For example, if multiple schedules are formed and provided to a user, the user can operate the naphtha cracking center based on one of them.

[0074] FIGS. 3a and 3b are drawings illustrating a reinforcement learning method using multiple agents according to one embodiment of the present disclosure.

[0075] Referring to FIG. 3a, an asynchronous multi-agent system is illustrated in which the start and duration times of the actions of each agent are different. For example, if the multi-agents consist of three first agents (310), second agents (320), and third agents (330) as in FIG. 3a, each agent determines a different action at a different time, and the determined action vector The changed state vector resulting from the actions being transmitted to the environment (340) and applied to the environment (340). It can be determined. Also, the action termination vector The reward generated as a result of applying actions to the environment (340) can be provided to the agents. In addition, the reward generated as a result of applying actions to the environment (340) It can also be provided to each agent. The actions of each agent can be determined asynchronously as shown in Fig. 3b.

[0076] In another embodiment, the actions of each agent may be transmitted to their respective environment (340) whenever an action is determined. For example, each action may be reflected in their respective environment (340) such that at the first time when the first agent (310) determines the first action, the first action is reflected in the environment (340), at the second time when the second action is determined, the second action is reflected in the environment (340), and at the third time when the second agent (310) determines the third action, the third action is reflected in the environment (340). However, since the reward is determined by assuming that the product is ultimately produced and sold, the reward may be determined and provided to each agent after the actions of the first agent (310), the second agent (320), and the third agent (330) are all reflected in the environment (340).

[0077] In one embodiment, the state, behavior, and reward of each agent may include the following information.

[0078] 1. First Agent

[0079] - Status: Naphtha receiving schedule (e.g., vessel schedule or receiving plans from other companies), current status of receiving tanks (e.g., naphtha inventory and properties per tank), constraints related to receiving tanks

[0080] - Action: Identifier of the receiving tank to store naphtha upon receipt, the amount of naphtha to be stored in the corresponding receiving tank, or the pipeline connection schedule (e.g., pipeline connection with Vessel A to Receiving Tank No. 1 from 2:00 PM to 10:00 PM)

[0081] - Reward: Returns considering whether constraints are satisfied

[0082] 2. Second agent

[0083] - Status: Stock and properties of each incoming tank, stock and properties of the mixing tank

[0084] - Action: Identifier of the receiving tank to hold at least some naphtha in the mixing tank, the amount of naphtha to be brought from that receiving tank to the mixing tank, or the pipe connection schedule (e.g., pipe connection from Receiving Tank No. 1 to the mixing tank from 8:00 AM to 2:00 PM)

[0085] Compensation: Returns considering whether constraints are satisfied

[0086] 3. Third agent

[0087] Status: Inventory and properties of mixing tanks, operating status by disassembly furnace

[0088] Action: Feed rate, COT, DS ratio, etc. to feed naphtha from the mixing tank into each cracking furnace

[0089] Compensation: Returns considering whether constraints are satisfied

[0090] However, this is merely one example, and it goes without saying that the state, behavior, and rewards of each agent can be adjusted differently.

[0091] In one embodiment, the naphtha cracking center scheduling system can acquire input information. For example, the input information can be acquired through user input based on a User Interface (UI).

[0092] In one embodiment, the input information may include constraints, naphtha receiving schedule information, tank inventory information, naphtha property information within the tank, mixing tank operation information, cracking furnace operation plan information, target production volume information for each specific product, raw material unit price information, product unit price information, etc. In another embodiment, some of the information, such as constraints, may be pre-set information, and in that case, may not be included in the input information because it has been set in advance.

[0093] In one embodiment, constraints may include physical constraints such as tank storage capacity criteria to be satisfied and the number of pipes that can be connected at once, stability constraints regarding stability, and operational constraints for complying with a set target production volume during a specific period (e.g., weekly or monthly). Additionally, naphtha receiving schedule information may include a vessel receiving schedule, tank information of another company, a naphtha receiving schedule for a specific future period, scheduled receiving date and time, receiving rate, receiving quantity, naphtha property information, and identification information based on the naphtha receiving method (e.g., vessel identifier, tank identifier of another company, pipe identifier by company, etc.).

[0094] In one embodiment, at least one of the input information may include information for a specific period or information at a specific point in time. For example, the tank inventory information and the naphtha properties information within the tank, respectively, may each include the naphtha inventory information of the corresponding tank at the time of the scheduling start and the naphtha properties information of the corresponding tank at the time of the scheduling start.

[0095] In one embodiment, the blending tank operation information may include one or more blending schedules, such as a blending start time, a blending end time, the name of the receiving tank to be blended, and a blending speed per receiving tank. For example, the blending schedule may include a recent blending schedule.

[0096] In one embodiment, the decomposition furnace operation plan information may include schedule information for each decomposition furnace for a future specified period. For example, the decomposition furnace operation plan information may include schedule information determined for each decomposition furnace for the next 30 days. In one embodiment, the decomposition furnaces may exist in various types. For example, if there are a first decomposition furnace, a second decomposition furnace, and a third decomposition furnace of different types, the decomposition furnace operation plan information may include schedule information determined for each of the first decomposition furnace, the second decomposition furnace, and the third decomposition furnace for the next 30 days. The decomposition furnace operation plan information may include a decomposition start time, a decomposition end time, an operating mode (or feed mode), decoking schedule information, COT (coil outlet temperature), coil outlet pressure, a predetermined speed (e.g., feed rate), a DS (Dilution Steam) ratio, etc.

[0097] In one embodiment, the target production volume information for a specific product may include the target production volume or rate for a specific product over a specific period in the future. For example, the daily target production volume of ethylene for the next 30 days, the daily target production volume of propylene for the next 30 days, etc., may be included in the target production volume information for a specific product.

[0098] In one embodiment, the raw material unit price information may include raw material unit price information at the time of information input, raw material unit price information for a specific period prior to the time of information input, and expected raw material unit price information for a specific period after the time of information input. For example, the raw material unit price information may include the expected daily price of raw materials for the next 30 days.

[0099] In one embodiment, the product unit price information may include unit price information at the time of inputting information for each product, unit price information for a specific period prior to the time of inputting information for each product, and expected unit price information for a specific period after the time of inputting information for each product. For example, the product unit price information may include the expected daily price of naphtha products for the next 30 days.

[0100] However, the above input information is merely an example and is not limited thereto, and various input information for scheduling the naphtha cracking center may be included.

[0101] In one embodiment, a naphtha cracking center scheduling system may determine receiving tank information using a first agent based on input information. In one embodiment, the first agent may be an agent trained using reinforcement learning. Based on input information including a ship receiving schedule, a naphtha receiving schedule from another company, real-time receiving tank inventory, naphtha properties information within the tank, cracking furnace operation plan information, etc., the first agent may determine a receiving tank to receive naphtha from at least one of the tanks of a ship and another company, and determine the amount, ratio, or schedule information of naphtha to be stored in the corresponding receiving tank. For example, based on the input information, the first agent may determine an identifier for at least one receiving tank to store naphtha among a plurality of receiving tanks, and determine naphtha receiving ratio or amount information for each tank corresponding to each identifier. Additionally, based on the input information, the first agent may determine naphtha receiving schedule information and information on the period for storing naphtha in the corresponding receiving tank for each tank corresponding to each identifier. The receiving schedule information may include date or time information for connecting the receiving tank to a vessel or another company's device via a pipe (e.g., receiving naphtha from Vessel B to Receiving Tank A from 2:00 PM to 6:00 PM). In one embodiment, pipes may be connected as a method to transfer naphtha from a vessel or another company's device to a receiving tank; however, since connecting pipes and performing other tasks is inconvenient if the schedule changes frequently, constraints such as a minimum connection time of n hours per pipe may exist. The first agent may determine the receiving tank information by taking these constraints into account.

[0102] In one embodiment, the naphtha cracking center scheduling system can obtain naphtha property information corresponding to each receiving tank after a predetermined amount of naphtha has been distributed to the receiving tank.

[0103] In one embodiment, all receiving tank information may be determined using a single first agent, receiving tank information may be determined using a different first agent for each receiving tank, or receiving tank information may be determined using a different first agent for each receiving tank group. That is, there may be one or more first agents.

[0104] In one embodiment, the naphtha cracking center scheduling system can determine mixing tank combination information using a second agent. In one embodiment, the second agent may be an agent trained using reinforcement learning. The second agent can determine the mixing tank combination information based on the inventory of each receiving tank, the properties of the naphtha stored in each receiving tank, etc.

[0105] In one embodiment, the blending tank combination information may include an identifier of at least one receiving tank among a plurality of receiving tanks to transfer naphtha into the blending tank, naphtha ratio information (or amount information) to be transferred to the blending tank for each of the at least one receiving tank identifiers, blending schedule information with the blending tank for each of the at least one receiving tanks, naphtha blending ratio information for each of the at least one receiving tanks, blending execution date information, etc. The naphtha blending ratio information may include ratio information or amount information of naphtha to be taken from each receiving tank.

[0106] In one embodiment, all mixing tank combination information may be determined using a single second agent, receiving tank information may be determined using a different second agent for each mixing tank, or receiving tank information may be determined using a different second agent for each mixing tank group. That is, there may be one or more second agents.

[0107] In one embodiment, the naphtha cracking center scheduling system may determine cracking furnace operation information using a third agent. In one embodiment, the third agent may be an agent trained using reinforcement learning. The third agent may determine cracking furnace operation information based on the inventory information of the mixing tank, the properties of the mixing tank, cracking furnace status information, etc. In one embodiment, the cracking furnace operation information may include cracking furnace mode information, cracking furnace identifier, feed rate, COT (Coil Outlet Temperature), DSR (Dilution Steam Ratio), heating time, cracking furnace operation schedule information, and one or more variables for cracking furnace operation.

[0108] In one embodiment, there may be multiple third agents. For example, different agents may be used for each cracking furnace mode. For example, the third agents include a 3-1 agent reinforced learning for the cracking furnace of mode A, a 3-2 agent reinforced learning for the cracking furnace of mode B, and a 3-3 agent reinforced learning for the cracking furnace of mode C, and the naphtha cracking center scheduling system may determine the operation information for the cracking furnace of mode A using the 3-1 agent for the cracking furnace of mode A, determine the operation information for the cracking furnace of mode B using the 3-2 agent for the cracking furnace of mode B, and determine the operation information for the cracking furnace of mode C using the 3-3 agent for the cracking furnace of mode C. In another embodiment, a single third agent may determine all cracking furnace operation information.

[0109] In one embodiment, the naphtha cracking center scheduling system can determine one or more scheduling information of a naphtha cracking center based on receiving tank information generated by a first agent, mixing tank combination information generated by a second agent, and cracking furnace operation information generated by a third agent. In one embodiment, the scheduling information may include receiving scheduling information, mixing scheduling information, cracking furnace scheduling information, expected production volume information, expected profit information, expected naphtha inventory information, expected properties information, constraint satisfaction test result information, scheduling graph, etc.

[0110] In one embodiment, the receiving scheduling information may include receiving schedule information for a predetermined period in the future. For example, the receiving scheduling information may include tank identification information to be received for the next two weeks, the start time of receiving for the tank, the end time of receiving for the tank, etc.

[0111] In one embodiment, the mixing scheduling information may include mixing schedule information for a predetermined period in the future. For example, the mixing scheduling information may include mixing tank identification information to be used for the next two weeks, the mixing start time of the mixing tank, the mixing end time of the mixing tank, and mixing speed information of the tank (e.g., mixing speed of Tank A: about 100 Ton / hour).

[0112] In one embodiment, the decomposition furnace scheduling information may include schedule information for each decomposition furnace for a predetermined period in the future. For example, the decomposition furnace identification information to be used for the next two weeks, the start time of decomposition for the corresponding decomposition furnace, the end time of decomposition for the corresponding decomposition furnace, the decomposition speed (e.g., target control speed determined by artificial intelligence, feed rate, etc.), COT, DS ratio, etc. may be included in the decomposition furnace scheduling information.

[0113] In one embodiment, the expected production volume information, expected revenue information, expected naphtha inventory information, expected properties information, etc., may also be expected information for a predetermined period in the future. For example, the expected production volume information may include daily expected production volume provided by product for the next two weeks. Additionally, the expected naphtha inventory information may include naphtha inventory or naphtha change information provided by tank for the next two weeks, and the expected properties information may include properties change information provided by tank for the next two weeks.

[0114] In one embodiment, the constraint satisfaction test result information may include evaluation information on how well the generated schedule satisfies the predetermined constraints.

[0115] Furthermore, the naphtha cracking center scheduling system can provide one or more scheduling information to the user through a UI / UX. For example, an overview of each of the above one or more scheduling information can be displayed and provided in the form of a graph or figure through the UI / UX, and summary information such as cumulative profit and constraint satisfaction can also be provided.

[0116] According to one embodiment of the present disclosure, a naphtha cracking center scheduling system may determine scheduling information for a naphtha cracking center using an asynchronous multi-agent system comprising a first agent, a second agent, and a third agent. For example, each agent may determine different information at different times.

[0117] FIG. 4 is a diagram showing the schedule of a naphtha cracking center according to one embodiment of the present disclosure.

[0118] Referring to FIG. 4, scheduling at a naphtha cracking center may mean planning a continuous and multi-stage process that converts raw naphtha into high-value-added products such as ethylene. For example, the process at a naphtha cracking center may include the following three interdependent stages.

[0119] 1) Unloading stage (410): A stage of unloading naphtha transported from a ship, etc., into a receiving tank.

[0120] 2) Blending step (430): A step of mixing naphtha from a selected receiving tank in a blending tank to achieve a target composition ratio.

[0121] 3) Cracking step (450): A step of producing ethylene, etc. by high-temperature decomposition of the mixed raw materials in a furnace.

[0122] These stages are so closely linked that the entire process can be halted if even one is out of schedule; therefore, a scheduling strategy that precisely coordinates the timing of each stage is essential.

[0123] To explain using the example in FIG. 4, when a new vessel arrives, an unloading step (410) must be performed. When the first vessel arrives, unloading from the first vessel into the first tank must be performed, and when the second vessel arrives, unloading from the second vessel into the second tank must be performed. Accordingly, the processor must schedule the unloading step (410) by anticipating the arrival time of the first vessel and the arrival time of the second vessel. In addition, based on the determined receiving tank information, the processor must determine the time to move from the receiving tank to the blending tank, the amount to be mixed, etc., and based on this blending information, determine the start time of operation of the disassembly furnace, the operating period, the temperature, etc.

[0124] That is, the processor must determine the timing and actions to be performed at each stage based on the fact that the unloading stage (410), blending stage (430), and cracking stage (450) are interconnected. However, these stages are so closely linked that if even one of them is out of schedule, the entire process may be halted; therefore, a scheduling strategy that precisely coordinates the timing of each stage is essential. Furthermore, each stage of the naphtha cracking center process is also affected by external factors such as operational constraints (e.g., equipment tolerances, safety standards), demand fluctuations, and ship arrival schedules, requiring long-term and robust scheduling. In one embodiment, when determining the schedule, the plant operating status (e.g., current tank inventory, equipment utilization rate, maintenance schedule, etc.), the ship arrival schedule (e.g., arrival time and shipment volume based on prior shipment data (shipment plan)), tank capacity (e.g., remaining space in each receiving tank and blending tank), raw material quality objectives (e.g., composition ratio after blending, sulfur content, evaporation point, etc.), and market conditions (e.g., external economic variables such as oil price, ethylene price, contract delivery date, etc.) must be comprehensively considered.

[0125] In one embodiment, the processor determines which receiving tank will receive the naphtha unloaded from each vessel in the unloading step (410), sets the tank combination and mixing ratio to satisfy the standard quality in the blending step (430), determines the blending sequence, and determines operating conditions such as the feed rate and coil outlet temperature per cracking furnace in the cracking step (450). The schedule determined in this way must be designed to maximize profitability while preventing conflicts between processes, and it is important that it possesses robust characteristics capable of responding to real-time fluctuations (vessel delays, equipment failures, etc.).

[0126] However, the mainstream approach has traditionally been to optimize by separating steps such as unloading, blending, and cracking. This step-by-step optimization method fails to adequately reflect the interactions between processes, making it difficult to fully resolve issues such as cascading delays and quality degradation that occur during actual operations. Furthermore, the NCC scheduling environment is extremely sensitive, where a single seemingly minor error can invalidate the entire schedule and ultimately halt operations. For instance, failing to start blending on time can cause the incoming tank to overflow, while an inappropriate blending ratio can generate off-spec feed, leading to a cascading impact on downstream processes. Due to these vulnerabilities, there is a problem in that reinforcement learning agents alone are insufficient to reliably generate a consistently valid and safe schedule.

[0127] In addition, reinforcement learning agents are generally trained based on a single scalar reward function, but in actual petrochemical processes, there may be a need to prioritize conflicting goals differently depending on the situation, such as profit maximization, process stability, and compliance with operating constraints. Since priorities fluctuate frequently depending on external factors such as market prices, delays in raw material arrival, and changes in equipment status, there is a problem in that it is difficult to sufficiently reflect these dynamic trade-offs using only a fixed reward function according to the method described with reference to FIGS. 1 to 3b.

[0128] One embodiment of the present disclosure aims to provide a method for determining information regarding a process consisting of multiple steps that solves these problems. Specifically, one embodiment of the present disclosure aims to provide a method for supporting an operator's decision-making in both long-term and short-term planning by generating, evaluating, and selecting multiple candidate schedules while simultaneously considering complex constraints.

[0129] In other words, the present disclosure relates to a system / method for generating and updating a schedule to reduce the risk of overall failure due to interdependencies between stages and to maximize process efficiency in a multi-stage process (e.g., manufacturing, logistics, IT pipeline, model learning pipeline, etc.). More specifically, the present disclosure relates to a technology in which artificial intelligence sequentially analyzes stage states, result information, etc., to branch and expand a pivot schedule, and combines reinforcement learning-based agent macro operations with a fitness estimator network (evaluator) to update the pivot schedule and determine the final schedule at each synchronization point.

[0130] FIG. 5 is a diagram illustrating a method for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0131] Referring to FIG. 5, at the initialization point (510), the processor can initialize the pivot schedule of each group into a blank macro operation sequence, reflecting the initial operating state of the NCC system.

[0132] In one embodiment, the processor may determine a planning period for creating a schedule. That is, the processor may determine a reference time for starting to create the schedule and a deadline for not extending the schedule further. For example, the processor may determine a planning period of three weeks from a specific point in time as the planning period for creating the schedule.

[0133] In one embodiment, the processor may create a plurality of groups. Referring to the example of FIG. 5, the processor may create group 1-1, group 1-2, group 2-1, group 3-1, etc.

[0134] In one embodiment, the processor may generate a plurality of groups corresponding to specific operating scenarios, including operating criteria, operating objectives, operating levels, constraints, etc. Operating objectives may include increased profitability, increased process stability, increased energy efficiency, compliance with quality specifications, savings in computing costs, reduced training time, reduced latency, reduced memory usage, and increased model accuracy. Operating criteria may include indicators representing process performance, such as profitability and process stability, and may be evaluated by a scalar goodness-of-fit function. Additionally, operating levels may include indicators representing the strictness of constraints and may be expressed as conservative, moderate, or stressed. Conservative may be selected when the range of constraints is narrow and safety is prioritized, while stressed may be selected under high-load operating requirements approaching equipment limits. Furthermore, constraints may include conditions regarding equipment operating limits, storage capacity, material property specifications, quality specifications, safety standards, memory capacity, the number of available server instances, network bandwidth, target inference latency, etc.

[0135] In one embodiment, the processor may configure multiple groups such that, as one moves toward lower groups, it satisfies higher (stricter) operating standards and operating levels. Since a schedule satisfying a higher level, i.e., a stricter level, will naturally also satisfy lower-level constraints, promising schedules can be transitioned between groups through these structured groups. For example, Group 1-2 may correspond to a group that must satisfy stricter operating standards and operating levels than Group 1-1, Group 2-1 than Group 1-2, and Group 3-1 than Group 2-1.

[0136] According to one embodiment of the present disclosure, such a hierarchical structure directly reflects real-world conditions in the model and can improve search efficiency through propagation from upper levels to lower levels. For example, in actual NCC operations, the intensity of constraints changes frequently depending on the state of the plant, market demand, etc., and the processor can select a stress level during a surge in demand and a maintenance level during stable operation.

[0137] - Pivot-based Branching operation

[0138] In one embodiment, the processor may duplicate the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups. For example, the processor may duplicate the pivot schedule (520) of the first-1 group in parallel to determine that all branch schedules of the first-1 group are identical to the pivot schedule (520) of the first-1 group. Additionally, the processor may duplicate the pivot schedule of the first-2 group in parallel to determine that all branch schedules of the first-2 group are identical to the pivot schedule of the first-2 group.

[0139] - Scenario-based Rollout Action

[0140] In one embodiment, the processor may perform a rollout process for each quarter schedule, taking into account the operational scenario of the corresponding group. A rollout process may refer to a stage of creating a complete schedule by deploying and simulating candidate schedules (or policies) for a predetermined planning period. This rollout process may be performed by an agent trained based on an artificial intelligence model.

[0141] That is, an agent trained based on an artificial intelligence model can extend a schedule for each of multiple quarter schedules by reflecting operational scenarios corresponding to each of multiple groups. For example, the processor can have the agent trained based on an artificial intelligence model continue to extend the schedule until the schedule for a predetermined planning period is fully completed. The processor can have the agent trained based on an artificial intelligence model construct a complete schedule by sequentially attaching macro actions to quarter schedules that are still empty or only partially defined. For example, the processor can complete a three-week production plan by having the agent trained based on an artificial intelligence model simulate actions in the unloading, blending, and cracking phases.

[0142] In one embodiment, the processor can generate a complete schedule in which all timesteps up to a predetermined planning period are filled, by taking as input the branch schedule up to the current point in time, the operating scenario of the group (objective, constraint, etc.), and macro action candidates (pre-learned policy, etc.).

[0143] In one embodiment, the processor may calculate a reward or penalty score using a fitness function of the group during or at the time of completion of the rollout process. Such reward or penalty score may be used when selecting a schedule to replicate during the synchronization evaluation phase.

[0144] - Synchronized Evaluation & Update operation

[0145] In one embodiment, the processor may update the group's pivot schedule by performing an evaluation at a predefined synchronization point (530, 540, 550). The synchronization point (530, 540, 550) may include a time when a common event occurs. For example, the processor may determine the time of a ship's arrival at the NCC as the synchronization point (530, 540, 550). Based on the fact that candidate schedules have progressed to the same operational time, the processor may select the superior schedule up to that time as the pivot schedule.

[0146] In one embodiment, the processor may evaluate the branch schedule of a first group and the branch schedule of another group with the same or higher constraint level as the first group at a first synchronization time using a fitness estimator network, and update the pivot schedule of the first group according to the evaluation result. In one embodiment, the fitness estimator network (e.g., a multilayer perceptron or a transformer) may receive various information as input and output a scalar fitness. For example, the input information may include a summary of the state of each branch schedule (e.g., inventory, capacity utilization, quality margin, safety margin, etc.), resource allocations, energy estimates, cost estimates, constraint violation penalties, scenario levels (e.g., maintenance, medium, stress, etc.), and operational purpose weight vectors. Learning is performed by regression on labels (e.g., realized revenue, total penalty, etc.) generated from a simulator or historical operating data, and may be fine-tuned by mini-batch based on the results of recent executions in the online phase. If the processor retains the pivot schedule of the first group at the first synchronization time when the score of the pivot schedule of the first group is the highest, and if the score of the pivot schedule of the first group is not the highest at the first synchronization time, it may replace the pivot schedule of the first group with the schedule having the highest score.

[0147] In one embodiment, at a predefined synchronization point (530, 540, 550), each group may evaluate (i) its own group's branch schedule and (ii) the branch schedule of another group having the same or higher operating level using its own fitness function. As a result of the evaluation, the best schedule may be updated as the new pivot schedule for that group.

[0148] - Multilayer recovery mechanism

[0149] In one embodiment, a recovery mechanism may be used. If a failed schedule exists, the processor may immediately replace it with a superior schedule within the group or a higher-level group and continue the search. That is, if any schedule is determined to have failed because it can no longer proceed throughout the rollout process, the recovery mechanism may be executed.

[0150] In one embodiment, if a specific branch schedule fails, the processor may replace the schedule with the schedule of highest fitness within the same group and continue the rollout process (Intra-group recovery). For example, the pivot schedule of Group 1-1 is replicated to become the first branch schedule and the second branch, and the schedule is expanded by reflecting the operational scenario of Group 1-1 for each of the first and second branch schedules. If, at some point, the first branch schedule fails, the processor evaluates the branch schedule of Group 1-1 and, if it determines that the second branch schedule has the highest fitness among the multiple branch schedules of Group 1-1, can replace the first branch schedule with the second branch schedule.

[0151] If all schedules belonging to a group fail, the processor may replace the entire group with a copy of the schedule with the highest fitness among other groups having the same or a higher level (inter-group recovery). For example, if the processor replicates the pivot schedule of Group 1-1 into the first and second branch schedules, and replicates the pivot schedule of Group 1-2 into the third and fourth branch schedules, and all branch schedules of Group 1-1 fail, the processor may determine the schedule with the highest fitness among the branch schedules of Group 1-2 as the pivot schedule of Group 1-1. In this case, Group 1-2 may be a group having requirements of a higher level (stricter level) than Group 1-1.

[0152] In one embodiment, if the schedule of all groups fails, the processor may perform a re-search starting from the pivot schedule of each group at the previous synchronization point.

[0153] This hierarchical recovery mechanism can increase the robustness of the plan by preventing premature search termination.

[0154] In one embodiment, after a predetermined planning period has ended, the processor may select one or more final schedules (560) by applying post-processing evaluation criteria to the completed pivot schedules.

[0155] For example, when the schedule is completed by a predetermined planning period, the processor may re-evaluate the completed pivot schedule based on final evaluation criteria. Unlike the fitness function used in the rollout process, the final evaluation criteria may include elements that can be precisely calculated only after the entire schedule is completed (e.g., cumulative profit, overall safety indicator, etc.). Based on the results of the re-evaluation, the processor may determine the final schedule. This final schedule (560) may include a schedule consisting of multiple stages and may be multiple. For example, the top three schedules may be provided to the user.

[0156] For example, when one embodiment of the present disclosure is used to determine the schedule of an NCC, the final schedule includes the operational schedule of the NCC, and the synchronization point may include the time of vessel arrival, the time of completion of a process step, the time of simulation time elapsed, etc. Additionally, the processor may perform a rollout process by extending the schedule based on a pre-learned reinforcement learning policy, by having at least one agent perform at least one of a naphtha receiving operation, a blending operation, and a cracking operation.

[0157] A candidate-based multi-scenario planning method according to one embodiment may structure candidates into multiple groups to overcome the limitations of existing single-population-based methods. According to one embodiment of the present disclosure, by maintaining various candidates, multiple pieces of information optimized under different objective functions (profit, stability, etc.) and constraint levels can be searched and preserved in parallel. Furthermore, according to one embodiment of the present disclosure, by utilizing a scenario hierarchy, a schedule satisfying high (strict) constraint levels automatically satisfies low constraint levels as well, thereby facilitating transitions between schedules.

[0158] In one embodiment, a structured candidate population can enable targeted search for various operational scenarios while maintaining the advantages of candidate-based techniques such as parallel exploration and escaping local optima. As a result, a set of information with superior quality and diversity compared to a single population explored uniformly can be generated. For example, in NCC scheduling, a schedule can be provided that simultaneously ensures comprehensiveness and robustness of production planning.

[0159] According to one embodiment of the present disclosure, by always maintaining a set of schedules capable of responding quickly to changes in priority or unexpected situations, robustness and adaptability that are difficult to achieve with an agent alone can be secured.

[0160] In addition, according to one embodiment of the present disclosure, the driver can immediately select the option most suitable for the current situation among a plurality of scenario-based schedules, thereby reducing the burden of decision-making.

[0161] In addition, according to one embodiment of the present disclosure, the risk of a single selection error disrupting the entire process can be significantly reduced.

[0162] In addition, according to one embodiment of the present disclosure, the schedule can be flexibly reconfigured even when abnormal events occur, such as market price fluctuations or equipment malfunctions.

[0163] In addition, according to one embodiment of the present disclosure, the limitations of a fixed reward structure can be compensated for by multi-objective, multi-constraint search according to one embodiment of the present disclosure while effectively utilizing behavior candidates generated by a reinforcement learning agent.

[0164] In other words, one embodiment of the present invention serves as a key element that bridges the potential of reinforcement learning-based scheduling with real-world plant requirements, and by simultaneously improving usability, safety, and efficiency, it can enable the introduction of various multi-stage processes into industrial sites.

[0165] Although an NCC scheduling method has been described as an example to explain one embodiment of the present disclosure, the method according to one embodiment of the present disclosure is not applicable only to NCC scheduling. For example, one embodiment of the present disclosure may also be used for data preprocessing operations, distributed model training operations, fine-tuning operations, inference operations, etc.

[0166] FIG. 6 is a flowchart illustrating a method for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0167] In operation 610, the processor may create a plurality of groups including a first group and a second group. In one embodiment, the plurality of groups may be created hierarchically. For example, the second group may be formed as a group having a higher level of operational scenario than the first group.

[0168] In one embodiment, each group can be mapped to an operational scenario defined by an objective function (e.g., profit, stability, latency minimization) and a constraint level (conservative, moderate, stress, etc.). This allows the processor to simultaneously explore different operational goals and constraints even within the same system.

[0169] In operation 630, the processor may duplicate pivot information for each of the multiple groups into multiple branch information for each of the multiple groups. In one embodiment, for each group, the processor may duplicate multiple branch information in parallel based on the pivot information of the group—a macro operation sequence that initially represents a blank or system initial state. For example, the processor may duplicate the pivot information of the first group into first branch information and second branch information, and duplicate the pivot information of the second group into third branch information and fourth branch information. This provides a basis for extensively exploring various alternative decision-making options at the same time.

[0170] In this case, the information may include schedules, logistics delivery routes, generator load curves, LLM training pipelines, etc. Additionally, macro actions may include unloading actions, blending actions, cracking actions, warehouse management actions, GPU training actions, validation actions, container deployment actions, etc.

[0171] In operation 650, the processor can expand information for each of the multiple branch information by reflecting the operational scenario corresponding to each of the multiple groups. That is, the processor can roll out information by reflecting the operational scenario of the corresponding group for each branch information. For example, in a Large Language Model (LM) pipeline, the processor can expand the steps in the order of data preprocessing, distributed learning, checkpoint saving, and inference deployment.

[0172] In one embodiment, if the processor cannot proceed with the scenario using the branch information of the first group, the first branch information may be replaced with the second branch information having the highest suitability among the plurality of branch information of the first group.

[0173] If there is no branch information to replace the multiple branch information in the first group, that is, if scenario progression is impossible with all the branch information in the first group, the processor may replace the first branch information with the third branch information having the highest suitability among the multiple branch information in the second group. In this case, the second group may be a group having a higher level of constraint than the first group.

[0174] If scenario progression is impossible with any of the information from all groups, including the first and second groups, the processor may perform a re-search based on the pivot information of each of the multiple groups at the previous synchronization point.

[0175] In one embodiment, the operation scenario may include operation objectives, constraints, etc. Operation objectives include increased profitability, increased process stability, increased energy efficiency, compliance with quality standards, reduced computing costs, reduced training time, reduced latency, reduced memory usage, increased model accuracy, etc., and constraints may include conditions regarding at least one of equipment operating limits, storage capacity, material property specifications, quality specifications, safety standards, memory capacity, the number of available server instances, network bandwidth, and target inference latency.

[0176] In operation 670, the processor may update the pivot information of the first group at the first synchronization point based on the expanded information. In one embodiment, when the expanded branch information reaches the first synchronization point, the processor may evaluate all branch information of the first group and branch information of other groups having the same or higher constraint level using a fitness function. Additionally, the processor may update the pivot information of the first group to the best branch information based on the result evaluated by the fitness function. The synchronization point may include, for example, common system-wide events such as ship arrival, facility turnaround, or completion of a learning epoch.

[0177] In operation 690, the processor may determine final information based on the updated pivot information of the first group. In one embodiment, after a predetermined planning period has ended, the processor may select one or more final information by applying post-processing evaluation criteria to the completed pivot information.

[0178] In one embodiment, when the time of plan completion is reached, the processor may re-evaluate the completed pivot information, including the updated pivot information of the first group, using a separate final evaluation criterion. Unlike the fitness function used during rollout, the final criterion may include elements that can be precisely calculated only at completion, such as cumulative revenue, total delay, and energy consumption. Information with the highest performance results from the evaluation may be selected as Final Info and transmitted to an operator UI (User Interface) or an automated execution module. In one embodiment, the Final Info may include information consisting of multiple steps.

[0179] FIG. 7 is a block diagram of a system for determining information about a process consisting of a plurality of steps according to one embodiment of the present disclosure.

[0180] Referring to FIG. 7, the system (700) (the system may be referred to as a server or device) may include a transceiver (710), memory (720), a database (730), and a processor (740). However, not all components shown in FIG. 7 are essential components of the system (700). The system (700) may be implemented with more components than those shown in FIG. 7, or with fewer components than those shown in FIG. 6. Furthermore, the transceiver (710), memory (720), and processor (740) may be implemented in the form of a single chip.

[0181] In one embodiment, the transceiver (710) may communicate with a terminal or other electronic device connected to the system (700) via wired or wireless connection. For example, the transceiver (710) may receive input information from a user terminal. In one embodiment, the input information may include naphtha receiving plan information, production target quantity for each product (e.g., ethylene production target quantity), constraints, status information, and scheduling start time. The naphtha receiving plan information may include the scheduled time of receiving, receiving speed, receiving quantity, naphtha properties information, and receiving type information (e.g., whether it is a ship or a tank of another company). The status information may include a predetermined mixing schedule and a predetermined cracking furnace schedule. The mixing schedule or cracking furnace schedule may include the start time, end time, tank name, and mixing or cracking speed for each tank, and the cracking furnace schedule may further include various variable information such as pressure and temperature information. Additionally, the status information may include naphtha inventory quantity and properties information for each tank at the time of scheduling start.

[0182] Various types of data, such as programs and files, such as applications, can be installed and stored in the memory (720). The processor (740) may access and use the data stored in the memory (720) or store new data in the memory (720). Additionally, one or more instructions may be stored in the memory (720). The processor (740) may execute one or more instructions stored in the memory.

[0183] The processor (740) controls the overall operation of the system (700) and may include at least one processor, such as a CPU, GPU, etc. The processor (740) may control other components included in the system (700) to perform operations for operating the system (700). For example, the processor (740) may determine incoming tank information using a first agent based on input information, determine mixing tank combination information using a second agent, and determine disassembly furnace operation information using a third agent.

[0184] In one embodiment, the processor (740) may generate one or more scheduling information for the naphtha cracking center based on the receiving tank information, the mixing tank combination information, and the cracking furnace operation information. Additionally, the transmitting and receiving unit (710) may transmit the scheduling information to the user terminal so that the scheduling information is displayed on the display of the user terminal. In one embodiment, the output scheduling information may include a receiving schedule for a future predetermined period (e.g., receiving start and end times, receiving tank identifier, etc.), a mixing schedule for a future predetermined period (e.g., mixing start and end times, mixing tank identifier, mixing speed, etc.), a cracking furnace schedule for a future predetermined period (e.g., cracking start and end times, target speed (feed rate) determined by an algorithm, COT, DS ratio, etc.), daily production volume and expected profit information by product for a future predetermined period, naphtha inventory quantity and change in properties for a future predetermined period, constraint check result information for the generated schedule, and a plot visualizing the generated schedule.

[0185] The process of generating and outputting scheduling information in this manner can be performed through the UI / UX of the user terminal. For example, when the processor (740) obtains input information entered by the user, it verifies whether there is sufficient data in the input information to generate output data, and if it is determined that the input information is valid, it can generate one or more scheduling information using an artificial intelligence scheduler based on the scheduling start date entered by the user. Additionally, the processor (740) can graph one or more scheduling information to provide information, and the user terminal can display this information in the form of a UI / UX.

[0186] The database (730) can store various training data for training a learning model. Additionally, the database (730) may store material information, phase information, simulation result information, etc., and in various embodiments, output data produced by the learning model may be stored. Although FIG. 7 is illustrated as including the database (730) in the system (700), the database (730) may be provided outside the device. In this case, the database (730) may be connected to the system (700) via wired or wireless connection.

[0187] Additionally, the learning model may be implemented outside the system (700) (e.g., cloud-based) or included inside the system (700).

[0188] One embodiment of the present disclosure may also be implemented in the form of a recording medium comprising computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include both computer storage media and communication media. A computer storage medium includes both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. A communication medium typically includes computer-readable instructions, data structures, or program modules and includes any information transmission medium.

[0189] The foregoing description of the present disclosure is for illustrative purposes only, and those skilled in the art will understand that modifications can be easily made to other specific forms without altering the technical spirit or essential features of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.

[0190] The scope of the present disclosure is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present disclosure.

Claims

1. In the system, One or more processors; and The system includes one or more memories that collectively store instructions that cause the system to perform operations when executed by the above-mentioned one or more processors, and the operations are: The operation of creating a plurality of groups including a first group and a second group; The operation of replicating the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups; For each of the plurality of branch schedules above, an operation to expand the schedule by having an agent trained based on an artificial intelligence model perform a macro operation reflecting an operation scenario corresponding to each of the plurality of groups; An operation of evaluating the branch schedule of the first group and the branch schedule of another group with the same or higher constraint level as the first group at the first synchronization time using a goodness-of-fit estimator network, and updating the pivot schedule of the first group according to the evaluation result; and A system comprising an operation to determine a final schedule based on the pivot schedule of the first group updated above.

2. In paragraph 1, the above final schedule is, A system comprising a schedule consisting of multiple stages.

3. In Paragraph 1, A system characterized in that the above-mentioned second group has a higher level of operational scenario than the above-mentioned first group.

4. In paragraph 1, the operation of replicating the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups is, The operation includes replicating the pivot schedule of the first group into a first branch schedule and a second branch schedule, and The operation of expanding the schedule by performing a macro operation reflecting the operation scenario for each of the above multiple branch schedules is, A system comprising the operation of replacing the first quarter schedule with the second quarter schedule when the first quarter schedule of the first group fails.

5. In paragraph 4, the above second quarter schedule is, A system determined by one or more processors as the schedule having the highest suitability among a plurality of branch schedules of the first group.

6. In paragraph 1, the operation of replicating the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups is, The operation of replicating the pivot schedule of the first group above into a first branch schedule and a second branch schedule; and The operation includes replicating the pivot schedule of the second group above into the third quarter schedule and the fourth quarter schedule, and The operation of expanding the schedule by performing a macro operation reflecting the operation scenario for each of the above multiple branch schedules is, An operation to determine that a plurality of branch schedules of the first group, including the first branch schedule and the second branch schedule, are all failures; and The method includes an operation of determining one of the branch schedules of the second group as the pivot schedule of the first group at the time when all of the multiple branch schedules of the first group fail, and The above second group is, A method characterized by being a group having a higher level of constraint than the first group above.

7. In paragraph 6, any one of the branch schedules of the second group determined by the pivot schedule of the first group is, A system determined by one or more processors as the schedule having the highest suitability among a plurality of branch schedules of the second group.

8. In Paragraph 1, A system comprising, when all schedules of the plurality of groups fail, an operation to perform a re-search starting from the pivot schedule of each of the plurality of groups at the previous synchronization point.

9. In paragraph 1, the operation of determining the final schedule is, A system comprising the operation of selecting one or more final schedules by applying post-processing evaluation criteria to completed pivot schedules after a predetermined planning period has ended.

10. In paragraph 1, the operation scenario corresponding to each of the plurality of groups is, A system including operational purposes and constraints.

11. In Paragraph 10, the above operational purpose is, It includes at least one of increased profitability, increased process stability, increased energy efficiency, compliance with quality standards, reduced computing costs, reduced training time, reduced latency, reduced memory usage, and increased model accuracy, and The above constraints are, A system comprising conditions for at least one of equipment operating limits, storage capacity, physical property specifications, quality specifications, safety standards, memory capacity, number of available server instances, network bandwidth, and target inference latency.

12. In paragraph 1, the above final schedule is, Including the operating schedule of the naphtha cracking center, The above first synchronization point is, A system comprising at least one of the time of arrival of a vessel, the time of completion of a process step, and the time of elapsed simulation time.

13. In Clause 12, the operation of expanding the schedule by performing a macro operation reflecting an operation scenario for each of the plurality of branch schedules is, A system comprising an operation to extend a schedule by having at least one agent perform at least one of a naphtha receiving operation, a blending operation, and a cracking operation based on a pre-learned reinforcement learning policy.

14. A method performed by one or more processors, The operation of creating a plurality of groups including a first group and a second group; The operation of replicating the pivot schedule of each of the plurality of groups into a plurality of branch schedules for each of the plurality of groups; For each of the plurality of branch schedules above, an operation to expand the schedule by having an agent trained based on an artificial intelligence model perform a macro operation reflecting an operation scenario corresponding to each of the plurality of groups; Based on the above extended schedule, at a first synchronization time, the branch schedule of the first group and the branch schedule of another group with the same or higher constraint level as the first group are evaluated by a goodness-of-fit estimator network, and the pivot schedule of the first group is updated according to the evaluation result; and A method comprising determining a final schedule based on the pivot schedule of the first group above.

15. A program stored on a computer-readable recording medium to execute the method of paragraph 14 on a computer.