Crude oil blending scheduling method and apparatus
By combining deep reinforcement learning and genetic algorithms, hyperparameters are automatically optimized to establish a crude oil blending and scheduling model. This solves the problems of large solution scale and poor reusability in existing technologies, achieving efficient crude oil blending and scheduling and improving the production stability and resource utilization efficiency of refineries.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- RICHFIT INFORMATION TECH
- Filing Date
- 2024-12-23
- Publication Date
- 2026-05-07
Smart Images

Figure CN2024141323_07052026_PF_FP_ABST
Abstract
Description
A method and apparatus for blending and scheduling crude oil
[0001] Related applications
[0002] This application claims priority to Chinese Patent Application No. 202411537638.9, filed on October 30, 2024, and incorporates the entire contents of the aforementioned patent application as part of this application. Technical Field
[0003] This application relates to the field of crude oil processing technology, specifically to a crude oil blending and scheduling method and apparatus. Background Technology
[0004] Crude oil blending is the process of mixing crude oils of different properties in a certain proportion to obtain a blended crude oil that meets specific requirements. With companies increasingly purchasing low-priced and heavy oils, ensuring the stability and efficiency of crude oil blending processing has become an unavoidable issue, thus placing higher demands on crude oil blending scheduling.
[0005] Currently, crude oil from different sources exhibits significant differences in properties, such as density and viscosity, which directly impact refining processes and product quality. Simultaneously, with the continued growth in global energy demand and the gradual depletion of oil resources, improving crude oil utilization efficiency is crucial. Crude oil blending allows for the mixing and refining of crude oils of different qualities to achieve property homogenization, forming stable blended crude oils. This ensures the stable operation of refinery distillation units, maximizes the utilization of all crude oil resources, and effectively reduces raw material costs. However, current crude oil dispatching schemes often employ rigorous mathematical programming methods, which suffer from problems such as excessively large solution scales, infeasibility, and poor reusability. These methods are ill-suited to the dynamic nature of actual production and fail to effectively fulfill the role of blending and dispatching. Therefore, there is an urgent need to research novel crude oil blending and dispatching technologies to more effectively address the challenges of crude oil property differences and resource utilization. Summary of the Invention
[0006] To address the problems in the prior art, this application provides a crude oil blending and scheduling method and apparatus.
[0007] Firstly, this application proposes a crude oil blending and scheduling method, including:
[0008] Obtain data on the status of crude oil blending resources;
[0009] Based on the crude oil blending resource status data and the crude oil blending scheduling model, a crude oil blending scheduling scheme is obtained; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
[0010] Secondly, this application provides a crude oil blending and scheduling device, comprising:
[0011] The first acquisition module is used to acquire crude oil blending resource status data;
[0012] The prediction module is used to obtain a crude oil blending scheduling scheme based on the crude oil blending resource status data and the crude oil blending scheduling model; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
[0013] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the program to implement the crude oil blending and scheduling method described in any of the above embodiments.
[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the crude oil blending and scheduling method described in any of the above embodiments.
[0015] Fifthly, this application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements the crude oil blending and scheduling method described in any of the above embodiments.
[0016] The crude oil blending and scheduling method and apparatus provided in this application acquire crude oil blending resource status data; based on the crude oil blending resource status data and the crude oil blending and scheduling model, a crude oil blending and scheduling scheme is obtained; the crude oil blending and scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training, thereby improving the efficiency of obtaining the crude oil blending and scheduling scheme. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0018] Figure 1 is a flowchart illustrating the crude oil blending and scheduling method provided in the first embodiment of this application.
[0019] Figure 2 is a flowchart illustrating the crude oil blending and scheduling method provided in the second embodiment of this application.
[0020] Figure 3 is a flowchart illustrating the crude oil blending and scheduling method provided in the third embodiment of this application.
[0021] Figure 4 is a flowchart illustrating the crude oil blending and scheduling method provided in the fourth embodiment of this application.
[0022] Figure 5 is a flowchart illustrating the crude oil blending and scheduling method provided in the fifth embodiment of this application.
[0023] Figure 6 is a flowchart of the establishment of the crude oil blending and scheduling model provided in the sixth embodiment of this application.
[0024] Figure 7 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the seventh embodiment of this application.
[0025] Figure 8 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the eighth embodiment of this application.
[0026] Figure 9 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the ninth embodiment of this application.
[0027] Figure 10 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the tenth embodiment of this application.
[0028] Figure 11 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the eleventh embodiment of this application.
[0029] Figure 12 is a schematic diagram of the physical structure of the electronic device provided in the twelfth embodiment of this application. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit this application. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.
[0031] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0032] To facilitate understanding of the technical solution provided in this application, the relevant content of the technical solution in this application will be explained below.
[0033] This application utilizes deep reinforcement learning as a tool to solve the crude oil blending and scheduling problem through a data-driven approach. Unlike rigorous mathematical programming methods (mechanism-based approaches), the data-driven method models the problem using data, freeing it from the constraints of precise formulas. This allows for broader applicability and solution stability while still meeting accuracy requirements. Deep reinforcement learning possesses powerful feature learning capabilities, enabling it to automatically learn useful features from raw data.
[0034] Research on refining and chemical production scheduling optimization based on deep reinforcement learning can leverage its advantages in decision-making and optimization problems to provide more flexible and intelligent scheduling schemes. However, due to the large number of parameters involved in the training process of deep reinforcement learning, a significant amount of time and effort is required for parameter tuning and optimization. Therefore, this application integrates the powerful global optimal solution search capability of genetic algorithms with deep reinforcement learning algorithms to form a novel optimization method for crude oil blending scheduling. The genetic algorithm automatically finds the optimal hyperparameter combination of deep reinforcement learning, thereby obtaining the optimal crude oil blending scheduling decision scheme.
[0035] The following describes the specific implementation process of the crude oil blending and scheduling method provided in this application embodiment, taking the server as the execution subject as an example. The execution subject of the crude oil blending and scheduling method provided in this application embodiment is not limited to the server.
[0036] Figure 1 is a flowchart illustrating the crude oil blending and scheduling method provided in the first embodiment of this application. As shown in Figure 1, the crude oil blending and scheduling method provided in this embodiment includes:
[0037] S101. Obtain crude oil blending resource status data;
[0038] Specifically, the server can obtain crude oil blending resource status data, which can serve as the basis for decision-making regarding crude oil blending production over a future period. This data includes, but is not limited to, the processing capacity of the atmospheric and vacuum distillation unit, the current processing plan, the tanker's arrival time, its estimated departure time, the amount and type of crude oil loaded in the tanker's oil tanks, the initial inventory of a certain crude oil in a certain terminal storage tank, and the initial inventory of a certain crude oil in a certain plant area storage tank. These parameters are set according to actual needs and are not limited in this embodiment. For example, the processing capacity of the atmospheric and vacuum distillation unit might be 1000 tons per hour. The processing plan might be, for example, a ratio of four crude oils (a:b:c:d) of 10:20:50:20, with a total processing volume of 80,000 tons. The processing plan is preset.
[0039] For example, the crude oil blending production of refinery A involves the dispatch of crude oil between tanker a, port b, and refinery A's plant area. The processing capacity of the atmospheric and vacuum distillation unit, the current processing plan, the arrival time of tanker a at port b, the expected departure time of tanker a from port b, the amount and type of crude oil loaded in the i-th oil tank of tanker a, the current inventory and type of crude oil in the j-th crude oil storage tank of port b, and the original inventory and type of crude oil in the k-th crude oil storage tank of refinery A's plant area can be obtained as the crude oil blending resource status data of refinery A.
[0040] S102. Based on the crude oil blending resource status data and the crude oil blending scheduling model, a crude oil blending scheduling scheme is obtained; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
[0041] Specifically, the server inputs crude oil blending resource status data into the crude oil blending scheduling model and can output a crude oil blending scheduling plan. Crude oil blending scheduling is a sequential decision-making process, and the plan can include the initial state and action decisions for each cycle. The initial state for each cycle can include the state of the tanker, the state of the tanker's oil tanks, the state of the terminal crude oil storage tanks, and the state of the plant's crude oil storage tanks. The action decisions for each cycle can include whether a tanker's oil tank is unloading oil from a storage tank, whether a storage tank is transferring oil to a loading tank, whether a loading tank is supplying oil to the atmospheric and vacuum distillation unit, whether the atmospheric and vacuum distillation unit is switching feeds, the amount of oil unloaded from a tanker to a storage tank, the amount of oil transferred from a storage tank to a loading tank, and the amount of oil transferred from a loading tank to the atmospheric and vacuum distillation unit. The crude oil blending scheduling plan can be used for crude oil scheduling in crude oil blending production.
[0042] In this embodiment, the crude oil blending and scheduling model is obtained through training based on crude oil blending sample data and corresponding actual scheduling actions. The crude oil blending sample data can be obtained from historical crude oil blending resource status data, and the actual scheduling actions corresponding to the crude oil blending sample data serve as labels for the crude oil blending sample data. During the training process of the crude oil blending and scheduling model, the hyperparameters of the model are optimized. The optimized hyperparameters are then used for model training. That is, this application does not pre-specify the model's hyperparameters during training, nor does it require manual adjustment of the model's hyperparameters by technicians. Instead, the adjustment of hyperparameters is integrated into the training process of the crude oil blending and scheduling model, automatically optimizing and adjusting the model's hyperparameters during training. Because the hyperparameters can be automatically optimized and adjusted, the training efficiency of the model is improved compared to the manual adjustment of hyperparameters in the prior art.
[0043] The crude oil blending and scheduling method provided in this application obtains crude oil blending resource status data; based on the crude oil blending resource status data and the crude oil blending and scheduling model, it obtains a crude oil blending and scheduling scheme; wherein, the crude oil blending and scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the training process of the crude oil blending and scheduling model, the hyperparameters of the model are optimized and the model training is accelerated, and the crude oil blending and scheduling model corresponding to the optimized hyperparameters is used for model training, which can improve the efficiency of obtaining the crude oil blending and scheduling scheme.
[0044] Figure 2 is a flowchart illustrating the crude oil blending and scheduling method provided in the second embodiment of this application. As shown in Figure 2, based on the above embodiments, the crude oil blending and scheduling model is further trained based on crude oil blending sample data and corresponding actual scheduling actions, including:
[0045] S201. Obtain crude oil blending sample data and corresponding actual scheduling actions;
[0046] Specifically, historical crude oil blending resource status data and corresponding crude oil blending scheduling actions can be collected. A preset number of historical crude oil blending resource status data can be obtained as crude oil blending sample data, and the crude oil blending scheduling actions corresponding to the preset number of historical crude oil blending resource status data can be obtained as the actual scheduling actions corresponding to the crude oil blending sample data.
[0047] S202. Initialize the original model and obtain the range of hyperparameter values;
[0048] Specifically, the server can randomly initialize the original model to obtain its parameters. The server can also obtain the range of hyperparameter values. In machine learning, hyperparameters typically refer to parameters that need to be set before model training; however, in this application, hyperparameters are adjusted during model training. These hyperparameters may include the learning rate, number of network layers, number of neurons, discount factor, and experience replay pool size, and are set according to actual needs; this application does not limit the specific settings. The range of hyperparameter values is set based on practical experience, and this application does not limit the specific settings. The original model can be a reinforcement learning model, such as a Deep Deterministic Policy Gradient (DDPG) model.
[0049] S203. Generate each set of hyperparameters based on their value range;
[0050] Specifically, the server randomly generates each set of hyperparameters based on their value range. Each set of hyperparameters includes the learning rate, number of network layers, number of neurons, discount factor, and experience replay pool size, etc., and is set according to actual needs; this embodiment does not limit this setting. The number of hyperparameter sets is selected based on practical experience; this embodiment does not limit this setting either.
[0051] S204. Based on the crude oil blending sample data, the corresponding actual scheduling actions, each group of hyperparameters, the fitness function, and the original model, a crude oil blending scheduling model is trained; where the fitness function is preset.
[0052] Specifically, in each training round, the hyperparameters of each group are screened and optimized to obtain optimized hyperparameters. Based on crude oil blending sample data and corresponding actual scheduling actions, the crude oil blending scheduling model corresponding to each optimized hyperparameter group is trained. Multiple training rounds are conducted until the training termination condition is met, resulting in a trained crude oil blending scheduling model to determine the crude oil blending scheduling scheme. The fitness function is preset and used to screen the hyperparameters.
[0053] This application incorporates hyperparameter optimization into the model training process, improving the efficiency and reliability of hyperparameter acquisition.
[0054] Figure 3 is a flowchart illustrating the crude oil blending and scheduling method provided in the third embodiment of this application. As shown in Figure 3, based on the above embodiments, the crude oil blending and scheduling model is further trained using crude oil blending sample data, corresponding actual scheduling actions, hyperparameters, fitness functions, and the original model, including:
[0055] S301. Based on the sample data of crude oil blending in this round and the model to be trained for each set of hyperparameters, obtain the prediction results of the crude oil blending scheduling action for this round corresponding to each set of hyperparameters.
[0056] Specifically, the server inputs the current round of crude oil blending sample data into the training model (the crude oil blending and scheduling model to be trained) corresponding to each set of hyperparameters, and can input the prediction results of the current round of crude oil blending and scheduling actions corresponding to each set of hyperparameters. In this embodiment, for the first round of training, the training model is the original model after initialization; from the second round of training onwards, the training model refers to the model obtained after the previous round of training. The training model corresponding to a set of hyperparameters is the model obtained by setting a set of hyperparameters as the hyperparameters of the training model.
[0057] S302. Based on the prediction results and fitness function of the current round of crude oil blending and scheduling actions corresponding to each set of hyperparameters, obtain the fitness value of the current round corresponding to each set of hyperparameters.
[0058] Specifically, based on the prediction results of the current round of crude oil blending and scheduling actions corresponding to each set of hyperparameters and the fitness function, the server can obtain the fitness value for each set of hyperparameters in the current round. In this embodiment, the fitness function can be constructed based on minimizing the cumulative cost.
[0059] S303. Optimize hyperparameters based on each group of hyperparameters and the fitness value of each group of hyperparameters in the current round to obtain the optimized hyperparameters of each group.
[0060] Specifically, the server optimizes the hyperparameters based on each set of hyperparameters and the corresponding fitness value in the current round, obtaining the optimized hyperparameters for each set. In this embodiment, a genetic algorithm can be used to optimize each set of hyperparameters, and the fitness function can be used as the fitness function of the genetic algorithm.
[0061] S304. Based on the sample data of crude oil blending in this round and the corresponding actual scheduling actions, conduct a round of model training for the training models corresponding to each set of optimized hyperparameters.
[0062] Specifically, for each set of optimized hyperparameters corresponding to the training model, a round of model training is performed based on the current round of crude oil blending sample data and the corresponding actual scheduling actions, updating the model parameters. The objective of model training is to maximize the cumulative reward, i.e., minimize the cumulative cost.
[0063] Figure 4 is a flowchart illustrating the crude oil blending and scheduling method provided in the fourth embodiment of this application. As shown in Figure 4, based on the above embodiments, further, hyperparameter optimization is performed according to each set of hyperparameters and the current-round fitness corresponding to each set of hyperparameters, resulting in optimized sets of hyperparameters including:
[0064] S401. Based on the fitness value and screening conditions corresponding to each group of hyperparameters in this round, obtain the retained hyperparameters for each group;
[0065] Specifically, the server filters each group of hyperparameters based on its current fitness value and the filtering criteria, obtaining the hyperparameter groups that meet the filtering criteria, which are then retained as the hyperparameters for each group. The filtering criteria are preset.
[0066] For example, the filtering criteria are the preset number of hyperparameters with the smallest fitness value.
[0067] S402. Perform cross operations based on the retained hyperparameters of each group and the fitness values corresponding to the retained hyperparameters of each group to obtain the offspring hyperparameters of each group.
[0068] Specifically, the server performs cross operations based on the retained hyperparameters of each group and the fitness values corresponding to the retained hyperparameters of each group to generate new hyperparameters and obtain the offspring hyperparameters of each group.
[0069] For example, the crossover operation selects the adaptive arithmetic crossover method. For every two sets of retained hyperparameters, a new set of hyperparameters can be generated as a set of child hyperparameters. A set of child hyperparameters can be obtained according to the formula child = λ·Parent1 + (1-λ)·Parent2, where Parent1 and Parent2 represent two different sets of retained hyperparameters, and λ is a random number between [0,1].
[0070] Each group of retained hyperparameters has a corresponding adaptive crossover probability, which can be calculated using the following formula:
[0071] Where P represents the adaptive crossover probability corresponding to a set of retained hyperparameters, P c Let f represent the basic crossover probability, and let f represent the fitness values corresponding to a set of retained hyperparameters. f represents the average fitness value. min P represents the minimum fitness value, k1 and k2 represent positive proportionality constants, and P represents the minimum fitness value. min P represents the minimum crossover probability. max This represents the maximum crossover probability. The base crossover probability P c The minimum crossover probability P min The maximum value of the crossover probability P max The settings should be configured according to actual needs; this application's embodiments do not impose limitations. Minimum fitness value f min The settings should be configured according to actual needs; this application's embodiments do not impose limitations. Average fitness The fitness value can be obtained by calculating the average of the fitness values corresponding to the retained hyperparameters of each group.
[0072] If the adaptive value f corresponding to a set of retained hyperparameters is higher than the average fitness This will reduce the corresponding adaptive crossover probability; if the adaptive value f corresponding to a set of retained hyperparameters is lower than the average fitness... But higher than the minimum fitness f min This will increase the corresponding adaptive crossover probability; if the adaptive value f corresponding to a set of retained hyperparameters is equal to the average fitness... Then the corresponding adaptive crossover probability is equal to the basic crossover probability P. c .
[0073] After the crossover operation, check whether the hyperparameters of the newly generated child sets exceed the predefined parameter range. If they do, take the corresponding boundary values and adjust their values to the predefined parameter range.
[0074] S403. Perform mutation operations on the hyperparameters of each group of offspring to obtain the optimized hyperparameters of each group.
[0075] Specifically, the server performs mutation operations on each group of offspring hyperparameters to obtain new hyperparameter groups, and uses the mutated offspring hyperparameters as the optimized hyperparameter groups.
[0076] Similar to crossover, mutation can introduce an error integral term to adaptively adjust probabilities. Multiplying the fitness value error by a ratio k yields the adaptive probability value. As the fitness value of an individual becomes more stable, the error decreases, the mutation amplitude decreases, and the mutation probability decreases. This protects the acquired high-quality genes from being destroyed and allows for more refined exploration near the optimal solution. The specific calculation formula is as follows:
[0077] The proportional term directly reflects the difference between the current fitness value and the optimal fitness value. The proportionality coefficient k p This determines the extent to which the error affects the probability of mutation.
[0078] The integral term reflects the cumulative effect of fitness error over time.
[0079] The total mutation probability is:
[0080] Where f represents the fitness value corresponding to a set of offspring hyperparameters, and P m It is the basic mutation probability, k p k represents the proportionality coefficient. i E represents the integral coefficient. cumuThis represents the cumulative error. The cumulative error needs to be maintained periodically to prevent numerical overflow. The specific process of obtaining f is similar to the calculation process of the fitness value of the parent corresponding to f, and will not be elaborated here.
[0081] After the mutation operation, check whether the newly generated offspring individuals exceed the predefined parameter range. If they do, take the corresponding boundary value and adjust their values to the predefined parameter range.
[0082] Based on the above embodiments, the further filtering condition includes that if the deviation between the fitness value of a set of hyperparameters in this round and the optimal fitness value is less than a threshold, then a set of hyperparameters is retained as a set of hyperparameters.
[0083] Specifically, the deviation between the fitness value of a set of hyperparameters in the current round and the optimal fitness value is calculated. The absolute value of this deviation is compared with a threshold. If the absolute value of the deviation is less than the threshold, then the set of hyperparameters is retained. If the absolute value of the deviation is greater than or equal to the threshold, then the set of hyperparameters is not retained.
[0084] For example, the deviation between the fitness value of the current round and the optimal fitness value corresponding to a set of hyperparameters can be expressed by |ff. best | / f best The calculation yields f, which represents the fitness value corresponding to a set of hyperparameters. best This represents the optimal fitness value.
[0085] Based on the above embodiments, the fitness function is further defined as the sum of the tanker's waiting costs at sea, unloading costs, tank inventory costs, and costs incurred in switching the feed to the atmospheric and vacuum distillation unit.
[0086] Specifically, the fitness function is fitness = C V +C uload +C inv +C set , where C V C represents the waiting costs for oil tankers at sea. uload C represents the cost of unloading oil. inv C represents the inventory cost of the can. set This indicates the cost incurred during the feed switching of the atmospheric and vacuum distillation unit.
[0087] Waiting costs for oil tankers at sea can be obtained by multiplying the waiting time by the waiting cost per unit time. Unloading costs can be obtained by multiplying the unloaded volume by the unloading cost per unit volume. Tank storage costs can be obtained by multiplying the amount of crude oil stored by the storage cost per unit of crude oil. Costs incurred during feed switching at the atmospheric and vacuum distillation unit are constant and set according to actual conditions.
[0088] Figure 5 is a flowchart illustrating the crude oil blending and scheduling method provided in the fifth embodiment of this application. As shown in Figure 5, based on the above embodiments, the crude oil blending and scheduling method provided in this application further includes:
[0089] S501. Conduct actual crude oil blending and scheduling according to the crude oil blending and scheduling plan.
[0090] Specifically, after obtaining the crude oil blending and scheduling plan, actual crude oil blending and scheduling can be carried out according to the plan. During the crude oil blending and scheduling process, the crude oil blending resource status data will change.
[0091] S502. If it is determined that the error between the current actual crude oil blending resource status data and the predicted current crude oil blending resource status data exceeds the error threshold, then a new crude oil blending scheduling plan is obtained based on the current actual crude oil blending resource status data and the crude oil blending scheduling model.
[0092] Specifically, the server can obtain the current actual crude oil blending resource status data. Based on the crude oil blending scheduling scheme, it can obtain the predicted current crude oil blending resource status data. By comparing the current actual crude oil blending resource status data with the predicted current crude oil blending resource status data, the error between the two can be obtained. This error is then compared with an error threshold. If the error exceeds the threshold, it indicates a significant difference between the predicted and actual crude oil blending resource status data. Continuing to use the crude oil blending scheduling scheme for actual crude oil blending will result in a worse scheduling effect. The current actual crude oil blending resource status data can be input into the crude oil blending scheduling model to obtain a new crude oil blending scheduling scheme, and then the actual crude oil blending scheduling can be performed based on this newly obtained scheme. The error between the current actual and predicted crude oil blending resource status data can be expressed as standard error.
[0093] The following uses the crude oil blending and scheduling of refinery B as an example to illustrate the specific implementation process of the crude oil blending and scheduling method provided in this application. The crude oil blending and scheduling of refinery B involves the scheduling of crude oil from tankers, terminals, and inter-refinery areas.
[0094] To perform crude oil blending and scheduling at Refinery B, a crude oil blending and scheduling model needs to be established. This embodiment of the application establishes a crude oil blending and scheduling model based on DDPG and a genetic algorithm. The model's hyperparameters include the learning rate, number of network layers, number of neurons, discount factor, and experience replay pool size. The fitness function is fitness = C. V +C uload +C inv +C set .
[0095] Figure 6 is a flowchart of the establishment of the crude oil blending and scheduling model provided in the sixth embodiment of this application. As shown in Figure 6, the steps for establishing the crude oil blending and scheduling model are as follows:
[0096] Step 1: Acquire Training Data. Training data includes crude oil blending sample data and corresponding actual scheduling actions. Crude oil blending sample data includes tanker arrival times for each time period, estimated departure times for tankers, the amount of crude oil loaded in the tanker's oil tanks, the initial inventory of type k crude oil in the crude oil storage tank at terminal i, and the initial inventory of type k crude oil in the crude oil storage tank (loading tank) at plant j. Actual scheduling actions include, for each time period, whether oil is unloaded from oil tank p to crude oil storage tank i at terminal i; whether oil is transferred from crude oil storage tank i to loading tank j at plant j; whether loading tank j to atmospheric and vacuum distillation unit l; whether feed switching occurs at atmospheric and vacuum distillation unit l; the amount of oil unloaded from oil tank p to crude oil storage tank i at terminal i; the amount of oil transferred from crude oil storage tank i to loading tank j at plant j; and the amount of oil transferred from loading tank j to atmospheric and vacuum distillation unit l.
[0097] Step 2: Initialize the DDPG model. Initialize the model parameters of the DDPG model.
[0098] Step 3: Obtain the hyperparameters for each set. Obtain the value ranges for the learning rate, number of network layers, number of neurons, discount factor, and experience replay pool size. The learning rate range is [0.0001, 0.1], the number of network layers range is [1, 50], the number of neurons range is [50, 500], the discount factor range is [0.1, 0.9], and the experience replay pool size range is [10000, 100000]. During the first round of model training, a set number of hyperparameter sets are randomly generated based on the value ranges of each hyperparameter. Each set of hyperparameters includes specific values for the learning rate, number of network layers, number of neurons, discount factor, and experience replay pool size. Starting from the second round of model training, obtain the optimized sets of hyperparameters.
[0099] Step 4: Calculate the fitness value corresponding to each set of hyperparameters. Set each set of hyperparameters as the hyperparameters of the model to be trained to obtain the model to be trained for each set of hyperparameters. Using the current round of crude oil blending sample data and the model to be trained for each set of hyperparameters, obtain the prediction result of the current round of crude oil blending scheduling actions corresponding to each set of hyperparameters. Based on the prediction result of the current round of crude oil blending scheduling actions corresponding to each set of hyperparameters and the fitness function, obtain the fitness value for each set of hyperparameters in this round. For the first round of training, the model to be trained is the original model after initialization; from the second round of training onwards, the model to be trained refers to the model obtained after the previous round of training.
[0100] Step 5: Determine whether the training is completed. The training stop condition in this application is that if, after k consecutive trainings, the absolute value of the difference between the fitness value corresponding to a set of hyperparameters and the average of the fitness values corresponding to each set of hyperparameters is less than the set threshold s, that is, the training end condition is |f i -f avg |<s, then the training ends and proceeds to Step 10; otherwise, the training has not ended and proceeds to Step 6. Here, f i represents the fitness value corresponding to the i-th set of hyperparameters obtained after the k-th training, and f avg represents the average of the fitness values corresponding to each set of hyperparameters obtained after the k-th training. s can take 0.01 or 0.1, and k is selected according to the actual situation.
[0101] In addition, an iteration threshold can be set. For example, the iteration threshold is equal to 5000. If the training still cannot meet the training end condition after 5000 iterations, the training also ends.
[0102] Step 6: Select each set of parent hyperparameters. Calculate the deviation between the fitness value corresponding to each set of hyperparameters in this round and the optimal fitness value. If the above deviation is less than the threshold, then this set of hyperparameters is used as a set of parent hyperparameters. If the above deviation is greater than or equal to the threshold, then this set of hyperparameters is discarded. The optimal fitness value is obtained by comparing the fitness values corresponding to each set of hyperparameters from the first round of training to the previous round of training, and taking the minimum fitness value as the optimal fitness value. It can be understood that for the first round of training, there is no optimal fitness value, and all sets of hyperparameters can be directly used as parent hyperparameters.
[0103] The deviation between the fitness value corresponding to each set of hyperparameters in this round and the optimal fitness value can be calculated by the formula |f - f best | / f best . Here, f represents the fitness value corresponding to a set of hyperparameters, f best represents the optimal fitness value, and ε represents the threshold.
[0104] Step 7: Perform the crossover operation. Perform the crossover operation on each set of parent hyperparameters (i.e., each set of retained hyperparameters) to obtain each set of child hyperparameters. The crossover operation selects the adaptive arithmetic crossover method. Calculate a set of child hyperparameters according to the formula child = λ·Parent1 + (1 - λ)·Parent2. The adaptive crossover probability of each set of parent hyperparameters can be calculated according to the formula Calculate and obtain.
[0105] Step 8: Perform the mutation operation. Perform the mutation operation on each set of child hyperparameters to obtain each set of optimized hyperparameters.
[0106] Step 9: Update the model parameters of the DDPG model. Using the optimized hyperparameters, update the hyperparameters of the policy network and Q-network of the DDPG model. Then, train the model based on the current training data and the updated hyperparameters. A soft update strategy is used to smoothly adjust the weights of the target network.
[0107] After obtaining the optimized hyperparameters, assign each optimized hyperparameter to the policy network and Q network of the DDPG model to obtain the DDPG model corresponding to each optimized hyperparameter.
[0108] For each set of hyperparameters corresponding to the DDPG model, actions are selected through the policy network, executed in the environment, and state transitions and rewards are observed. Data is stored in an experience buffer. A batch of experience data is randomly sampled from the experience buffer, and a Q-network is used to calculate the Q-value as an evaluation of the action. The target Q-value is calculated from the action generated by the target policy network in the next state and the prediction value of the target Q-network. The update formula for the target Q-value is: Q(s′,a′)=r+γ·Q target (s′,A target (s′))
[0109] Where s′ is the next state, r is the reward, γ is the discount factor, and A target It is a target policy network, Q target It is a target evaluation network.
[0110] The state space of the DDPG algorithm is defined as follows:
[0111] Among them, T vrr (v) represents the arrival time of tanker v, T vlea (v) represents the estimated departure time of tanker v, V v (p) represents the amount of crude oil loaded in the p-th oil tanker of tanker v. This represents the initial inventory of crude oil of type k in the crude oil storage tank at terminal i. Let R(k) represent the initial inventory of crude oil of type k in the crude oil storage tank (feeding tank) of the j-th plant area, P represent the processing capacity of the atmospheric and vacuum distillation unit, R(k) represent the proportion of crude oil of type k in the current processing plan, and Q represent the total processing volume of the current processing plan.
[0112] The action space of the DDPG algorithm is defined as: a = {X(p,i,n),Y(i,j,n),Z(j,l,n),CO(l,n),Bx(p,i,n),By(i,j,n),Bz(j,l,n)}
[0113] Where X(p,i,n) represents whether the p-th oil tank unloads oil to the ith crude oil storage tank at the terminal during the nth time period, and is a 0-1 variable; Y(i,j,n) represents whether the ith crude oil storage tank at the terminal ...
[0114] The constraints for crude oil dispatching include material balance constraints, component material balance constraints, crude oil property constraints, time constraints, and general rule constraints, which are set according to actual needs and are not limited in the embodiments of this application. If, during the training of the sub-model, the constraints for crude oil dispatching are not met, the values can be adjusted to make the modified values meet the constraints for crude oil dispatching.
[0115] The mean-square error (MSE) loss between the predicted Q-value and the target Q-value of the evaluation network is calculated. Backpropagation and the steepest descent method are used to update the weights of the evaluation network to minimize the loss.
[0116] The evaluation network is used to assess the value of the action in the current state and calculate the action gradient. The gradient of the action is updated by the output of the policy network to maximize the output of the evaluation network, thereby completing the weight update of the policy network.
[0117] After obtaining the parameters of the policy network and the target network, the weights of the two target networks can be updated using a soft update method. The weight update formula is expressed as: θ target =λ·θ orgin +(1-λ)·θ target
[0118] Where θ target θ represents the target network weights. orgin λ represents the weights of the corresponding main network (i.e., the policy network and the target network), and λ∈(0,1] represents the update ratio.
[0119] Step 10: Model training complete. Based on the final optimal set of hyperparameters and the last updated DDPG model, the crude oil blending and scheduling model is obtained.
[0120] After obtaining the crude oil blending and scheduling model, the crude oil blending resource status data of refinery B is obtained. The crude oil blending resource status data of refinery B is input into the crude oil blending and scheduling model, and the crude oil blending and scheduling scheme of refinery B can be output.
[0121] This patent utilizes the DDPG algorithm to solve the crude oil blending and scheduling problem. Facing a vast decision space containing both continuous and discrete variables, it employs a data-driven approach to achieve rapid solutions, avoiding the infeasibility issues encountered in traditional rigorous mathematical programming. During the solution process, a genetic algorithm is used to optimize the hyperparameter settings of DDPG, achieving automatic hyperparameter adjustment and avoiding the tedious process of manual adjustment. This improves training speed and model convergence, and enables the handling of randomness and uncertainty in the production process, achieving more flexible and intelligent crude oil blending decisions.
[0122] Figure 7 is a structural schematic diagram of the crude oil blending and scheduling device provided in the seventh embodiment of this application. As shown in Figure 7, the crude oil blending and scheduling device provided in this embodiment includes a first acquisition module 701 and a prediction module 702, wherein:
[0123] The first acquisition module 701 is used to acquire crude oil blending resource status data; the prediction module 702 is used to obtain crude oil blending scheduling scheme based on crude oil blending resource status data and crude oil blending scheduling model; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
[0124] Specifically, the first acquisition module 701 can acquire crude oil blending resource status data, which is the basic data for making action decisions on crude oil blending production in the future. Crude oil blending resource status data includes, but is not limited to, tanker arrival time, estimated tanker departure time, crude oil loading volume and type in tanker tanks, initial inventory of a certain crude oil in a certain terminal storage tank, and initial inventory of a certain crude oil in a certain plant storage tank, etc., and can be set according to actual needs; this application embodiment does not limit this.
[0125] The prediction module 702 inputs crude oil blending resource status data into the crude oil blending scheduling model and can output a crude oil blending scheduling plan. Crude oil blending scheduling is a sequential decision-making process. The crude oil blending scheduling plan can include the initial state and action decisions for each cycle. The initial state for each cycle can include the state of the tanker, the state of the tanker's oil tanks, the state of the terminal crude oil storage tanks, and the state of the plant's crude oil storage tanks, etc. The action decisions for each cycle can include whether a tanker's oil tank is unloading oil from a storage tank in the current cycle, whether a storage tank is transferring oil to a loading tank in the current cycle, whether a loading tank is transferring oil to the atmospheric and vacuum distillation unit in the current cycle, whether a feeding switch occurs in the current cycle for the current cycle, the amount of oil unloaded from a tanker to a storage tank in the current cycle, the amount of oil transferred from a storage tank to a loading tank in the current cycle, and the amount of oil transferred from a loading tank to the atmospheric and vacuum distillation unit in the current cycle, etc. The crude oil blending scheduling plan can be used for crude oil scheduling in crude oil blending production.
[0126] The crude oil blending and scheduling model is trained based on crude oil blending sample data and corresponding actual scheduling actions. The crude oil blending sample data can be obtained from historical crude oil blending resource status data, and the actual scheduling actions corresponding to the crude oil blending sample data serve as labels for the sample data. During model training, the model's hyperparameters are optimized, and the model with the optimized hyperparameters is then trained. In other words, this application does not pre-specify the model's hyperparameters during training, nor does it require manual adjustment by technicians. Instead, hyperparameter adjustment is integrated into the model training process, automatically optimizing and adjusting the model's hyperparameters during training. Because the hyperparameters can be automatically optimized and adjusted, the training efficiency is improved compared to the manual adjustment of hyperparameters in existing technologies.
[0127] The crude oil blending and scheduling device provided in this application acquires crude oil blending resource status data; and obtains a crude oil blending and scheduling scheme based on the crude oil blending resource status data and the crude oil blending and scheduling model. The crude oil blending and scheduling model is obtained through training based on crude oil blending sample data and corresponding actual scheduling actions. During model training, the hyperparameters of the model are optimized and the model training is accelerated. The optimized hyperparameters are then used for model training, which improves the efficiency of obtaining the crude oil blending and scheduling scheme.
[0128] Figure 8 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the eighth embodiment of this application. As shown in Figure 8, based on the above embodiments, the crude oil blending and scheduling device provided in this application further includes a second acquisition module 703, an initialization module 704, a generation module 705, and a training module 706, wherein:
[0129] The second acquisition module 703 is used to acquire crude oil blending sample data and corresponding actual scheduling actions; the initialization module 704 is used to initialize the original model and acquire the value range of hyperparameters; the generation module 705 is used to generate each set of hyperparameters according to the value range of hyperparameters; the training module 706 is used to train and obtain the crude oil blending scheduling model according to the crude oil blending sample data, corresponding actual scheduling actions, each set of hyperparameters, fitness function and original model; wherein, the fitness function is preset.
[0130] Figure 9 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the ninth embodiment of this application. As shown in Figure 9, based on the above embodiments, the training module 706 further includes a prediction unit 7061, an acquisition unit 7062, an optimization unit 7063, and a training unit 7064, wherein:
[0131] The prediction unit 7061 is used to obtain the crude oil blending and scheduling scheme corresponding to each set of hyperparameters based on the sample data of crude oil blending in this round and the training model corresponding to each set of hyperparameters; the acquisition unit 7062 is used to obtain the fitness value corresponding to each set of hyperparameters in this round based on the crude oil blending and scheduling scheme and fitness function corresponding to each set of hyperparameters in this round; the optimization unit 7063 is used to optimize the hyperparameters based on each set of hyperparameters and the fitness value corresponding to each set of hyperparameters in this round, and obtain the optimized hyperparameters for each set; the training unit 7064 is used to perform one round of model training on the training model corresponding to each set of hyperparameters based on the sample data of crude oil blending in this round and the corresponding actual scheduling actions.
[0132] Figure 10 is a schematic diagram of the structure of the crude oil blending and scheduling device provided in the tenth embodiment of this application. As shown in Figure 10, based on the above embodiments, the optimization unit 7063 further includes a screening subunit 70631, a cross-operation subunit 70632, and a mutation operation subunit 70633, wherein:
[0133] The screening subunit 70631 is used to obtain the retained hyperparameters of each group based on the fitness value of each group in this round and the screening conditions; the crossover operation subunit 70632 is used to perform crossover operation based on the retained hyperparameters of each group and the fitness value corresponding to the retained hyperparameters of each group to obtain the offspring hyperparameters of each group; the mutation operation subunit 70633 is used to perform mutation operation on the offspring hyperparameters of each group to obtain the optimized hyperparameters of each group.
[0134] Based on the above embodiments, the further filtering condition includes that if the difference between the fitness value of a set of hyperparameters in the current round and the optimal fitness is less than a threshold, then the set of hyperparameters is regarded as a set of retained hyperparameters.
[0135] Based on the above embodiments, the fitness function is further defined as the sum of the tanker's waiting costs at sea, unloading costs, tank inventory costs, and costs incurred in switching the feed to the atmospheric and vacuum distillation unit.
[0136] Figure 11 is a structural schematic diagram of the crude oil blending and scheduling device provided in the eleventh embodiment of this application. As shown in Figure 11, based on the above embodiments, the crude oil blending and scheduling device provided in this application further includes a scheduling module 707 and a judgment module 708, wherein:
[0137] The scheduling module 707 is used to perform actual crude oil blending scheduling according to the crude oil blending scheduling scheme; the judgment module 708 is used to re-obtain the crude oil blending scheduling scheme based on the current actual crude oil blending resource status data and the crude oil blending scheduling model after determining that the error between the current actual crude oil blending resource status data and the predicted current crude oil blending resource status data exceeds the error threshold.
[0138] The embodiments of the apparatus provided in this application can be used to execute the processing flow of the above-described method embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the above-described method embodiments.
[0139] Figure 12 is a schematic diagram of the physical structure of the electronic device provided in the twelfth embodiment of this application. As shown in Figure 12, the electronic device 600 may include a processor 100 and a memory 140. The memory 140 is coupled to the processor 100. The processor 100 can call the logical instructions in the memory 140 to execute the methods provided in the above-described method embodiments, such as: acquiring crude oil blending resource status data; obtaining a crude oil blending scheduling scheme based on the crude oil blending resource status data and the crude oil blending scheduling model; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated to obtain the model corresponding to the optimized hyperparameters for model training.
[0140] This embodiment discloses a computer program product, which includes a computer program / instructions stored on a non-transitory computer-readable storage medium. When the computer program / instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments, such as: acquiring crude oil blending resource status data; obtaining a crude oil blending scheduling scheme based on the crude oil blending resource status data and the crude oil blending scheduling model; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated to obtain the model corresponding to the optimized hyperparameters for model training.
[0141] This embodiment provides a computer-readable storage medium storing a computer program / instruction that causes a computer to execute the methods provided in the above-described method embodiments, such as: acquiring crude oil blending resource status data; obtaining a crude oil blending scheduling scheme based on the crude oil blending resource status data and a crude oil blending scheduling model; wherein the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during model training, the hyperparameters of the model are optimized and the model training is accelerated to obtain the model corresponding to the optimized hyperparameters for model training.
[0142] As shown in Figure 12, the electronic device 600 may further include: a communication module 110, an input unit 120, an audio processor 130, a display 160, and a power supply 170. It is worth noting that the electronic device 600 is not necessarily required to include all the components shown in Figure 12; furthermore, the electronic device 600 may also include components not shown in Figure 12, as can be found in existing technologies. It is important to note that this figure is exemplary; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.
[0143] As shown in Figure 12, the processor 100, sometimes also referred to as a controller or operation control, may include a microprocessor or other processor device and / or logic device. The processor 100 receives input and controls the operation of various components of the electronic device 600.
[0144] The memory 140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The processor 100 may execute the program stored in the memory 140 to perform information storage or processing, etc.
[0145] Input unit 120 provides input to processor 100. Input unit 120 may be, for example, a keypad or touch input device. Power supply 170 provides power to electronic device 600. Display 160 displays images and text. Display 160 may be, for example, an LCD display, but is not limited thereto.
[0146] Memory 140 can be a solid-state memory, such as read-only memory (ROM), random access memory (RAM), SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of memory 140 are sometimes referred to as EPROM, etc. Memory 140 can also be some other type of device. Memory 140 includes a buffer 141 (sometimes referred to as buffer memory). Memory 140 may include an application / function storage unit 142 for storing application programs and function programs or processes for executing the operation of electronic device 600 by processor 100.
[0147] The memory 140 may also include a data storage unit 143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 144 of the memory 140 may include various drivers for the electronic device's communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0148] The communication module 110 includes a transmitter / receiver that transmits and receives signals via antenna 111. The communication module 110 is coupled to processor 100 to provide input signals and receive output signals, which can be the same as in a conventional mobile communication terminal.
[0149] Based on different communication technologies, multiple communication modules 110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 110 is also coupled to a speaker 131 and a microphone 132 via an audio processor 130 to provide audio output via the speaker 131 and receive audio input from the microphone 132, thereby realizing typical telecommunications functions. The audio processor 130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 130 is coupled to the processor 100, enabling on-device recording via the microphone 132 and on-device playback of stored sound via the speaker 131.
[0150] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0151] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0152] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0153] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0154] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0155] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A crude oil blending and scheduling method, characterized in that, include: Obtain data on the status of crude oil blending resources; as well as Based on the crude oil blending resource status data and the crude oil blending scheduling model, a crude oil blending scheduling scheme is obtained; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the training process of the crude oil blending scheduling model, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
2. The method according to claim 1, characterized in that, The crude oil blending scheduling model, trained based on crude oil blending sample data and corresponding actual scheduling actions, includes: Obtain crude oil blending sample data and corresponding actual scheduling actions; Initialize the original model and obtain the range of hyperparameter values; Based on the range of hyperparameter values, generate each set of hyperparameters; and The crude oil blending scheduling model is trained based on crude oil blending sample data, corresponding actual scheduling actions, hyperparameters of each group, fitness functions, and the original model; wherein the fitness function is preset.
3. The method according to claim 2, characterized in that, The process of training the crude oil blending scheduling model based on crude oil blending sample data, corresponding actual scheduling actions, hyperparameters, fitness functions, and the original model includes: Based on the sample data of this round of crude oil blending and the model to be trained corresponding to each set of hyperparameters, the prediction results of the crude oil blending scheduling action for this round corresponding to each set of hyperparameters are obtained. Based on the prediction results and fitness function of the current round of crude oil blending and scheduling actions corresponding to each set of hyperparameters, the fitness value of the current round corresponding to each set of hyperparameters is obtained; Hyperparameter optimization is performed based on each set of hyperparameters and their corresponding fitness values in the current round, resulting in optimized sets of hyperparameters; and Based on the sample data of crude oil blending in this round and the corresponding actual scheduling actions, a round of model training is carried out on the training models corresponding to each set of optimized hyperparameters.
4. The method according to claim 3, characterized in that, The hyperparameter optimization is performed based on each set of hyperparameters and the corresponding fitness of each set of hyperparameters in the current round, resulting in the following optimized sets of hyperparameters: Based on the fitness value and screening conditions corresponding to each group of hyperparameters in this round, the retained hyperparameters of each group are obtained; Cross-operation is performed based on the retained hyperparameters of each group and the corresponding fitness values of each group to obtain the hyperparameters of the offspring generation; and Mutation operations are performed on the hyperparameters of each group of offspring to obtain the optimized hyperparameters of each group.
5. The method according to claim 4, characterized in that, The filtering criteria include: If the deviation between the fitness value of a set of hyperparameters in the current round and the optimal fitness value is less than a threshold, then the set of hyperparameters is retained as a set of hyperparameters.
6. The method according to claim 3, characterized in that, The fitness function is the sum of the tanker's waiting costs at sea, unloading costs, tank inventory costs, and costs incurred in switching feed to the atmospheric and vacuum distillation unit.
7. The method according to claim 1, characterized in that, Also includes: Actual crude oil blending and scheduling shall be carried out according to the crude oil blending and scheduling scheme. as well as If the error between the current actual crude oil blending resource status data and the predicted current crude oil blending resource status data exceeds the error threshold, then a new crude oil blending scheduling plan is obtained based on the current actual crude oil blending resource status data and the crude oil blending scheduling model.
8. A crude oil blending and scheduling device, characterized in that, include: The first acquisition module is used to acquire crude oil blending resource status data; as well as The prediction module is used to obtain a crude oil blending scheduling scheme based on the crude oil blending resource status data and the crude oil blending scheduling model; wherein, the crude oil blending scheduling model is obtained by training based on crude oil blending sample data and corresponding actual scheduling actions; during the model training process, the hyperparameters of the model are optimized and the model training is accelerated, and the model corresponding to the optimized hyperparameters is used for model training.
9. The apparatus according to claim 8, characterized in that, Also includes: The second acquisition module is used to acquire crude oil blending sample data and corresponding actual scheduling actions. The initialization module is used to initialize the original model and obtain the range of hyperparameter values; The generation module is used to generate each set of hyperparameters based on the range of hyperparameter values. as well as The training module is used to train the crude oil blending scheduling model based on crude oil blending sample data, corresponding actual scheduling actions, hyperparameters of each group, fitness functions, and the original model; wherein the fitness function is preset.
10. The apparatus according to claim 9, characterized in that, The training module includes: The prediction unit is used to obtain the crude oil blending and scheduling scheme for each set of hyperparameters based on the sample data of this round of crude oil blending and the model to be trained for each set of hyperparameters. The acquisition unit is used to obtain the fitness value of each set of hyperparameters in this round based on the current round crude oil blending and scheduling scheme and fitness function corresponding to each set of hyperparameters. The optimization unit is used to optimize hyperparameters based on each set of hyperparameters and the corresponding fitness value in the current round, to obtain the optimized hyperparameters for each set; and The training unit is used to train the model corresponding to each set of optimized hyperparameters in one round based on the sample data of crude oil blending in this round and the corresponding actual scheduling actions.
11. The apparatus according to claim 10, characterized in that, The optimization unit includes: The filtering subunit is used to obtain the retained hyperparameters for each group based on the fitness value and filtering conditions corresponding to each group of hyperparameters in this round. The crossover subunit is used to perform crossover operations based on the retained hyperparameters of each group and the fitness values corresponding to the retained hyperparameters of each group to obtain the hyperparameters of each group of offspring; and The mutation operation subunit is used to perform mutation operations on the hyperparameters of each group of offspring to obtain the optimized hyperparameters of each group.
12. The apparatus according to claim 11, characterized in that, The selection criteria include: if the deviation between the fitness value of a set of hyperparameters in this round and the optimal fitness value is less than a threshold, then the set of hyperparameters is selected as a set of retained hyperparameters.
13. The apparatus according to claim 10, characterized in that, The fitness function is the sum of the tanker's waiting costs at sea, unloading costs, tank inventory costs, and costs incurred in switching feed to the atmospheric and vacuum distillation unit.
14. The apparatus according to claim 8, characterized in that, Also includes: The scheduling module is used to perform actual scheduling of crude oil blending according to the crude oil blending and scheduling scheme. as well as The judgment module is used to re-obtain the crude oil blending and scheduling scheme based on the current actual crude oil blending and scheduling model after determining that the error between the current actual crude oil blending and scheduling resource status data and the predicted current crude oil blending and scheduling data exceeds the error threshold.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Deep reinforcement learning model training method and device based on hyper-parameter optimization
CN113723615A
Deep reinforcement learning task scheduling method and device for offshore unmanned equipment
CN115309521A
Method for controlling deposition process of mud cabin of trailing suction dredger based on reinforcement learning
CN117369262A
Material scheduling method based on deep reinforcement learning improved genetic algorithm and related device
CN117455143A
Intelligent fuel dispatching system and method
CN118569618A