Coking process intelligent control method and control system based on reinforcement learning
By constructing an intelligent control method based on LSTM and reinforcement learning, the problem of insufficient data processing in the coking production process was solved, achieving high-precision and stable automatic control, reducing reliance on personnel, and improving production safety and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-14
AI Technical Summary
The existing coking production process control system has limited data access and low level of data intelligence processing, resulting in poor control accuracy, stability and predictability. It is also highly dependent on personnel, posing safety hazards and low production efficiency.
An intelligent control method based on Long Short-Term Memory Neural Network (LSTM) and reinforcement learning is constructed. By collecting and processing historical data, a production environment simulation model is built to achieve intelligent control of the coking process.
It improves control precision and stability, reduces reliance on personnel, enhances production safety and efficiency, and enables intelligent control of complex chemical production processes.
Smart Images

Figure CN121857320A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial process control technology, specifically relating to an intelligent control method and control system for coking processes based on reinforcement learning, which is particularly suitable for complex process control scenarios involving multiple variables, nonlinearity, and strong coupling in coking production. Background Technology
[0002] The benzene / cooling / desulfurization / ammonium sulfate removal process in chemical production is a complex system engineering project, easily affected by factors such as production operating conditions, equipment status, the skill level of technical operators, energy media, safety and environmental protection, external environment, and market economy. Currently, this leads to a disconnect between planning and control optimization. The chemical production process control system uses conventional feedback control, which is "post-event adjustment" and lacks feedforward predictive control. Secondly, due to issues with control loop parameter settings and the influence of equipment and instrument valves, the system is essentially in a manual adjustment state. Currently, manual adjustment results in relatively stable production, but changes in coking load, equipment status, and energy media cause fluctuations. Manual operation is complex and labor-intensive, and differences in manual adjustment experience directly affect the stability of the production process. Thirdly, the high-altitude and cold-climate chemical production areas in Northeast China pose a risk of hazardous gas leaks, making these areas high-risk zones where personnel exposure time should not be prolonged. Therefore, the chemical production safety management and control system has bottleneck problems such as many process parameters, large data fluctuations, black box control, and lagging adjustment. The original control system has limited access to data, low level of data intelligent processing, poor predictability of control effects, and cannot achieve coordinated optimization and predictive control of the process, resulting in the process control not meeting expectations.
[0003] Chemical production process control is a complex systems engineering project. Production processes are easily affected by external factors such as operating conditions, equipment status, external environment, and market economy, leading to a disconnect between planning and control optimization. This makes it difficult to achieve continuous and effective coordination and balanced production throughout the process, resulting in problems such as poor control quality, increased labor intensity in certain sections, and increased energy consumption. Based on the characteristics of process control in coking chemical products, this study investigates the process flows of benzene washing, cooling drum, desulfurization, and ammonium sulfate systems. A mathematical model for predicting, analyzing, and controlling chemical production processes is constructed, ultimately enabling coordinated and optimized control of each process.
[0004] Therefore, this invention presents a production environment simulation model based on a long short-term memory neural network to upgrade the intelligence level of the coking process, reduce the time spent in dangerous areas of coking, reduce direct dependence on personnel, and improve employee safety. Furthermore, it proposes an intelligent control method for the coking process based on reinforcement learning. Summary of the Invention
[0005] The purpose of this invention is to address the problems of limited data access, low level of intelligent data processing, and high dependence on personnel in existing control systems, which lead to poor control accuracy, stability, predictability, efficiency, and safety. Therefore, this invention proposes an intelligent control method and control system for the coking process based on reinforcement learning.
[0006] The specific process of an intelligent control method for coking process based on reinforcement learning is as follows:
[0007] Step 1: Collect historical data of the coking production process. Based on the box plot outlier detection method, remove outliers from the historical data to obtain the historical data after removing outliers. Normalize the historical data after removing outliers to obtain the normalized historical data. The historical data of the coking production process includes temperature, pressure, and flow rate.
[0008] Step 2: Construct an improved LSTM neural network model; the improved LSTM neural network model includes, in sequence: LSTM neural network model, ReLU activation function layer, Dropout layer, and Dense;
[0009] The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0010] The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0011] The improved LSTM neural network model is trained based on the historical data after normalization in step one, and the trained improved LSTM neural network model is obtained.
[0012] Step 3: Use the trained improved LSTM neural network model as the environment for the reinforcement learning control model;
[0013] The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate; the action taken by the agent is a normalized control variable.
[0014] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0015] Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model;
[0016] Step 4: Collect online data of the coking production process, and remove outliers from the online data using the box plot outlier detection method to obtain the online data after removing outliers;
[0017] The online data after removing outliers is normalized to obtain normalized online data.
[0018] Online data for the coking production process includes temperature, pressure, and flow rate;
[0019] The normalized online data is input into the trained reinforcement learning control model, which then outputs the temperature at the next moment.
[0020] Preferably, in step one, historical data of the coking production process is collected, and outlier detection based on a box plot is used to remove outlier data, resulting in historical data after outlier removal; the historical data after outlier removal is then normalized to obtain normalized historical data; the historical data of the coking production process includes temperature, pressure, and flow rate; the specific process is as follows:
[0021] Step 1: Collect historical data on the coking production process;
[0022] Steps 1 and 2: Divide the collected historical data into fixed time intervals. A subset;
[0023] Step 13: For each subset, calculate the first quartile, third quartile, interquartile range, upper boundary, and lower boundary; treat data outside the upper and lower boundary intervals of each subset as outliers, remove outliers, and obtain the historical data after removing outliers;
[0024] Step 14: Normalize the historical data after removing outliers to obtain normalized historical data.
[0025] Preferably, in step one, historical data of the coking production process is collected; the specific process is as follows:
[0026] 1) Collect temperature data from January to October of the previous historical year from the Distributed Control System (DCS) during the coking production process. ;
[0027] in, This represents the temperature data for January of the previous historical year. This represents the temperature data for February of the previous historical year. This represents the temperature data for March of the previous historical year. This represents the temperature data for April of the previous historical year. This represents the temperature data for May of the previous historical year. This represents the temperature data for June of the previous historical year. This represents the temperature data for July of the previous historical year. This represents the temperature data for August of the previous historical year. This represents the temperature data for September of the previous historical year. This represents the temperature data for October of the previous historical year.
[0028] 2) Collect pressure data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. ;
[0029] in, This represents the stress data for January of the previous historical year. This represents the stress data for February of the previous historical year. This indicates the stress data for March of the previous historical year. This represents the stress data for April of the previous historical year. This represents the stress data for May of the previous historical year. This represents the stress data for June of the previous historical year. This represents the stress data for July of the previous historical year. This represents the stress data for August of the previous historical year. This indicates the stress data for September of the previous historical year. This represents the stress data for October of the previous historical year;
[0030] 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. ;
[0031] in, This represents traffic data for January of the previous historical year. This represents traffic data for February of the previous historical year. This represents traffic data for March of the previous historical year. This represents traffic data for April of the previous historical year. This represents traffic data for May of the previous historical year. This represents traffic data for June of the previous historical year. This represents traffic data for July of the previous historical year. This represents traffic data for August of the previous historical year. This represents traffic data for September of the previous historical year. This represents the traffic data for October of the previous historical year.
[0032] Preferably, in steps one and two, the collected historical data is divided into segments according to fixed time intervals. A subset; the specific process is as follows:
[0033] 1) Collect temperature data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. Divided into according to fixed time intervals A subset of temperature data ;
[0034] in, This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of temperature data after being divided at fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of temperature data after being divided according to fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data;
[0035] 2) Collect pressure data from the distributed control system (DCS) of the previous historical year (January to October) for the coking production process. Divided into according to fixed time intervals A subset of stress data ;
[0036] in, This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data;
[0037] 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. Divided into according to fixed time intervals Subset ;
[0038] in, This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data.
[0039] Preferably, in steps one and three, the first quartile, third quartile, interquartile range, upper boundary, and lower boundary are calculated for each data subset; data outside the upper and lower boundary intervals of each subset are considered outliers, and outliers are removed to obtain historical data after outlier removal; the specific process is as follows:
[0040] 1) For a subset of temperature data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain the historical data after outlier removal; specifically:
[0041] subset First quartile , Subset The third quartile , Subset interquartile range , ;
[0042] subset upper boundary Subset lower boundary
[0043] express The 25th percentile of the data; express The 75th percentile of the data;
[0044] Representing a subset The upper boundary; Representing a subset The lower boundary;
[0045] 2) For a subset of stress data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain historical pressure data after outlier removal; specifically:
[0046] subset First quartile , Subset The third quartile , Subset interquartile range , ;
[0047] subset upper boundary Subset lower boundary
[0048] express The 25th percentile of the data; express The 75th percentile of the data;
[0049] Representing a subset The upper boundary; Representing a subset The lower boundary;
[0050] 3) For subsets of traffic data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. These outliers are removed to obtain the historical traffic data after outlier removal. Specifically:
[0051] subset First quartile Subset The third quartile Subset interquartile range ;
[0052] subset upper boundary Subset lower boundary
[0053] express The 25th percentile of the data; express The 75th percentile of the data;
[0054] Representing a subset The upper boundary; Representing a subset The lower boundary.
[0055] Preferably, in step one four, the historical data after removing outlier data is normalized to obtain normalized historical data; the specific process is as follows:
[0056] The historical temperature data after removing outliers is normalized to obtain the normalized historical temperature data.
[0057] The historical stress data after removing outliers is normalized to obtain the normalized historical stress data.
[0058] The historical traffic data after removing outliers is normalized to obtain normalized historical traffic data.
[0059] Preferably, in step two, an improved LSTM neural network model is constructed; the improved LSTM neural network model includes, in sequence: an LSTM neural network model, a ReLU activation function layer, a Dropout layer, and a Dense layer; the Dropout layer is a regularization layer; the Dense layer is a fully connected layer.
[0060] The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0061] The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0062] Train the improved LSTM neural network model to obtain a well-trained improved LSTM neural network model;
[0063] The specific process is as follows:
[0064] The historical data after the normalization process in step one is divided into a training set and a test set;
[0065] The training set is input into the improved LSTM neural network model, with mean squared error as the loss function, and the Adam optimizer is used to optimize the parameters of the improved LSTM neural network model; the test set is input into the improved LSTM neural network model, and the improved LSTM neural network model outputs predicted values and actual values. When the coefficient reaches 0.98, a well-trained improved LSTM neural network model is obtained.
[0066] Preferably, in step three, the trained improved LSTM neural network model is used as the environment for the reinforcement learning control model;
[0067] The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate; the action taken by the agent is a normalized control variable.
[0068] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0069] Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model;
[0070] The specific process is as follows:
[0071] The trained improved LSTM neural network model is used as the environment for the reinforcement learning control model;
[0072] The current state and the actions taken by the agent are inputs to a reinforcement learning control model;
[0073] The input states for the reinforcement learning control model are temperature, pressure, and flow rate;
[0074] The action taken by the agent inputting the reinforcement learning control model is the normalized valve opening of the flow control valve; the agent is the flow control valve.
[0075] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0076] The action taken by the agent in the next moment, as output by the reinforcement learning control model, is converted into the valve opening of the flow control valve after inverse normalization.
[0077] The reward is calculated based on the next state output by the reinforcement learning control model.
[0078] The reinforcement learning control model is trained until the maximum number of iterations is reached, at which point a well-trained reinforcement learning control model is obtained.
[0079] Preferably, the reward is calculated based on the next-time state output by the reinforcement learning control model; the specific process is as follows:
[0080] like ,award ;like ,award ;like ,award ;other: .
[0081] A reinforcement learning-based intelligent control system for a coking process is used to execute a reinforcement learning-based intelligent control method for a coking process.
[0082] The beneficial effects of this invention are as follows:
[0083] This invention proposes a complete reinforcement learning-based industrial process control model architecture, enabling intelligent control of complex chemical production processes. This architecture constructs a high-precision production simulation environment model and trains a reinforcement learning framework on it. Compared to traditional industrial PID control strategies, this two-layer architecture of "production environment simulation model + reinforcement learning control model" achieves automatic control of the control system. It provides self-optimization and self-learning capabilities for various industrial control modules, significantly improving control accuracy and stability, reducing dwell time in hazardous areas of coking plants, reducing direct dependence on personnel, and enhancing the efficiency and safety of coking processes. Attached Figure Description
[0084] Figure 1 This is a schematic diagram of the structure of an intelligent control method for coking process based on reinforcement learning proposed in this invention.
[0085] Figure 2 This is a schematic diagram of the production environment simulation model proposed in this invention;
[0086] Figure 3a This is a schematic diagram showing the distribution of crude benzene reflux flow rate after preprocessing. FT_82208.AV represents the crude benzene reflux flow rate.
[0087] Figure 3b This is a schematic diagram of the distribution of preprocessed column top temperature data. TT_82207.AV represents the column top temperature, and TT is the temperature.
[0088] Figure 3c This is a schematic diagram showing the distribution of preprocessed tower top pressure data. PT_82208.AV represents the tower top pressure, and PT is the pressure.
[0089] Figure 3d This is a schematic diagram showing the distribution of preprocessed tower top pressure data; PT_82209.AV represents the tower bottom pressure.
[0090] Figure 3e This is a schematic diagram showing the distribution of preprocessed bottom temperature data. TT_82210.AV represents the bottom temperature.
[0091] Figure 3f This is a schematic diagram showing the distribution of approximate crude benzene temperature data after preprocessing. Jinsicubenwendu represents the approximate crude benzene temperature.
[0092] Figure 4a This is a diagram showing the comparison between the predicted and actual values of the tower top temperature in an LSTM-based production environment simulation model. TT_82207.AV represents the tower top temperature.
[0093] Figure 4b This is a diagram showing the comparison between the predicted and actual values of the tower top pressure in an LSTM-based production environment simulation model. PT_82208.AV represents the tower top pressure.
[0094] Figure 4c This is a diagram showing the comparison between the predicted and actual values of the bottom pressure in the LSTM-based production environment simulation model. PT_82209.AV represents the bottom pressure.
[0095] Figure 4d This is a diagram showing the comparison between the predicted and actual values of the bottom temperature in a production environment simulation model based on LSTM. TT_82210.AV represents the bottom temperature.
[0096] Figure 4e This is a diagram showing the comparison between the predicted and actual values of the approximate crude benzene temperature in a production environment simulation model based on LSTM. Jinsicubenwendu represents the approximate crude benzene temperature.
[0097] Figure 5 A schematic diagram illustrating the results of reinforcement learning stochastic control;
[0098] Figure 6 This is a schematic diagram of the control results based on a production environment simulation model and a reinforcement learning control model.
[0099] Figure 7 This is a schematic diagram comparing the predicted and actual values of the controlled parameter based on a production environment simulation model and a reinforcement learning control model. Detailed Implementation
[0100] Specific Implementation Method 1: The specific process of this implementation method for intelligent control of the coking process based on reinforcement learning is as follows:
[0101] Step 1: Collect historical data for each stage of the coking production process. Use a box plot outlier detection method to remove outliers from the historical data to obtain the historical data after outlier removal. Normalize the historical data after outlier removal to obtain the normalized historical data and eliminate the influence of dimensions. The historical data for each stage of the coking production process includes temperature, pressure, and flow rate.
[0102] Step 2: Construct an improved LSTM neural network model; the improved LSTM neural network model includes, in sequence: LSTM neural network model, ReLU activation function layer, Dropout layer, and Dense;
[0103] The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0104] The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0105] The improved LSTM neural network model is trained based on the historical data after normalization in step one, and the trained improved LSTM neural network model is obtained.
[0106] The Adam optimizer was selected, and mean squared error was chosen as the loss function. This is suitable for regression problems and measures the difference between the model's predicted and actual values. 80% of the processed data was used as the training set, and 20% as the test set. The input and output formats were modified to meet the model's requirements, and the data was normalized before use. The model was trained using the training set and tested using the test set. The R² coefficients for the actual and predicted values of each output parameter were all around 0.98, indicating a good fit. The test results are shown in the figure below. Figure 4a , Figure 4b , Figure 4c , Figure 4d , Figure 4e ;
[0107] As shown.
[0108] The data was divided into a training set (80%) and a test set (20%). Mean squared error (MSE) was used as the loss function, and the Adam optimizer was employed for parameter optimization. Finally, the weights of the improved LSTM neural network model trained from the data collected in step one were obtained, which were used to provide an environment for subsequent reinforcement learning.
[0109] Step 3: Use the trained improved LSTM neural network model as the environment for the reinforcement learning control model;
[0110] The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate; the action taken by the agent is a normalized control variable, such as valve opening.
[0111] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0112] Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model;
[0113] Step 4: Collect online data from each stage of the coking production process, and remove outliers from the online data using a box plot outlier detection method to obtain the online data after removing outliers;
[0114] The online data after removing outliers is normalized to obtain normalized online data, thus eliminating the influence of units.
[0115] Online data for each stage of the coking production process includes temperature, pressure, and flow rate;
[0116] The normalized online data is input into the trained reinforcement learning control model, which then outputs the temperature at the next moment.
[0117] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step one, historical data of each stage of the coking production process is collected, and outlier detection methods based on box plots are used to remove outlier data, resulting in historical data after outlier removal; the historical data after outlier removal is then normalized to obtain normalized historical data, eliminating the influence of dimensions; the historical data of each stage of the coking production process includes temperature, pressure, and flow rate; the specific process is as follows:
[0118] Step 1: Collect historical data for each stage of the coking production process;
[0119] Steps 1 and 2: Divide the collected historical data into fixed time intervals (e.g., 2 months). Five (e.g.) subsets;
[0120] Step 13: For each subset, calculate the first quartile, third quartile, interquartile range, upper boundary, and lower boundary; treat data outside the upper and lower boundary intervals of each subset as outliers, remove outliers, and obtain the historical data after removing outliers;
[0121] Step 14: Normalize the historical data after removing outliers to obtain normalized historical data.
[0122] The other steps and parameters are the same as in Specific Implementation Method 1.
[0123] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that: in step one, historical data of each stage of the coking production process is collected; the specific process is as follows:
[0124] 1) Collect temperature data from January to October of the previous historical year from the Distributed Control System (DCS) during the coking production process. ;
[0125] Data from the coking production process is stored in the Distributed Control System (DCS), so data is retrieved from the DCS.
[0126] in, This represents the temperature data for January of the previous historical year (it can be a single temperature data point or multiple temperature data points). This represents the temperature data for February of the previous historical year. This represents the temperature data for March of the previous historical year. This represents the temperature data for April of the previous historical year. This represents the temperature data for May of the previous historical year. This represents the temperature data for June of the previous historical year. This represents the temperature data for July of the previous historical year. This represents the temperature data for August of the previous historical year. This represents the temperature data for September of the previous historical year. This represents the temperature data for October of the previous historical year.
[0127] 2) Collect pressure data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. ;
[0128] in, This represents the pressure data for January of the previous historical year (it can be one temperature data point or multiple temperature data points). This represents the stress data for February of the previous historical year. This indicates the stress data for March of the previous historical year. This represents the stress data for April of the previous historical year. This represents the stress data for May of the previous historical year. This represents the stress data for June of the previous historical year. This represents the stress data for July of the previous historical year. This represents the stress data for August of the previous historical year. This indicates the stress data for September of the previous historical year. This represents the stress data for October of the previous historical year;
[0129] 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. ;
[0130] in, This represents the flow data for January of the previous historical year (it can be one temperature data point or multiple temperature data points). This represents traffic data for February of the previous historical year. This represents traffic data for March of the previous historical year. This represents traffic data for April of the previous historical year. This represents traffic data for May of the previous historical year. This represents traffic data for June of the previous historical year. This represents traffic data for July of the previous historical year. This represents traffic data for August of the previous historical year. This represents traffic data for September of the previous historical year. This represents the traffic data for October of the previous historical year.
[0131] Other steps and parameters are the same as in specific implementation method one or two.
[0132] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that: in steps one and two, the collected historical data is divided into fixed time intervals (e.g., 2 months). Five (e.g.) subsets; the specific process is as follows:
[0133] 1) Collect temperature data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. Divided into fixed time intervals (e.g., 2 months) A subset of temperature data (e.g., 5 months) ;
[0134] in, This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of temperature data after being divided at fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of temperature data after being divided according to fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data;
[0135] 2) Collect pressure data from the distributed control system (DCS) of the previous historical year (January to October) for the coking production process. Divided into fixed time intervals (e.g., 2 months) Five (for example) subsets of stress data ;
[0136] in, This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data;
[0137] 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. Divided into fixed time intervals (e.g., 2 months) 5 subsets ;
[0138] in, This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data.
[0139] The other steps and parameters are the same as those in specific implementation methods one through three.
[0140] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One through Four in that: in steps one and three, the first quartile, third quartile, interquartile range, upper boundary, and lower boundary are calculated for each data subset; data outside the upper and lower boundary intervals of each subset are considered outliers, and outliers are removed to obtain historical data after outlier removal; the specific process is as follows:
[0141] 1) For a subset of temperature data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain the historical data after outlier removal; specifically:
[0142] subset First quartile , Subset The third quartile , Subset interquartile range , ;
[0143] subset upper boundary Subset lower boundary
[0144] express The 25th percentile of the data; express The 75th percentile of the data;
[0145] Representing a subset The upper boundary; Representing a subset The lower boundary;
[0146] 2) For a subset of stress data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain historical pressure data after outlier removal; specifically:
[0147] subset First quartile , Subset The third quartile , Subset interquartile range , ;
[0148] subset upper boundary Subset lower boundary
[0149] express The 25th percentile of the data; express The 75th percentile of the data;
[0150] Representing a subset The upper boundary; Representing a subset The lower boundary;
[0151] 3) For subsets of traffic data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. These outliers are removed to obtain the historical traffic data after outlier removal. Specifically:
[0152] subset First quartile Subset The third quartile Subset interquartile range ;
[0153] subset upper boundary Subset lower boundary
[0154] express The 25th percentile of the data; express The 75th percentile of the data;
[0155] Representing a subset The upper boundary; Representing a subset The lower boundary.
[0156] The other steps and parameters are the same as those in specific implementation methods one through four.
[0157] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that: in step one four, the historical data after removing outlier data is normalized to obtain normalized historical data, thus eliminating the influence of dimensions; the specific process is as follows:
[0158] The historical temperature data after removing outliers is normalized to obtain normalized historical temperature data, thus eliminating the influence of dimensions.
[0159] The historical stress data after removing outliers is normalized to obtain normalized historical stress data, thus eliminating the influence of dimensions.
[0160] The historical traffic data after removing outliers is normalized to obtain normalized historical traffic data, thus eliminating the influence of units.
[0161] The processed data is used to train a production environment simulation model.
[0162] The other steps and parameters are the same as those in specific implementation methods one through five.
[0163] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One to Six in that: in step two, an improved LSTM neural network model is constructed; the improved LSTM neural network model includes, in sequence: LSTM neural network model, ReLU activation function layer, Dropout layer, and Dense layer;
[0164] The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0165] The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time;
[0166] Train the improved LSTM neural network model to obtain a well-trained improved LSTM neural network model;
[0167] The specific process is as follows:
[0168] The historical data after the normalization process in step one is divided into a training set and a test set;
[0169] The training set is input into the improved LSTM neural network model, with mean squared error (MSE) as the loss function, and the Adam optimizer is used to optimize the parameters of the improved LSTM neural network model; the test set is input into the improved LSTM neural network model, and the output of the improved LSTM neural network model is compared with the actual values. When the coefficient reaches 0.98, a well-trained improved LSTM neural network model is obtained.
[0170] Dropout layers are regularization layers; Dense layers are fully connected layers.
[0171] LSTM layers handle sequence dependencies through gating mechanisms, resolving long-term dependency issues. The number of neuron units in the LSTM layer is set to 50, and the ReLU activation function is used.
[0172] A Dropout layer is introduced to prevent overfitting by randomly dropping neurons, thus improving the model's generalization ability. The dropout rate is set to 20%, meaning that 20% of the neurons are randomly masked in each training iteration. The Dropout layer operates on the output of the LSTM layer to reduce noise interference and ensure the model's robustness on the test set. A Dense layer is added as the output layer to map the features extracted by the LSTM to the final prediction target.
[0173] Furthermore, this model uses the Adam optimizer with adaptively adjusted learning rate. The loss function is the mean squared error function, which the optimizer minimizes during training to update the weights. The network structure of the production environment simulation model is as follows: Figure 2 As shown.
[0174] The data was divided into a training set (80%) and a test set (20%). Mean squared error (MSE) was used as the loss function, and the Adam optimizer was employed for parameter optimization. Finally, the weights of the LSTM neural network model trained from the data collected in step one were obtained, which were used to provide an environment for subsequent reinforcement learning.
[0175] The Adam optimizer was selected, and mean squared error was chosen as the loss function. This is suitable for regression problems and measures the difference between the model's predicted and actual values. 80% of the processed data was used as the training set, and 20% as the test set. The input and output formats were modified to meet the model's requirements, and the data was normalized before use. The model was trained using the training set and tested using the test set. The R² coefficients for the actual and predicted values of each output parameter were all around 0.98, indicating a good fit. The test results are shown in the figure below. Figure 4a , Figure 4b , Figure 4c , Figure 4d , Figure 4e As shown.
[0176] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0177] Specific Implementation Method Eight: This implementation method differs from one of Specific Implementation Methods One to Seven in that: in step three, the trained improved LSTM neural network model is used as the environment for the reinforcement learning control model;
[0178] The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate.
[0179] The actions taken by the intelligent agent are normalized control quantities, such as valve opening degree;
[0180] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0181] Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model;
[0182] The specific process is as follows:
[0183] The trained improved LSTM neural network model is used as the environment for the reinforcement learning control model;
[0184] The current state and the actions taken by the agent are inputs to a reinforcement learning control model;
[0185] The input states for the reinforcement learning control model are temperature, pressure, and flow rate;
[0186] The action taken by the agent inputting the reinforcement learning control model is the normalized valve opening of the flow control valve; the agent is the flow control valve.
[0187] Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent.
[0188] The action taken by the agent in the next moment, as output by the reinforcement learning control model, is converted into the valve opening of the flow control valve after inverse normalization.
[0189] The reward is calculated based on the next state output by the reinforcement learning control model.
[0190] The reinforcement learning control model is trained until the maximum number of iterations is reached, at which point a well-trained reinforcement learning control model is obtained.
[0191] In practical applications, the obtained intelligent agent model is deployed into the actual production system. Data is read from sensors in real time, and after undergoing the same normalization process, it is input into the model. The control actions output by the model are converted into actual valve control signals after inverse normalization, thereby achieving closed-loop control.
[0192] Define a state space, including key parameters for each process stage, such as temperature, pressure, and flow rate, and discretize continuous variables. Consider physical constraints in actual operation and define adjustable operating variables, such as valve opening. Design a reward function. The basic reward is set as the deviation between the control target and the setpoint; the penalty term is set as drastic changes in operating variables, exceeding the safe range, etc., comprehensively considering multiple objective weighting factors such as control accuracy.
[0193] The reinforcement learning control model is trained until the maximum number of iterations is reached, at which point a well-trained reinforcement learning control model is obtained; the specific process is as follows:
[0194] n_steps=2048, batch_size=64, n_epochs=10, learning_rate=3e-4, etc.
[0195] The other steps and parameters are the same as those in specific implementation methods one through seven.
[0196] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One through Eight in that: the reward is calculated based on the next-moment state output by the reinforcement learning control model; the specific process is as follows:
[0197] like ,award Serious abnormality, severe punishment;
[0198] like ,award The optimal range yields the highest reward.
[0199] like ,award Acceptable range, moderate reward;
[0200] other: For deviations from the range, the penalty is proportional to the deviation.
[0201] The other steps and parameters are the same as those in specific implementation methods one through eight.
[0202] Specific Implementation Method 10: This implementation method provides an intelligent control system for a coking process based on reinforcement learning, which is used to execute an intelligent control method for a coking process based on reinforcement learning.
[0203] The beneficial effects of the present invention are verified using the following embodiments:
[0204] Example 1:
[0205] This invention proposes a two-layer modeling method based on a reinforcement learning-based intelligent control method for coking processes: a production environment simulation model and a reinforcement learning control model. To verify its effectiveness, this invention collected environmental parameters affecting the top temperature of the benzene stripping tower from Company A from January to October 2024: top pressure, bottom temperature, bottom pressure, approximate crude benzene temperature, and crude benzene reflux rate. First, data preprocessing was performed on each environmental parameter by month using a box plot-based outlier removal method, deleting a total of 550,295 data sets, leaving 4,184,310 data sets. The results are as follows: Figure 3a , Figure 3b , Figure 3c , Figure 3d , Figure 3e , Figure 3f .
[0206] A two-layer LSTM neural network model is constructed, adding an LSTM layer, a Dropout layer, and finally a Dense output layer. The LSTM layer captures the temporal features of the time series data, while the Dense layer provides the final prediction result. The Adam optimizer is selected, and mean squared error is chosen as the loss function, suitable for regression problems, to measure the difference between the model's predicted and actual values. 80% of the processed data is used as the training set, and 20% as the test set. The input and output formats are modified to meet the model's requirements, and the data is normalized before use. The model is trained using the training set and tested using the test set. The actual and predicted values of each output parameter are compared. The coefficients are all around 0.98, indicating a good fit. The test results are shown in the figure below. Figure 4a , Figure 4b , Figure 4c , Figure 4d , Figure 4e As shown.
[0207] Based on the aforementioned production environment simulation model, a simulation environment capable of interacting with a reinforcement learning model was constructed. The goal of this reinforcement learning model is to train a model that, by controlling the crude benzene reflux flow rate, enables the column top temperature to respond rapidly within the simulation environment and maintain a temperature range between 93°C and 98°C.
[0208] Without model training, the crude benzene reflux flow rate was controlled by setting random numbers within a reasonable range, and the temperature change curve at the top of the column was obtained as follows. Figure 5 The temperature change at the top of the tower in the figure is relatively unstable.
[0209] In reinforcement learning schemes, a reinforcement learning model can be trained by defining a reward policy, and the model's performance is as follows: Figure 6 After applying reinforcement learning, the temperature can quickly approach a reasonable temperature range and remain stable.
[0210] Rewards are key to guiding the learning of intelligent agents. The model constructed using the top temperature of a benzene removal tower as an example is shown below:
[0211] like ,award Serious abnormality, severe punishment;
[0212] like ,award The optimal range yields the highest reward.
[0213] like ,award Acceptable range, moderate reward;
[0214] other: For deviations from the range, the penalty is proportional to the deviation.
[0215] To study the feasibility of reinforcement learning models, we used actual collected data, assuming the current data is... At that time, the red curve corresponds to... Real-time tower top temperature data.
[0216] Will The environmental data at any given time is input into the reinforcement learning model. The reinforcement learning model outputs an action, which is applied to the simulated environment. The simulated environment then outputs... The environmental values at each time point are shown, with the tower top temperature curve represented by the green curve. Assuming the simulated environment can well characterize the actual production environment, then from... Figure 7 As can be seen, compared with the current temperature control method, the reinforcement learning model can make the temperature at the top of the tower more stable and keep it within a reasonable temperature range.
[0217] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A method for intelligent control of a coking process based on reinforcement learning, characterized in that: The specific process of the method is as follows: Step 1: Collect historical data of the coking production process, and remove outliers from the historical data based on the box plot outlier detection method to obtain the historical data after removing outliers; The historical data after removing outliers is normalized to obtain normalized historical data. Historical data on the coking production process includes temperature, pressure, and flow rate; Step 2: Construct an improved LSTM neural network model; The improved LSTM neural network model includes, in sequence: LSTM neural network model, ReLU activation function layer, Dropout layer, and Dense; The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time; The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time; The improved LSTM neural network model is trained based on the historical data after normalization in step one, and the trained improved LSTM neural network model is obtained. Step 3: Use the trained improved LSTM neural network model as the environment for the reinforcement learning control model; The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate; the action taken by the agent is a normalized control variable. Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent. Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model; Step 4: Collect online data of the coking production process, and remove outliers from the online data using the box plot outlier detection method to obtain the online data after removing outliers; The online data after removing outliers is normalized to obtain normalized online data. Online data for the coking production process includes temperature, pressure, and flow rate; The normalized online data is input into the trained reinforcement learning control model, which then outputs the temperature at the next moment.
2. The intelligent control method for coking process based on reinforcement learning according to claim 1, characterized in that: In step one, historical data of the coking production process is collected, and outlier data is removed from the historical data based on the box plot outlier detection method to obtain the historical data after removing outlier data. The historical data, after removing outliers, is normalized to obtain normalized historical data. Historical data for the coking production process includes temperature, pressure, and flow rate. The specific process is as follows: Step 1: Collect historical data on the coking production process; Steps 1 and 2: Divide the collected historical data into fixed time intervals. A subset; Step 13: For each subset, calculate the first quartile, third quartile, interquartile range, upper boundary, and lower boundary; treat data outside the upper and lower boundary intervals of each subset as outliers, remove outliers, and obtain the historical data after removing outliers; Step 14: Normalize the historical data after removing outliers to obtain normalized historical data.
3. The intelligent control method for coking process based on reinforcement learning according to claim 2, characterized in that: The step described above involves collecting historical data on the coking production process; the specific process is as follows: 1) Collect temperature data from January to October of the previous historical year from the distributed control system (DCS) during the coking production process. ; in, This represents the temperature data for January of the previous historical year. This represents the temperature data for February of the previous historical year. This represents the temperature data for March of the previous historical year. This represents the temperature data for April of the previous historical year. This represents the temperature data for May of the previous historical year. This represents the temperature data for June of the previous historical year. This represents the temperature data for July of the previous historical year. This represents the temperature data for August of the previous historical year. This represents the temperature data for September of the previous historical year. This represents the temperature data for October of the previous historical year. 2) Collect pressure data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. ; in, This represents the stress data for January of the previous historical year. This represents the stress data for February of the previous historical year. This represents the stress data for March of the previous historical year. This represents the stress data for April of the previous historical year. This represents the stress data for May of the previous historical year. This represents the stress data for June of the previous historical year. This represents the stress data for July of the previous historical year. This indicates the stress data for August of the previous historical year. This indicates the stress data for September of the previous historical year. This represents the stress data for October of the previous historical year; 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. ; in, This represents traffic data for January of the previous historical year. This represents traffic data for February of the previous historical year. This represents traffic data for March of the previous historical year. This represents traffic data for April of the previous historical year. This represents traffic data for May of the previous historical year. This represents traffic data for June of the previous historical year. This represents traffic data for July of the previous historical year. This represents traffic data for August of the previous historical year. This represents traffic data for September of the previous historical year. This represents the traffic data for October of the previous historical year.
4. The intelligent control method for coking process based on reinforcement learning according to claim 3, characterized in that: In steps one and two, the collected historical data is divided into segments according to fixed time intervals. A subset; the specific process is as follows: 1) Collect temperature data from the distributed control system (DCS) of the previous historical year, covering the period from January to October, during the coking production process. Divided into according to fixed time intervals A subset of temperature data ; in, This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of temperature data after being divided at fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of temperature data after being divided according to fixed time intervals; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data; This indicates the temperature data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of temperature data; 2) Collect pressure data from the distributed control system (DCS) of the previous historical year (January to October) for the coking production process. Divided into according to fixed time intervals A subset of stress data ; in, This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of pressure data after being divided at fixed time intervals; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data; This indicates the pressure data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of stress data; 3) Collect flow data from the distributed control system (DCS) of the previous historical year, from January to October, during the coking production process. Divided into according to fixed time intervals Subset ; in, This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The second subset of traffic data after being divided at fixed time intervals; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data; This indicates the flow data collected from the distributed control system (DCS) during the coking production process from January to October of the previous historical year. The first, divided according to a fixed time interval A subset of traffic data.
5. The intelligent control method for coking process based on reinforcement learning according to claim 4, characterized in that: In steps one and three, the first quartile, third quartile, interquartile range, upper boundary, and lower boundary are calculated for each data subset; data outside the upper and lower boundary intervals of each subset are considered outliers, and these outliers are removed to obtain the historical data after outlier removal; the specific process is as follows: 1) For a subset of temperature data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain the historical data after outlier removal; specifically: subset First quartile , ; subset The third quartile , ; subset interquartile range , ; subset upper boundary ; subset lower boundary express The 25th percentile of the data; express The 75th percentile of the data; Representing a subset The upper boundary; Representing a subset The lower boundary; 2) For a subset of stress data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. Outliers are removed to obtain historical pressure data after outlier removal; specifically: subset First quartile , ; subset The third quartile , ; subset interquartile range , ; subset upper boundary ; subset lower boundary express The 25th percentile of the data; express The 75th percentile of the data; Representing a subset The upper boundary; Representing a subset The lower boundary; 3) For subsets of traffic data Each subset Calculate subsets First quartile, subset The third quartile, subset interquartile range, subset upper boundary and subset The lower boundary; for each subset Data outside the upper and lower boundary intervals are considered outliers. These outliers are removed to obtain the historical traffic data after outlier removal. Specifically: subset First quartile ; subset The third quartile Subset interquartile range ; subset upper boundary ; subset lower boundary express The 25th percentile of the data; express The 75th percentile of the data; Representing a subset The upper boundary; Representing a subset The lower boundary.
6. The intelligent control method for coking process based on reinforcement learning according to claim 5, characterized in that: In step one four, the historical data after removing outlier data is normalized to obtain normalized historical data; the specific process is as follows: The historical temperature data after removing outliers is normalized to obtain the normalized historical temperature data. The historical stress data after removing outliers is normalized to obtain the normalized historical stress data. The historical traffic data after removing outliers is normalized to obtain the normalized historical traffic data.
7. The intelligent control method for coking process based on reinforcement learning according to claim 6, characterized in that: In step two, an improved LSTM neural network model is constructed. The improved LSTM neural network model includes, in sequence: LSTM neural network model, ReLU activation function layer, Dropout layer, and Dense layer; the Dropout layer is a regularization layer; the Dense layer is a fully connected layer. The input to the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time; The output of the improved LSTM neural network model is: Temperature, pressure, and flow rate at any given time; Train the improved LSTM neural network model to obtain a well-trained improved LSTM neural network model; The specific process is as follows: The historical data after the normalization process in step one is divided into a training set and a test set; The training set is input into the improved LSTM neural network model, with mean squared error as the loss function, and the Adam optimizer is used to optimize the parameters of the improved LSTM neural network model; the test set is input into the improved LSTM neural network model, and the output of the improved LSTM neural network model is compared with the actual values. When the coefficient reaches 0.98, a well-trained improved LSTM neural network model is obtained.
8. The intelligent control method for coking process based on reinforcement learning according to claim 7, characterized in that: In step three, the trained improved LSTM neural network model is used as the environment for the reinforcement learning control model. The inputs to the reinforcement learning control model are the state and the action taken by the agent; the state is temperature, pressure, and flow rate; the action taken by the agent is a normalized control variable. Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent. Train the reinforcement learning control model to obtain a well-trained reinforcement learning control model; The specific process is as follows: The trained improved LSTM neural network model is used as the environment for the reinforcement learning control model; The current state and the actions taken by the agent are inputs to a reinforcement learning control model; The input states for the reinforcement learning control model are temperature, pressure, and flow rate; The action taken by the agent inputting the reinforcement learning control model is the normalized valve opening degree of the flow control valve; The intelligent agent is a flow control valve; Reinforcement learning controls the model to output the state of the agent in the next time step and the actions to be taken by the agent. The action taken by the agent in the next moment, as output by the reinforcement learning control model, is converted into the valve opening of the flow control valve after inverse normalization. The reward is calculated based on the next state output by the reinforcement learning control model. The reinforcement learning control model is trained until the maximum number of iterations is reached, at which point a well-trained reinforcement learning control model is obtained.
9. The intelligent control method for coking process based on reinforcement learning according to claim 8, characterized in that: The reward is calculated based on the next-moment state output by the reinforcement learning control model; the specific process is as follows: like ,award ;like ,award ;like ,award ; other: .
10. An intelligent control system for a coking process based on reinforcement learning, characterized in that, The system is used to execute the intelligent control method for coking process based on reinforcement learning as described in any one of claims 1 to 9.