Reboiler switching control method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202311853059.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2043-12-28
AI Technical Summary
[0004]然而,前馈技术通常需要建立复杂的数学模型,这不仅造成了高昂的成本,而且还限制了其在不同场景之间的灵活性和可迁移性
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the reboiler switching control method according to any embodiment of the present invention.
Smart Images

Figure CN117797489B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chemical equipment technology, and in particular to a reboiler switching control method, device, electronic equipment and storage medium. Background Technology
[0002] Refining towers are mainly used to extract high-quality fuel oil and lubricating oil from crude oil. Their operating principle primarily utilizes the differences in boiling points of different components in the crude oil for separation. To ensure the separated products meet quality requirements, a hot water reboiler can be put into operation once the refining tower system is operating stably. During the process of putting the hot water reboiler into operation, the steam reboiler should be gradually shut down to reduce steam consumption.
[0003] In related technologies, feedforward technology is commonly used to control hot water reboilers and steam reboilers. Feedforward technology uses a PID controller to control the opening degree of the reboiler.
[0004] However, feedforward techniques typically require the establishment of complex mathematical models, which not only incurs high costs but also limits their flexibility and portability across different scenarios. If the system's dynamic model changes, previously tuned PID parameters become unusable. Furthermore, most PID parameter adjustments at present are based on manual tuning or expert systems, both of which rely heavily on manual parameter selection and domain knowledge, significantly reducing the ease of use, flexibility, and versatility of PID controllers. Summary of the Invention
[0005] This invention provides a reboiler switching control method, device, electronic equipment, and storage medium to achieve accurate control of the opening changes of the hot water reboiler and steam reboiler during the operation of the refining tower system. This achieves the effect of reducing labor costs and ensuring the stability of the refining tower system, while also improving the switching control efficiency between the hot water reboiler and steam reboiler.
[0006] According to one aspect of the present invention, a reboiler switching control method is provided, the method comprising:
[0007] During the operation of the purification tower system to be controlled, the operation process data corresponding to the purification tower system to be controlled is acquired. The operation process data includes the tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the purity value sequence of each separated object, the first opening value sequence of the hot water reboiler, the second opening value sequence of the steam reboiler, and the liquid flow rate value sequence of each target pipeline.
[0008] The operating process data is processed based on the pre-trained reboiler opening prediction model to obtain the first opening change of the hot water reboiler at the current time and the second opening change of the steam reboiler at the current time; wherein, the reboiler opening prediction model is trained based on a reinforcement learning algorithm.
[0009] The first opening change and the second opening change are processed based on a preset opening value processing method.
[0010] According to another aspect of the present invention, a reboiler switching control device is provided, the device comprising:
[0011] The process data acquisition module is used to acquire the operation process data corresponding to the purification tower system under control during the operation of the purification tower system under control. The operation process data includes the tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the purity value sequence of each separated object, the first opening value sequence of the hot water reboiler, the second opening value sequence of the steam reboiler, and the liquid flow rate value sequence of each target pipeline.
[0012] The reboiler opening change determination module is used to process the operation process data based on the pre-trained reboiler opening prediction model to obtain the first opening change of the hot water reboiler at the current time and the second opening change of the steam reboiler at the current time; wherein, the reboiler opening prediction model is trained based on a reinforcement learning algorithm.
[0013] The opening change processing module is used to process the first opening change and the second opening change based on a preset opening value processing method.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the reboiler switching control method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the reboiler switching control method according to any embodiment of the present invention.
[0019] The technical solution of this invention acquires the operation process data corresponding to the purification tower system during its operation. Further, based on a pre-trained reboiler opening prediction model, the operation process data is processed to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment. Finally, the first and second opening changes are processed based on a preset opening value processing method. This solves the problems of inaccurate prediction of the reboiler opening and complex control processes in related technologies. It achieves accurate control of the opening changes of the hot water and steam reboilers during the operation of the purification tower system, thereby reducing labor costs, ensuring the stability of the purification tower system, and improving the switching control efficiency between the hot water and steam reboilers.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a reboiler switching control method provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart of a reboiler switching control method provided in Embodiment 2 of the present invention;
[0024] Figure 3 This is a schematic diagram of a reboiler switching control device according to Embodiment 3 of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the reboiler switching control method of the present invention. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Example 1
[0029] Figure 1 This is a flowchart of a reboiler switching control method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where the opening degree of the reboiler is switched and controlled during the operation of a purification column system. This method can be executed by a reboiler switching control device, which can be implemented in hardware and / or software, and can be configured in a terminal and / or server. Figure 1 As shown, the method includes:
[0030] S110. During the operation of the purification tower system to be controlled, acquire the operation process data corresponding to the purification tower system to be controlled.
[0031] It should be noted that the technical solution of this invention can be applied to the scenario of reboiler switching control in a refining tower system. During stable operation of the refining tower system, the hot water reboiler is completely closed, while the steam reboiler is fully open. To improve the product separation efficiency and reduce plant production costs while ensuring continuous and stable operation of the refining tower system, the hot water reboiler can be put into use during stable operation. That is, the hot water reboiler can be gradually turned on while the steam reboiler is gradually turned off until both are fully open and completely closed. This ensures stable operation of the refining tower system and that the separated product meets quality requirements.
[0032] In this embodiment, the purification tower system to be controlled can be understood as the purification tower system to be controlled. The purification tower system can be a system for the precise separation of gasoline and hydrocarbon compounds. The purification tower system may include a purification tower, at least one reboiler, at least one pipeline, and other equipment. The purification tower is typically a vertical columnar container filled with packing material, with gaps between the packing. The function of the purification tower is to carry out processes such as absorption, adsorption, and precipitation of substances within the tower, separating different components in the mixture through the action of the packing material. The reboiler is a device that re-vaporizes the liquid. In the reboiler, the material expands or even vaporizes due to heat, its density decreases, and it leaves the vaporization space, smoothly returning to the tower. The gas and liquid phases returning to the tower rise through the trays, while the liquid phase falls to the bottom. The pipeline in the purification tower system can be a pipeline connected to the purification tower. These pipelines may include pipelines that can both extract liquid from the tower and release liquid back into the tower, or pipelines that can only extract liquid from the tower. Operational process data can be understood as the data generated during the operation of the purification tower system under control. Operational process data can include the operational data corresponding to each device in the purification tower system. Operational process data can be sequential data with each data acquisition time as a node. Specifically, the operational process data includes the sequence of tower bottom temperature values at the current time and at least one historical time prior to the current time, the sequence of purity values for each separated object, the sequence of the first opening value of the hot water reboiler, the sequence of the second opening value of the steam reboiler, and the sequence of liquid flow rates for each target pipeline.
[0033] In this embodiment, the entire operation of the purification tower system under control can be divided into multiple data acquisition times based on a pre-set data acquisition frequency. These data acquisition times can then be considered as individual times. Furthermore, for each time point, once the current time is determined, process data can be collected to obtain the corresponding process data value. The reboiler temperature value sequence can include the reboiler temperature values corresponding to each time point. The reboiler temperature can be understood as the temperature of the reboiler in the purification tower. The separation object can be understood as the object separated from the mixed liquid. For example, assuming the mixed liquid to be separated is crude oil, the separation object can be fuel oil, solvent oil, lubricating oil, lubricating grease, paraffin wax, asphalt, and liquefied petroleum gas, etc. The purity value sequence can include the purity values corresponding to each time point. The purity value can be the purity corresponding to the corresponding category object. The hot water reboiler can be a reboiler installed in the purification tower system under control. The first opening value sequence can include the first opening value corresponding to each time point. The first opening value is the angle value of the hot water reboiler valve opening. The steam reboiler can be a reboiler installed in the purification column system to be controlled. The second opening value sequence can include the second opening value corresponding to each time moment. The second opening value is the angle at which the steam reboiler valve is opened. The target pipeline can be a pipeline in the purification column system to be controlled that is associated with the reboiler switching control process. The target pipeline can be at least a portion of all pipelines installed in the purification column system to be controlled. The liquid flow rate value sequence can include the liquid flow rate corresponding to each time moment.
[0034] In practical applications, data analysis can be performed on the operation of the purification tower system to be controlled, so that the opening of the hot water reboiler and steam reboiler in the system can be controlled based on the data analysis results. Therefore, during the operation of the purification tower system, operating data at various times can be collected, thereby obtaining the corresponding operating process data of the purification tower system.
[0035] S120. Based on the pre-trained reboiler opening prediction model, process the operation process data to obtain the first opening change of the hot water reboiler at the current moment and the second opening change of the steam reboiler at the current moment.
[0036] The reboiler opening prediction model is trained based on a reinforcement learning algorithm.
[0037] In this embodiment, the reboiler opening prediction model can be understood as a policy network determined by a reinforcement learning algorithm. Those skilled in the art will understand that the policy network can be a deep neural network model that deploys policy functions, where the input object can be the state of the agent's environment, and the output object can be the decision action determined based on the input state. The reboiler opening prediction model can be a neural network model of any structure. Optionally, the model structure of the reboiler opening prediction model can include, but is not limited to, a gated recurrent neural network (GRU) and a long short-term memory network (LSTM). The first opening change can be understood as the increase or decrease in the first opening value, that is, the increase or decrease in the opening value based on the first opening value at the current time. Similarly, the second opening change can be understood as the increase or decrease in the second opening value, that is, the increase or decrease in the opening value based on the first opening value at the current time.
[0038] In practical applications, to predict the valve opening values of hot water reboilers and steam reboilers, after obtaining the operational data of the purification tower system to be controlled, the acquired operational data can be input into the reboiler opening prediction model. Then, based on the reboiler opening prediction model, the operational data can be processed to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment. It should be noted that the first and second opening changes can be quantities with both direction and magnitude, and the direction can be represented by "+" and "-".
[0039] It should be noted that during reboiler switching, the direction of reboiler parameter changes must be monotonic. The valve opening of the hot water reboiler only increases and never decreases, while the valve opening of the steam reboiler only decreases and never increases. In other words, the direction of change of the first opening for the hot water reboiler and the direction of change of the second opening for the steam reboiler are opposite. That is, if the hot water reboiler is a hot water reboiler, the first opening value increases, and the change is an increase; conversely, if the steam reboiler is a steam reboiler, the second opening value decreases, and the change is a decrease.
[0040] S130. Process the first opening change and the second opening change based on the preset opening value processing method.
[0041] In this embodiment, the preset opening value processing method can be a pre-set method that processes the opening change determined based on the model. The preset opening value processing method can be any processing method, and can be either indirect control with visual display or direct control, etc.
[0042] In practical applications, after obtaining the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment, the first opening change and the second opening change can be processed according to the preset opening value processing method.
[0043] In this embodiment, the preset opening value processing method can include multiple methods. For different processing methods, the corresponding opening change processing process is different. The following will describe these opening value processing methods respectively.
[0044] Optionally, the preset opening value processing method includes indirect control and visual display, processing the first opening change and the second opening change based on the preset data processing method, including: visually displaying the first opening change and the second opening change based on the display interface of the target terminal.
[0045] In this embodiment, the target terminal can be understood as a terminal device used to process operational data, a device used to deploy a reboiler opening prediction model, or a terminal device belonging to the target user; this embodiment does not specifically limit this. The target user can be a technician used to monitor the purification tower system under control. Optionally, the target terminal can be a mobile terminal or a PC, etc.
[0046] In practical applications, after obtaining the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment, the first and second opening changes can be visualized on the target terminal's display interface.
[0047] Optionally, the preset opening value processing method includes direct control, which processes the first opening change and the second opening change based on a preset data processing method, including: controlling the hot water reboiler to adjust its opening based on the first opening change and obtaining the first opening value of the hot water reboiler at the next moment; and controlling the steam reboiler to adjust its opening based on the second opening change and obtaining the second opening value of the steam reboiler at the next moment.
[0048] In this embodiment, direct control can be understood as directly adjusting the opening of the reboiler valve, that is, adjusting the opening of the reboiler valve without human intervention.
[0049] In practical applications, after obtaining the first change in the opening degree of the hot water reboiler at the current moment, the opening degree of the hot water reboiler can be adjusted based on this first change in opening degree. Furthermore, the valve opening of the hot water reboiler is adjusted according to the direction and magnitude of the change in the first opening degree. Further, the first opening degree value of the hot water reboiler at the next moment can be obtained. Similarly, after obtaining the second change in the opening degree of the steam reboiler at the current moment, the opening degree of the steam reboiler can be adjusted based on this second change in opening degree. Furthermore, the valve opening of the steam reboiler is adjusted according to the direction and magnitude of the change in the second opening degree. Further, the second opening degree value of the steam reboiler at the next moment can be obtained.
[0050] It should be noted that, when the preset opening value processing method is direct control, before adjusting the opening of the corresponding reboiler based on the opening change, it can be determined whether the opening value obtained after the adjustment reaches the preset control threshold. If not, the opening of the corresponding reboiler can be directly adjusted based on the opening change; if it does, the preset opening value processing method can be switched to indirect control, thereby ensuring the system stability of the purification column system under control.
[0051] Based on this, and in addition to the above technical solutions, the method further includes: when the preset opening value processing method is direct control, determining the first opening value of the hot water reboiler at the current moment based on the first opening value sequence, and determining the first opening value of the hot water reboiler at the next moment based on the first opening change and the first opening value; and determining the second opening value of the steam reboiler at the current moment based on the second opening value sequence, and determining the second opening value of the steam reboiler at the next moment based on the second opening change and the second opening value; and when it is detected that the first opening value at the next moment reaches the first preset control threshold, and / or the second opening value at the next moment reaches the second preset control threshold, switching the preset opening value processing method to indirect control and displaying it visually.
[0052] The technical solution of this invention acquires the operation process data corresponding to the purification tower system during its operation. Further, based on a pre-trained reboiler opening prediction model, the operation process data is processed to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment. Finally, the first and second opening changes are processed based on a preset opening value processing method. This solves the problems of inaccurate prediction of the reboiler opening and complex control processes in related technologies. It achieves accurate control of the opening changes of the hot water and steam reboilers during the operation of the purification tower system, thereby reducing labor costs, ensuring the stability of the purification tower system, and improving the switching control efficiency between the hot water and steam reboilers.
[0053] Example 2
[0054] Figure 2 This is a flowchart of a reboiler switching control method provided in Embodiment 2 of the present invention. Based on the aforementioned embodiments, before processing the operating process data based on the reboiler opening prediction model, simulation training samples can be constructed, and the reboiler opening prediction model can be trained based on the simulation training samples. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or similar to those in the above embodiments will not be repeated here.
[0055] like Figure 2 As shown, the method includes:
[0056] S210. For each training round, obtain the online simulation training samples of the purification tower system to be controlled in the current training round.
[0057] The online training samples are determined based on the simulation environment model corresponding to the purification tower system to be controlled.
[0058] In this embodiment, the simulation environment model can be understood as a model representing the operation of the purification tower system to be controlled, or as a model obtained after modeling the actual operating environment of the purification tower system to be controlled. Environment modeling can be a simulation modeling method in the field of reinforcement learning, which refers to abstracting the real environment into a computable model so that the agent can understand and predict the process. The environment model can include a state transition function (state transition model) and a reward function. In this embodiment, the simulation environment model can be a model established to better plan the operation of the purification tower system to be controlled. A training round can be understood as a training process constructed from the initial state of the simulation environment model until the maximum number of training rounds is reached. A training round can be understood as one training process. After each training round, the simulation environment model can be initialized and returned to the initial state for the next round of training. Then, after training for multiple rounds and detecting that the loss function in the model converges, the trained model can be obtained. In this embodiment, the initial state can be that the purification tower system to be controlled is running smoothly, the hot water reboiler is closed, and the steam reboiler is fully open.
[0059] In this embodiment, the online training samples can be understood as sample data collected in real time during the operation of the simulation environment model corresponding to the purification tower system to be controlled. In other words, the data included in the online training samples is online data. Specifically, the online training samples include the sample tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the sample purity value sequence of each separation object, the first sample opening value sequence of the hot water reboiler, the second sample opening value sequence of the steam reboiler, the sample liquid flow rate value sequence of each target pipeline, and the historical reward feedback information sequence corresponding to at least one historical time.
[0060] In this embodiment, the sample column reboiler temperature value sequence may include the sample column reboiler temperature value corresponding to each time point. The sample column reboiler temperature value is the reboiler temperature of the purification column in the simulation environment model. The sample purity value sequence may include the sample purity value corresponding to each time point. The sample purity value is the purity value of the corresponding separated object in the simulation environment model. The first sample opening value sequence may include the first sample opening value corresponding to each time point. The first sample opening value is the opening value of the hot water reboiler in the simulation environment model. The second sample opening value sequence may include the second sample opening value corresponding to each time point. The second sample opening value is the opening value of the steam reboiler in the simulation environment model. The sample liquid flow rate value sequence may include the sample liquid flow rate value corresponding to each time point. The sample liquid flow rate value is the flow rate value of the liquid in the corresponding target pipeline in the simulation environment model. The historical reward feedback information sequence may include reward feedback information corresponding to multiple time points. Reward feedback information is the numerical value obtained by the agent after performing an action, which can characterize the quality of the action. The reward feedback information can be determined based on the reward function deployed in the simulation environment model. Those skilled in the art will understand that in the field of reinforcement learning, reinforcement learning algorithms can be used to train a policy network so that the trained policy network can judge the current state of the agent to obtain the target decision action at the current time. For example, a flexible actor-critic algorithm can be used to train the policy network. This algorithm can include a judge network (e.g., a state value network and / or an action state value network) and a policy network. The judge network acts as a "critic," not directly taking an action but evaluating its quality; the policy network acts as an "actor," determining the decision action based on the input state. The policy network is essentially a neural network model that can directly predict the most appropriate policy to execute by observing the environmental state. Executing this policy yields the maximum expected reward, which can be used as the reward. For example, during the training of the policy network, the state s at time t is obtained. t , will s t The input is fed into the policy network, which can then determine the decision action 'a' based on the input state. t , for s t Implement decision-making action a t The new state is obtained, which is the state s corresponding to time t+1. t+1 And obtain the reward r corresponding to time t. t Furthermore, based on the reward r t The parameters of the evaluation network and policy network in reinforcement learning can be adjusted to obtain the trained policy network.
[0061] In practical applications, scenario initialization is first performed to ensure the simulation environment model corresponding to the purification tower system under control is in its initial state, i.e., the system is running smoothly, the hot water reboiler is closed, and the steam reboiler is fully open. Further, during the simulation environment model's operation, the generated operational data is collected. This allows for the acquisition of sample tower bottom temperatures at various times, and the construction of a sample tower bottom temperature sequence based on these collected values. Simultaneously, the sample purity values for each separated object at various times are obtained, along with a sequence of sample purity values for each separated object. Furthermore, the first sample opening value of the hot water reboiler at various times is obtained, along with a sequence of the first sample opening value for the hot water reboiler. Similarly, the second sample opening value of the steam reboiler at various times is obtained, along with a sequence of the second sample opening value for the steam reboiler. Finally, the sample liquid flow rate values for each target pipeline at various times are also obtained, along with a sequence of the sample liquid flow rate values for each target pipeline.
[0062] Furthermore, regarding the interaction between the simulation environment model and the hot water reboiler and steam reboiler, after the hot water reboiler and steam reboiler perform an action of changing their opening value, reward feedback information corresponding to that action can be obtained based on the reward function pre-deployed in the simulation environment model. Then, a historical reward feedback information sequence can be constructed based on the reward feedback information corresponding to each time point.
[0063] S220. A reboiler opening prediction model is trained based on various online training samples from simulations and reinforcement learning algorithms.
[0064] In practical applications, the reboiler opening prediction model can be trained based on the online training samples and reinforcement learning algorithms corresponding to each training round, so that after multiple training rounds, the reboiler opening prediction model can be obtained as a final trained model.
[0065] Optionally, a reboiler opening prediction model is trained based on each online simulation training sample and a reinforcement learning algorithm, including: for each training round, inputting the sample tower temperature value sequence, each sample purity value sequence, the first sample opening value sequence, the second sample opening value sequence, and each sample liquid flow rate value sequence from the online simulation training samples corresponding to the current training round into the reboiler opening prediction model to obtain the probability distribution of the actual changes in the hot water reboiler and the steam reboiler at the current moment; for each first opening change and each second opening change in the probability distribution of the actual changes, inputting the current first opening change, the current second opening change, the sample tower temperature value sequence, each sample purity value sequence, the first sample opening value sequence, the second sample opening value sequence, and each sample liquid flow rate value sequence into the state-action value model to obtain the expected reward corresponding to the current moment; based on the probability distribution of the actual changes... The reboiler opening prediction model is updated using various expected rewards and strategy gradient algorithms. Based on the probability distribution of actual changes, the first actual opening change of the hot water reboiler and the second actual opening change of the steam reboiler at the current moment are determined. The sample tower bottom temperature sequence, the purity sequence of each sample, the first sample opening sequence, the second sample opening sequence, the liquid flow rate sequence of each sample, the first actual opening change, and the second actual opening change are processed according to a pre-set multi-index reward function to determine the reward feedback information corresponding to the purification tower system to be controlled at the current moment, and the historical reward feedback information sequence is updated based on the reward feedback information. The state-action value model is updated using the reward feedback information and a time-series difference algorithm. Training ends when the preset training objective corresponding to the reboiler opening prediction model is reached, resulting in the reboiler opening prediction model.
[0066] The probability distribution of actual changes includes the probability information of the first opening changes corresponding to the hot water reboiler and the probability information of the second opening changes corresponding to the steam reboiler. The state-action value network can be understood as a deep neural network that takes the current state and the current decision action as inputs to evaluate the decision action taken in the current state. The state-action value network can be a neural network including a state-action value function. The input to the state-action value network can be the current state and the current decision action, and the output can be the value corresponding to taking the decision action in the current state, that is, the expected reward corresponding to the decision action at the current moment. The policy gradient algorithm is a class of algorithms for solving reinforcement learning problems. It is a gradient-based optimization algorithm that helps machine learning models optimize in decision-making environments to obtain the best results. The idea of the policy gradient algorithm is to first represent the policy as a continuous function related to the reward, and then use the optimization method of the continuous function to find the optimal policy. The optimization objective is to maximize the continuous function. Multi-index reward functions include the observation index safety reward function, the observation index stability reward function, and the control index reward function. Optionally, the multi-index reward function can be determined by weighted summation of the observed index safety reward function, observed index stability reward function, and control index reward function. Observed indices include sample column temperature, sample purity, first sample opening value, second sample opening value, and sample liquid flow rate. Control indices include the first and second actual opening changes. Temporal Difference (TD) algorithm is an algorithm used to estimate the value function of a policy, which can adaptively adjust the policy by learning the difference between the current state and future states. The core idea of the TD algorithm is the update of the state value function. The preset training objective can be a pre-set termination condition for the policy network training process. Optionally, the preset training objective can include maximizing the objective function value corresponding to the policy gradient algorithm or reaching a preset number of training iterations.
[0067] In practical applications, for each training round, the model parameters in the reboiler opening prediction model and the state-action value model can be updated based on the above process. This results in a fully trained reboiler opening prediction model.
[0068] In practical applications, for each training round, the sequence of sample tower temperature, purity, first sample opening, second sample opening, and liquid flow rate from the online training samples corresponding to the current training round can be input into the reboiler opening prediction model to obtain the probability distribution of actual changes in the hot water reboiler and steam reboiler at the current moment. Then, for each first and second opening change in the probability distribution of actual changes, the current first and second opening changes, the sample tower temperature, purity, first and second opening opening sequences, and liquid flow rate sequences are input into the state-action value model to obtain the expected reward corresponding to the current moment. Furthermore, based on the probability distribution of actual changes, the expected rewards, and the policy gradient algorithm, the parameters of the reboiler opening prediction model are updated.
[0069] Furthermore, the probability distribution of actual changes can be sampled, and the first actual opening change of the hot water reboiler and the second actual opening change of the steam reboiler at the current moment can be determined from the first opening change and the second opening change included in the probability distribution of actual changes.
[0070] Furthermore, based on a pre-set reward function, the sequence of sample column bottom temperature values, the sequence of purity values for each sample, the sequence of opening values for the first sample, the sequence of opening values for the second sample, the sequence of liquid flow rates for each sample, the change in the first actual opening, and the change in the second actual opening are processed to determine the reward feedback information corresponding to the purification column system to be controlled at the current moment, and the historical reward feedback information sequence is updated based on the reward feedback information. Then, based on the reward feedback information and the time-series difference algorithm, the parameters of the state-action value model are updated.
[0071] Finally, the training ends when the preset training target corresponding to the reboiler opening prediction model is reached, and the reboiler opening prediction model is obtained.
[0072] S230. During the operation of the purification tower system to be controlled, acquire the operation process data corresponding to the purification tower system to be controlled.
[0073] S240. Based on the pre-trained reboiler opening prediction model, process the operation process data to obtain the first opening change of the hot water reboiler at the current moment and the second opening change of the steam reboiler at the current moment.
[0074] S250. Process the first opening change and the second opening change based on the preset opening value processing method.
[0075] The technical solution of this invention acquires the operation process data corresponding to the purification tower system during its operation. Further, based on a pre-trained reboiler opening prediction model, the operation process data is processed to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment. Finally, the first and second opening changes are processed based on a preset opening value processing method. This solves the problems of inaccurate prediction of the reboiler opening and complex control processes in related technologies. It achieves accurate control of the opening changes of the hot water and steam reboilers during the operation of the purification tower system, thereby reducing labor costs, ensuring the stability of the purification tower system, and improving the switching control efficiency between the hot water and steam reboilers.
[0076] Example 3
[0077] Figure 3 This is a schematic diagram of a reboiler switching control device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a process data acquisition module 310, an opening change determination module 320, and an opening change processing module 330.
[0078] The process data acquisition module 310 is used to acquire the operation process data corresponding to the purification tower system under control during the operation of the purification tower system under control. The operation process data includes the tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the purity value sequence of each separated object, the first opening value sequence of the hot water reboiler, the second opening value sequence of the steam reboiler, and the liquid flow rate value sequence of each target pipeline. The opening change determination module 320 is used to process the operation process data based on a pre-trained reboiler opening prediction model to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current time. The reboiler opening prediction model is trained based on a reinforcement learning algorithm. The opening change processing module 330 is used to process the first opening change and the second opening change based on a preset opening value processing method.
[0079] The technical solution of this invention acquires the operation process data corresponding to the purification tower system during its operation. Further, based on a pre-trained reboiler opening prediction model, the operation process data is processed to obtain the first opening change of the hot water reboiler and the second opening change of the steam reboiler at the current moment. Finally, the first and second opening changes are processed based on a preset opening value processing method. This solves the problems of inaccurate prediction of the reboiler opening and complex control processes in related technologies. It achieves accurate control of the opening changes of the hot water and steam reboilers during the operation of the purification tower system, thereby reducing labor costs, ensuring the stability of the purification tower system, and improving the switching control efficiency between the hot water and steam reboilers.
[0080] Optionally, the preset opening value processing method includes indirect control and visual display, and the opening change processing module 330 includes: a visual display unit.
[0081] A visualization display unit is used to visualize the first opening change and the second opening change based on the display interface of the target terminal.
[0082] Optionally, the preset opening value processing method includes direct control, and the opening change processing module 330 includes: a hot water reboiler opening adjustment unit and a steam reboiler opening adjustment unit.
[0083] A hot water reboiler opening adjustment unit is used to control the hot water reboiler to adjust its opening based on the first opening change, and to obtain the first opening value of the hot water reboiler at the next moment; and,
[0084] A steam reboiler opening adjustment unit is used to control the steam reboiler to adjust its opening based on the second opening change, and to obtain the second opening value of the steam reboiler at the next moment.
[0085] Optionally, the device further includes: a first opening value determination module, a second opening value determination module, and a processing mode switching module.
[0086] The first opening value determination module is used to, when the preset opening value processing method is direct control, determine the first opening value of the hot water reboiler at the current time based on the first opening value sequence, and determine the first opening value of the hot water reboiler at the next time based on the first opening change and the first opening value; and,
[0087] The second opening value determination module is used to determine the second opening value of the steam reboiler at the current time based on the second opening value sequence, and to determine the second opening value of the steam reboiler at the next time based on the second opening change and the second opening value.
[0088] The processing mode switching module is used to switch the processing mode of the preset opening value to indirect control and display it visually when the first opening value corresponding to the next moment reaches the first preset control threshold and / or the second opening value corresponding to the next moment reaches the second preset control threshold.
[0089] Optionally, the device further includes a training sample acquisition module and a model training module.
[0090] The training sample acquisition module is used to acquire the online simulation training samples of the purification tower system to be controlled in the current training round for each training round. The online simulation training samples are determined based on the simulation environment model corresponding to the purification tower system to be controlled. The online simulation training samples include the sample tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the sample purity value sequence of each separation object, the first sample opening value sequence of the hot water reboiler, the second sample opening value sequence of the steam reboiler, the sample liquid flow rate value sequence of each target pipeline, and the historical reward feedback information sequence corresponding to the at least one historical time.
[0091] The model training module is used to train the reboiler opening prediction model based on the online training samples of each simulation and the reinforcement learning algorithm.
[0092] Optionally, the model training module includes: a probability distribution determination unit, a first expected reward determination unit, a second expected reward determination unit, a reboiler opening prediction model update unit, an actual opening change determination unit, a reward feedback information determination unit, a state-action value model update unit, and a reboiler opening prediction model determination unit.
[0093] The probability distribution determination unit is used to input the sample tower temperature value sequence, the purity value sequence, the first sample opening value sequence, the second sample opening value sequence, and the liquid flow rate value sequence of the online training samples corresponding to the current training round into the reboiler opening prediction model to obtain the probability distribution of the actual changes of the hot water reboiler and the steam reboiler at the current time. The probability distribution of the actual changes includes the probability information of the first opening change of the hot water reboiler and the probability information of the second opening change of the steam reboiler.
[0094] The expected reward determination unit is used to input the current first opening change, the current second opening change, the sample tower temperature value sequence, the purity value sequence of each sample, the first sample opening value sequence, the second sample opening value sequence, and the liquid flow rate value sequence of each sample into the state action value model for each first opening change and each second opening change in the probability distribution of the actual change, so as to obtain the expected reward corresponding to the current moment.
[0095] The reboiler opening prediction model update unit is used to update the parameters of the reboiler opening prediction model based on the probability distribution of the actual change, the expected rewards, and the strategy gradient algorithm; and...
[0096] The actual opening change determination unit is used to determine the first actual opening change of the hot water reboiler and the second actual opening change of the steam reboiler at the current time based on the probability distribution of the actual change.
[0097] The reward feedback information determination unit is used to process the sample tower bottom temperature value sequence, the purity value sequence of each sample, the first sample opening value sequence, the second sample opening value sequence, the liquid flow rate value sequence of each sample, the first actual opening change, and the second actual opening change according to a pre-set multi-index reward function, to determine the reward feedback information corresponding to the purification tower system to be controlled at the current moment, and to update the historical reward feedback information sequence based on the reward feedback information;
[0098] The state-action value model update unit is used to update the parameters of the state-action value model based on the reward feedback information and the temporal difference algorithm.
[0099] The reboiler opening prediction model determination unit is used to end the training when the preset training target corresponding to the reboiler opening prediction model is reached, and obtain the reboiler opening prediction model.
[0100] Optionally, the multi-index reward function includes an observation index safety reward function, an observation index stability reward function, and a control index reward function. The observation index includes the sample column temperature value, sample purity value, first sample opening value, second sample opening value, and sample liquid flow rate value. The control index includes the first actual opening change and the second actual opening change.
[0101] The reboiler switching control device provided in the embodiments of the present invention can execute the reboiler switching control method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method.
[0102] Example 4
[0103] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0104] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0105] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0106] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the reboiler switching control method.
[0107] In some embodiments, the reboiler switching control method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the reboiler switching control method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the reboiler switching control method by any other suitable means (e.g., by means of firmware).
[0108] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0110] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0113] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0114] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0115] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A reboiler switching control method, characterized in that, include: During the operation of the purification tower system to be controlled, the operation process data corresponding to the purification tower system to be controlled is acquired. The operation process data includes the tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the purity value sequence of each separated object, the first opening value sequence of the hot water reboiler, the second opening value sequence of the steam reboiler, and the liquid flow rate value sequence of each target pipeline. The operating process data is processed based on the pre-trained reboiler opening prediction model to obtain the first opening change of the hot water reboiler at the current time and the second opening change of the steam reboiler at the current time; wherein, the reboiler opening prediction model is trained based on a reinforcement learning algorithm. The first opening change and the second opening change are processed based on a preset opening value processing method.
2. The method according to claim 1, characterized in that, The preset opening value processing method includes indirect control and visual display. The processing of the first opening change and the second opening change based on the preset opening value processing method includes: The first opening change and the second opening change are visualized on the target terminal's display interface.
3. The method according to claim 1, characterized in that, The preset opening value processing method includes direct control. The processing of the first opening change and the second opening change based on the preset opening value processing method includes: Based on the first change in opening, the opening of the hot water reboiler is adjusted, and the first opening value of the hot water reboiler at the next moment is obtained; and, The opening of the steam reboiler is adjusted based on the second opening change, and the second opening value of the steam reboiler at the next moment is obtained.
4. The method according to claim 3, characterized in that, Also includes: When the preset opening value processing method is direct control, the first opening value of the hot water reboiler at the current moment is determined based on the first opening value sequence, and the first opening value of the hot water reboiler at the next moment is determined based on the first opening change and the first opening value. as well as, The second opening value of the steam reboiler at the current moment is determined based on the second opening value sequence, and the second opening value of the steam reboiler at the next moment is determined based on the second opening change and the second opening value. If the first opening value corresponding to the next moment reaches the first preset control threshold, and / or the second opening value corresponding to the next moment reaches the second preset control threshold, the preset opening value processing method is switched to indirect control and displayed visually.
5. The method according to claim 1, characterized in that, Also includes: For each training round, obtain the online simulation training samples corresponding to the purification tower system to be controlled in the current training round. The online simulation training samples are determined based on the simulation environment model corresponding to the purification tower system to be controlled. The online simulation training samples include the sample tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the sample purity value sequence of each separation object, the first sample opening value sequence of the hot water reboiler, the second sample opening value sequence of the steam reboiler, the sample liquid flow rate value sequence of each target pipeline, and the historical reward feedback information sequence corresponding to the at least one historical time. The reboiler opening prediction model is trained based on the simulated online training samples and reinforcement learning algorithms.
6. The method according to claim 5, characterized in that, The training of the reboiler opening prediction model based on the aforementioned online training samples and reinforcement learning algorithms includes: For each training round, the sequence of sample tower temperature, the sequence of each sample purity, the sequence of the first sample opening, the sequence of the second sample opening, and the sequence of each sample liquid flow rate from the online training samples corresponding to the current training round are input into the reboiler opening prediction model to obtain the probability distribution of the actual changes of the hot water reboiler and the steam reboiler at the current time. The probability distribution of the actual changes includes the probability information of each first opening change of the hot water reboiler and the probability information of each second opening change of the steam reboiler. For each first opening change and each second opening change in the probability distribution of the actual change, the current first opening change, the current second opening change, the sample tower temperature value sequence, the purity value sequence of each sample, the first sample opening value sequence, the second sample opening value sequence, and the liquid flow rate value sequence of each sample are input into the state action value model to obtain the expected reward corresponding to the current moment. Based on the probability distribution of the actual change, the expected rewards, and the strategy gradient algorithm, the parameters of the reboiler opening prediction model are updated; and, Based on the probability distribution of the actual changes, determine the first actual opening change of the hot water reboiler at the current moment and the second actual opening change of the steam reboiler at the current moment; The sample tower bottom temperature value sequence, the purity value sequence of each sample, the first sample opening value sequence, the second sample opening value sequence, the liquid flow rate value sequence of each sample, the first actual opening change, and the second actual opening change are processed according to the pre-set multi-index reward function to determine the reward feedback information corresponding to the purification tower system to be controlled at the current moment, and the historical reward feedback information sequence is updated based on the reward feedback information. Based on the reward feedback information and the temporal difference algorithm, the parameters of the state-action value model are updated; Training ends when the preset training target corresponding to the reboiler opening prediction model is reached, thus obtaining the reboiler opening prediction model.
7. The method according to claim 6, characterized in that, The multi-index reward function includes a safety reward function for observed indices, a stability reward function for observed indices, and a reward function for control indices. The observed indices include the sample column temperature, sample purity, first sample opening value, second sample opening value, and sample liquid flow rate. The control indices include the first actual opening change and the second actual opening change.
8. A reboiler switching control device, characterized in that, The reboiler switching control device is used to execute the reboiler switching control method according to any one of claims 1-7, including: The process data acquisition module is used to acquire the operation process data corresponding to the purification tower system under control during the operation of the purification tower system under control. The operation process data includes the tower bottom temperature value sequence corresponding to the current time and at least one historical time before the current time, the purity value sequence of each separated object, the first opening value sequence of the hot water reboiler, the second opening value sequence of the steam reboiler, and the liquid flow rate value sequence of each target pipeline. The reboiler opening change determination module is used to process the operation process data based on the pre-trained reboiler opening prediction model to obtain the first opening change of the hot water reboiler at the current time and the second opening change of the steam reboiler at the current time; wherein, the reboiler opening prediction model is trained based on a reinforcement learning algorithm. The opening change processing module is used to process the first opening change and the second opening change based on a preset opening value processing method.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the reboiler switching control method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the reboiler switching control method according to any one of claims 1-7.
Citation Information
Patent Citations
Reboiler temperature control method for rectification tower
JP1993189062A