An event-triggered energy system optimization method based on a time-event dual-scale framework

By combining a time-event dual-scale framework with neural Kalman blotting and reinforcement learning, the problem of unstable state representation in industrial energy systems is solved, achieving efficient event-driven scheduling and improving the stability of learning and the economy of scheduling.

CN120409843BActive Publication Date: 2025-10-28DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510904954.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-28
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing industrial energy system optimization methods suffer from a lack of interpretability in state representation and event-driven scheduling, unstable learning performance, and failure to effectively extract data distribution in the time dimension, resulting in insufficient stability and reliability of reinforcement learning optimization decisions.

Method used

We adopt a time-event dual-scale framework approach, combining a neural Kalman model and reinforcement learning. By constructing hidden state representations and linking time-varying state matrices, we establish dynamic connections between states and achieve time-driven and event-driven optimized scheduling.

Benefits of technology

It improves the stability and reliability of reinforcement learning, realizes the learning mechanism of state estimation-decision-parameter update, enhances the convergence and long-term reward perception of RL policies, and meets the scheduling needs of industrial energy systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409843B_ABST
    Figure CN120409843B_ABST
Patent Text Reader

Abstract

This invention discloses an event-triggered energy system optimization method based on a time-event dual-scale framework, belonging to the field of information technology. First, it constructs an optimization problem for the operation of an industrial energy system. Second, it constructs reinforcement learning elements to build a simulation environment. Third, based on a neural Kalman model, it obtains state representations with explicit temporal distribution characteristics. Then, it employs a proximal policy optimization method to construct a reinforcement learning framework, which, together with the simulation environment and the neural Kalman model, constructs and trains a time-driven optimization model based on hidden state representations. Finally, it constructs an event-driven optimization model and combines it with the neural Kalman model through time-varying state matrix linking to construct and train an event-driven optimization model based on state transition matrix linking. Ultimately, it achieves event-triggered optimal scheduling through time-event dual-scale collaborative computation. This invention can obtain more economical energy scheduling strategies, meeting the actual scheduling needs of industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information technology and relates to an event-triggered energy system optimization method based on a time-event dual-scale framework, specifically a time-event dual-scale energy optimization method that combines neural Kalman state representation and reinforcement learning. Background Technology

[0002] In industrial practice, energy production and consumption are closely related to the production process, resulting in energy data lacking explicit patterns or distributional characteristics over time. State representation of this type of data is beneficial for the optimization effect of learning-based methods, improving the efficiency and effectiveness of the learning process. On the other hand, considering that the formulation of scheduling strategies for steel energy systems is usually event-driven (e.g., by-product gas storage is about to exceed safety limits), event-driven optimization is used to make decisions, and scheduling operations are executed when the event occurs. Therefore, to ensure the long-term safe and economical operation of the system and meet the actual scheduling needs of industry, researching event-driven models has significant practical implications.

[0003] To uncover the features and latent states of industrial data, existing research mainly uses end-to-end deep neural networks to achieve hierarchical feature representation of complex data, including: using Temporal Convolutional Networks (TCNs) to extract hidden information and long-term temporal relationships (Wang Y., Chen J., Chen X. (2020). Short-term load forecasting for industrial customers based on TCN-Light GBM[J]. IEEE Transactions on PowerSystems, 36(3):1984-1997), and using deep contrastive learning methods to characterize uncertain states in energy optimization environments (Wang T., Zhao J, Wang W. (2023). A condition knowledge representation and feedback learning framework for dynamic optimization of integrated energy systems[J]. IEEE Transactions on Cybernetics, DOI: 10.1109 / TCYB.2023.3234077.). However, the aforementioned method of using neural network layers with "black box" characteristics for state representation has drawbacks such as lack of interpretability and poor learning stability when facing industrial scenarios. Furthermore, it fails to extract the data distribution of feature states in the time dimension, which reduces the stability and reliability of reinforcement learning (RL) when making optimization decisions.To address the time dependence between adjacent potential states, several scholars have conducted research on this issue in recent years, including: proposing weighted linear dynamic systems for feature representation and soft sensor application in nonlinear dynamic industrial processes (Yuan X., Wang Y., Yang C. (2017). Weighted linear dynamic system for feature representation and soft sensor application in nonlinear dynamic industrial processes[J]. IEEE Transactions on Industrial Electronics, 65(2): 1508-1517); utilizing neural networks to assist Kalman filtering for state estimation of dynamic systems, especially for partially known dynamics (Revach G., Shlezinger N., Ni X. (2022). KalmanNet: Neural network aided Kalman filtering for partially known dynamics[J]. IEEE Transactions on Signal Processing, 70:1532-1547) or even scenarios with invisible conditions (Chen C., Lu C., Wang B. (2021). DynaNet: Neural Kalman dynamical model for motion estimation and prediction[J]. IEEE Transactions on Neural Networks, 65(2): 1508-1517); and using neural networks to assist Kalman filtering for state estimation of dynamic systems, especially for partially known dynamics (Revach G., Shlezinger N., Ni X. (2022). KalmanNet: Neural network aided Kalman filtering for partially known dynamics[J]. IEEE Transactions on Signal Processing, 70:1532-1547). Networks and Learning Systems, 32(12): 5479-5491.). However, the above methods mainly focus on state estimation and lack coordination with decision feedback under the RL framework, making it difficult to realize the learning mechanism of "state estimation - decision - parameter update".Furthermore, existing event-driven models (such as (Li Y., Zhang H., Liang X.(2018). Event-triggered-based distributed cooperative energy management for multi-energy systems[J]. IEEE Transactions on Industrial Informatics, 15(4):2008-2022.) and (Zhang N., Sun Q., Yang L. (2021). Event-triggered distributed hybrid control scheme for the integrated energy system[J]. IEEE Transactions on Industrial Informatics, 18(2): 835-846.)) only focus on communication efficiency and ignore the potential dynamic relationship between states before and after an event occurs. This results in incomplete modeling of the dynamic relationship between state-action-reward sequences during RL learning, which is detrimental to the convergence of RL policies and long-term reward perception. Summary of the Invention

[0004] The purpose of this invention is to provide an event-triggered energy system optimization method based on a time-event dual-scale framework. This method employs Neural Kalman blotting to extract temporal correlations, improving the interpretability of state representations and thus enhancing the stability and reliability of RL learning. Furthermore, the combination of time-varying state matrix linking and Neural Kalman blotting in this invention establishes a potential dynamic relationship between states before and after an event, enabling event-driven RL to perceive the long-term rewards of optimized scheduling, thereby obtaining a more economical energy scheduling strategy. This invention meets the actual scheduling needs of industry and can provide decision-making guidance for industrial energy optimization scheduling.

[0005] The technical solution of this invention is as follows:

[0006] An event-triggered energy system optimization method based on a time-event dual-scale framework includes: first, constructing an optimization problem for the operation of an industrial energy system; second, constructing reinforcement learning elements including states, actions, and rewards to build a simulation environment; third, obtaining state representations with explicit time distribution characteristics based on a neural Kalman model; then, constructing a reinforcement learning framework using a proximal policy optimization method, and jointly constructing and training a time-driven optimization model based on hidden state representations with the simulation environment and the neural Kalman model; finally, constructing an event-driven optimization model using another reinforcement learning framework and the simulation environment, and combining it with the neural Kalman model through time-varying state matrix linking to construct and train an event-driven optimization model based on state transition matrix linking, ultimately achieving event-triggered optimal scheduling through time-event dual-scale collaborative computation. Specifically, the method includes the following steps:

[0007] Step 1: Construct an optimization problem for the operation of an industrial energy system:

[0008] Based on the optimization objectives of industrial energy systems and the constraints to be considered in actual operation, an optimization problem for the operation of industrial energy systems is constructed. The optimization objective is to minimize the economic operating cost within a specific timeframe. Including energy dissipation costs Energy purchase costs Carbon emission costs and equipment start-up and shutdown consumption The constraints include energy balance constraints, mass balance constraints, equipment safety operation constraints, and equipment output limits; the optimization variable is the flow rate of the energy medium input to each adjustable device in the industrial energy system.

[0009] Step 2, Construct reinforcement learning elements:

[0010] Reinforcement learning elements are constructed based on the industrial energy system optimization problem to build a simulation environment for interaction with the reinforcement learning framework. The reinforcement learning elements include: states... ,action and instant rewards Among them, state Includes energy data from industrial energy systems, using historical energy data as the training set; actions Input the adjustment amount of the energy medium flow rate for each adjustable device in the industrial energy system; instant reward. It consists of the objective and constraints in the optimization problem in step 1.

[0011] Step 3, construct the hidden state representation based on the neural Kalman model:

[0012] Due to the use of build status Energy data does not exhibit explicit patterns or distributional characteristics over time, therefore its state... The lack of Markov property affects the optimization effect. A hidden state representation based on a neural Kalman model is constructed, starting from the state... Starting with the formed time series, the hidden states with explicit time distribution characteristics, i.e., Markov state transition properties, are calculated using a neural Kalman filter. Forming a hidden state sequence ,in, This indicates the interaction cycle of reinforcement learning. Indicates the moment of interaction.

[0013] The neural Kalman model is In the formula, The iterative process of the Kalman filter is described as shown in equation (1). Let be the hidden state error covariance matrix. The hidden state transition matrix is... The state observation matrix, The covariance matrix of the observation noise, Let be the hidden state noise covariance matrix.

[0014] (1)

[0015] In the formula, Indicates Kalman gain, Represents the identity matrix. This represents the transpose of a matrix.

[0016] Unlike conventional Kalman models, the parameters in neural Kalman models... and It is determined through neural network learning. This is because each state in the hidden state sequence... Following the quasi-linear relationship in equation (3), therefore, the transition matrix It contains historical hidden state information. Long Short-Term Memory (LSTM) networks in recurrent neural networks can capture the temporal dependencies in time series data and learn complex mappings from historical states to the current state. Therefore, LSTM is used for computation... and As shown in equation (2), its input is and , For LSTM in The internal hidden variables at time step. and the neural network weights in LSTM These parameters will be used as model parameters for the neural Kalman model and updated during subsequent learning.

[0017] (2)

[0018] (3)

[0019] Step 4: Construct a time-driven optimization model based on hidden state representation:

[0020] A reinforcement learning framework is constructed using the Proximal Policy Optimization (PPO) method. This framework includes a policy network and an evaluation network. The policy network comprises two identical networks: a new policy network and an old policy network. These two networks have the same structure and parameter settings. The reinforcement learning framework, along with the simulation environment built in step 2 and the neural Kalman model constructed in step 3, together form a time-driven optimization model based on hidden state representation, which is then trained. Specifically:

[0021] First, from the state Initially, the hidden states with Markov state transition properties were calculated using a neural Kalman model. Hidden state Generate actions using the old policy network .

[0022] Actions generated by the old policy network For multidimensional actions that satisfy a Beta distribution, i.e. ,in, This indicates that the input is Next, the old strategy network selects actions. The probability distribution, distribution parameters and Through the neural layers that constitute the policy network and approximate:

[0023] (4)

[0024] (5)

[0025] in, Activation function , Activation function ; and These are the parameter variables of the old policy network. Indicates action Dimensions , This represents the number of probability distributions output by the policy network. This represents the maximum range of motion. Each variable in The sampling is obtained by sampling based on the probability distribution generated by the old policy network. The specific sampling formula is as follows:

[0026] (6)

[0027] in, This is a Beta distribution disturbance used to control the policy exploration capability.

[0028] Subsequently, the simulation environment constructed from historical energy data in step 2 is based on the actions... , the current state Transform into the state of the next moment And receive instant rewards Complete an interaction, and Feedback is fed into the neural Kalman model to generate latent variables for the next time step. .

[0029] conduct After each interaction completes one interaction cycle, the result is obtained. Using a set of interactive data, a new policy network and an evaluation network are trained based on the interactive data, and the policy parameters are updated using gradient descent. With evaluation network parameters Among them, strategy parameters Including parameters of the new policy network and the parameters of the neural Kalman model in step 3. and ,Right now The parameter update process is as follows:

[0030] The reward function of the new policy network is defined as follows:

[0031] (7)

[0032] (8)

[0033] in, This represents the probability distribution generated by the network that satisfies the old policy. of The expected value of the network reward under the new strategy is calculated. for The importance ratio of the new and old strategy networks at any given moment. The shearing mechanism in PPO represents the guarantee. exist Within the range, Hyperparameters for controlling the degree of shearing This indicates the distribution generated by the new policy network. Select Action The probability, This indicates the distribution generated by the old policy network. Select Action The probability, The dominant function is approximated by the time difference method, i.e.:

[0034] (9)

[0035] in, Representing state Following the distribution Take action below The expected cumulative reward that can be obtained can be approximated by the time difference method as follows: , This indicates that the evaluation network is based on the input. Output at time Indicates time difference, Monte Carlo estimates, This represents the discount factor.

[0036] The evaluation network uses a multilayer perceptron (MLP) network. Approximation, i.e. The output of the evaluation network is used to calculate the advantage function using equation (9). When training the evaluation network, the reward function of the evaluation network is defined as:

[0037] (10)

[0038] (11)

[0039] (12)

[0040] in, It is the target value. Indicates the time step from the current time. Start to interaction cycle The Each time step In time step TD error.

[0041] After a parameter update based on the reward functions shown in equations (7) and (10), the updated new policy network parameters are passed one-to-one to the old policy network. After the old policy network is updated, it is used for data collection in the next interaction cycle, generating new interaction data and performing the next parameter update. Real-time rewards are continuously accumulated during the interaction process. When the accumulated instant reward value converges, the training is completed and the optimal parameters are determined. Based on the trained time-driven model, the optimal action can be automatically selected according to the current state collected in real time by the industrial energy system as a time-driven scheduling strategy.

[0042] Step 5: Construct an event-driven optimization model based on the linking of state transition matrices:

[0043] Another reinforcement learning framework and simulation environment are used to construct an event-driven optimization model. The simulation environment is constructed according to step 2, and the reinforcement learning framework still adopts the proximal policy optimization method in reinforcement learning. That is, the output settings of its policy network and evaluation network have the same structure as the reinforcement learning framework based on the time-driven optimization model based on hidden state representation in step 4. Through time-varying state transition matrix linking, the neural Kalman state representation is combined with the event-driven model to construct an event-driven optimization model based on state transition matrix linking. This model is used to obtain the event-driven state, and the optimized event-driven policy is obtained through time-event dual-scale policy co-computation. Specifically:

[0044] First, the time-driven optimization model based on hidden state representation trained in step 4 is trained at a fixed time period. The rolling computation time-driven scheduling strategy is not executed.

[0045] When a scheduled event is triggered, the event-driven state at the time the previous scheduled event occurred is used as a reference. By combining the time-driven states between the previous scheduling event and the current scheduling event, the event-driven state after the occurrence of the current scheduling event is calculated. The calculation process is as follows:

[0046] Assume the previous scheduling event is separated from the current scheduling event by a certain interval. Time period Then, according to equation (3), we can obtain:

[0047] (13)

[0048] Without disrupting the time-varying relationship of the state variables, when the event is triggered, the equations in equation (13) are multiplied together to obtain... In order to integrate State information for adjacent time periods. ,make This yields the event-driven state transition relationships, which are the state transition matrix links:

[0049] (14)

[0050] As can be seen from the above formula, the event-driven state still retains the linear time transition characteristics of the time-driven part, that is, the time distribution characteristics are carried over to the event-driven model.

[0051] Will Actions generated by the old policy network in the input event-driven optimization model , will obtain Given the event-driven optimization model in a simulation environment, the state is obtained. With instant rewards ;Will Feedback is fed into the time-driven optimization model based on hidden state representation to update the corresponding state. The time-driven optimization model based on hidden state representation continues with a fixed time period. Rolling computation continues until the next scheduling event occurs, enabling collaborative computation of a time- and event-based dual-scale strategy. For events that occur... The secondary scheduling event, based on the collaborative computation of a time-event dual-scale strategy, yields... Group interactive data, using this Train the event-driven optimization model using group interaction data and update its new policy network parameters. ), old strategy network parameters ( and evaluation network parameters The parameter update process is the same as in step 4.

[0052] To reduce scheduling frequency while maximizing learning effectiveness, the reward of the event-driven optimization model is defined as... Among them, time-driven reward value , for adjacent events arrive Within range The sum of event-driven reward values Adjacent events arrive Instant rewards for event-driven optimization models within a given scope The sum of, and also: when the event has not occurred, For the simulation environment in the event-driven model, based on Calculations show that when the event occurs, In an event-driven model, the simulation environment is determined by the actions taken. Calculated. Accumulated continuously during the interaction process. When the cumulative Training is complete upon convergence, and the optimal parameters for the event-driven model are determined.

[0053] The trained event-driven model based on state transition matrix linking, together with the trained time-driven optimization model based on hidden state representation, achieves time-event dual-scale collaborative computation. This allows the optimal action to be calculated based on the current state collected in real time by the industrial energy system when the corresponding event occurs, serving as an event-driven scheduling strategy.

[0054] The beneficial effects of this invention are as follows: This invention employs a neural Kalman grammar to establish explicit state transition relationships, making the state representation process interpretable and improving the stability and reliability of reinforcement learning (RL) in optimization decision-making. Simultaneously, the synergy between state representation and the decision feedback of the RL framework realizes a learning mechanism of "state estimation—decision—parameter update." Furthermore, the time-event dual-scale collaborative computation method establishes the potential dynamic relationship between states before and after an event occurs, improving the convergence of the RL policy and its long-term reward perception capability. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the application of the present invention.

[0056] Figure 2 This is a schematic diagram illustrating the principle of the present invention.

[0057] Figure 3 This is a schematic diagram of the optimization of the time-driven optimization model based on hidden state representation.

[0058] Figure 4 This is a schematic diagram showing the convergence of reward values ​​and the deviation of reward values ​​between the example and PPO-clip, PPO-TCN, A3C, and TRPO.

[0059] Figure 5 The diagrams show a comparison between time-driven and event-driven strategies during the training process of the time-event dual-scale strategy. (a) is a comparison diagram of reward value convergence; (b) is a magnified comparison diagram of the two convergence curves in the later stage of training in (a); and (c) is a convergence curve of the difference in reward value.

[0060] Figure 6 The diagrams show the optimization effects of capacity scheduling or operation of coke oven gas holders and blast furnace gas holders. (a) is the time-driven scheduling effect diagram of coke oven gas holders; (b) is the event-driven scheduling effect diagram of coke oven gas holders; (c) is the operation optimization effect diagram of blast furnace gas holders during the transition period from peak to valley electricity prices; and (d) is the operation optimization effect diagram of blast furnace gas holders during the transition period from valley to peak electricity prices.

[0061] Figure 7 The optimized self-generated electricity generation varies with the purchase and sale prices, where (a) represents the change in self-generated electricity generation during the transition period from peak to off-peak electricity prices; and (b) represents the change in self-generated electricity generation during the transition period from off-peak to peak electricity prices.

[0062] Figure 8 Box plots are shown for the operating cost values ​​under various scenarios; (a) is the cost box plot for electricity prices at peak times; (b) is the cost box plot for electricity prices at valley times; (c) is the cost box plot for electricity prices during transition periods; and (d) is the overall cost box plot. Detailed Implementation

[0063] The generation and consumption of industrial energy are influenced by factors such as production processes and manual operations, resulting in reinforcement learning states built based on energy data lacking obvious distribution characteristics. To meet the supply and demand balance of various energy media and ensure the stable operation of the entire production process, on-site operators need to understand the changes in energy generation and consumption and their storage status in real time, and make scheduling actions based on events. If a latent state representation with temporal distribution characteristics and interpretability can be constructed, scheduling schemes can be well formulated to achieve energy conservation and emission reduction. Therefore, obtaining a state representation with explicit temporal distribution characteristics and implementing event-driven dynamic scheduling can improve the learning efficiency and model performance of RL, and meet the actual scheduling needs on-site. To better understand the technical route and implementation scheme of this invention, this embodiment uses a steel energy system as a background and actual industrial data as a basis to illustrate the specific implementation steps of this invention. The flowchart of this invention is as follows: Figure 1 The schematic diagram is as follows Figure 2 It includes the following steps:

[0064] Step 1: Construct an optimization problem for the operation of an industrial energy system:

[0065] Taking the steel energy system as an example, an optimization problem for the operation of the steel energy system is constructed based on the optimization objectives and the constraints that need to be considered in actual operation. The optimization variable is the flow rate of the energy medium input to each adjustable device in the steel energy system.

[0066] Step 1.1: The objective of the steel energy system optimization problem is to minimize the economic operating cost within a specific time frame. Including energy dissipation costs Energy purchase costs Carbon emission costs and equipment start-up and shutdown consumption The objective function is expressed as follows:

[0067] (15)

[0068] (16)

[0069] (17)

[0070] (18)

[0071] (19)

[0072] in, Let be the objective function. To optimize the time step of the problem, These are the unit venting costs for coal gas and steam, respectively. They are respectively The amount of gas and steam released at any given time; The cost of purchasing thermal coal, , These are the costs of purchasing and selling electricity, respectively. This refers to the consumption of purchased coal. These represent the electricity load demand and the power generation from the self-owned power plant, respectively. For adjustable equipment The start-stop consumption cost, To allow for adjustment of the total number of devices, Indicates adjustable equipment The running status, To close, To enable; The average unit price of carbon emissions, This represents the carbon emission factor.

[0073] Step 1.2, Construct the constraints for the steel energy system optimization problem:

[0074] The constraints that need to be considered in the optimization of the steel energy system operation include energy balance constraints, quality balance constraints, equipment safety operation constraints, and equipment output limits.

[0075] 1) Energy Balance Constraints. Various energy media in the steel production process can be converted using energy conversion equipment to meet the required energy balance. The energy balance constraints of energy conversion equipment mainly include the balance of the inflow and outflow energy media:

[0076] (20)

[0077] Where, constant and It is an efficiency parameter, obtained through a data-driven regression method; and These represent energy conversion equipment. At the point of time The flow rate of the input and output energy medium.

[0078] 2) Mass balance constraint:

[0079] (twenty one)

[0080] in, They represent The amount of gas produced and consumed at any given time. , To meet the gas load required for production, The first The power plant, the first Gas consumption of each cogeneration unit This indicates the total number of power plants. This indicates the total number of combined heat and power (CHP) units; Indicates gas holder The cabinet capacity, This indicates the total number of gas holders.

[0081] Similarly, the steam mass balance constraint is:

[0082] (twenty two)

[0083] in, This indicates that the steam pressure level is The A power plant or combined heat and power unit in The amount of steam generated at any given time The steam generation pressure level is The total number of power plants and combined heat and power units, This refers to the amount of steam emitted. For pressure level Steam load.

[0084] The power balance constraints are:

[0085] (twenty three)

[0086] In the formula, Indicates the amount of electricity purchased externally. This indicates the amount of electricity sold externally.

[0087] 3) Equipment safety operation constraints and output limits. The steel energy system includes various energy conversion and storage devices. The safety operation constraints and output limits for these devices are as follows:

[0088] (twenty four)

[0089] (25)

[0090] (26)

[0091] in, Indicates energy conversion equipment exist Constant effort The first The lower and upper limits of the output of an energy conversion or energy storage device. The first The lower and upper limits of the input to an energy conversion or energy storage device. Energy conversion equipment The slope of the increase and decrease in output. For energy storage devices exist The operational level at any given moment, Energy storage devices The lower and upper limits of the operating level.

[0092] Step 2, Construct reinforcement learning elements:

[0093] Based on the optimization problem of a steel energy system, reinforcement learning elements are constructed to build a simulation environment for interaction with the reinforcement learning framework. The reinforcement learning elements include: states... ,action and instant rewards .

[0094] 1) Status Including the production process and Because electricity prices vary at different times, the current time Time period It is also considered a key state factor:

[0095] (27)

[0096] Historical energy data collected at the minute level is used as the training set for model training in subsequent steps.

[0097] 2) Actions :action These are the decision variables in the optimization problem of step 1. The dynamic scheduling of the steel energy system controls the time periods... and This ensures that the electricity and steam generated by power plants and combined heat and power (CHP) units meet load demands and input / output constraints. As a supplementary input to power plants and combined heat and power units, it can and Once determined, the action is based on the difference between the flow rate of the energy medium output by the energy conversion equipment and its load demand. Defined as:

[0098] (28)

[0099] 3) Instant rewards Including system operating costs and penalties for exceeding limits :

[0100] (29)

[0101] in, and These represent the weight coefficients in the reward function, and Larger than ; The calculation formula is shown in equation (15). To optimize the constraint violation rate of the problem.

[0102] Step 3, construct the hidden state representation based on the neural Kalman model (the principle is as follows). Figure 3 (as shown)

[0103] Due to the use of build status Energy data does not exhibit explicit patterns or distributional characteristics over time, therefore its state... The lack of Markov property affects the optimization effect. A hidden state representation based on a neural Kalman model is constructed to represent the state... The hidden states with Markov state transition characteristics are calculated using a neural Kalman filter. This forms a hidden state sequence. Specifically:

[0104] The neural Kalman model is In the formula, The iterative process of the Kalman filter is described as shown in equation (1).

[0105] Unlike conventional Kalman models, the parameters in neural Kalman models... and It is determined through neural network learning. This is because each state in the hidden state sequence... Following the quasi-linear relationship in equation (3), therefore, the transition matrix It contains historical hidden state information. Long Short-Term Memory (LSTM) networks in recurrent neural networks can capture the temporal dependencies in time-series data and learn complex mappings from historical states to the current state. Therefore, LSTM is used to compute... and As shown in equation (2), its input is and For LSTM in The internal hidden variables at time step. and the neural network weights in LSTM These parameters will be used as model parameters for the neural Kalman model and will be continuously updated during the subsequent learning process.

[0106] Step 4: Construct a time-driven optimization model based on hidden state representation:

[0107] A reinforcement learning framework is constructed using the proximal policy optimization method from reinforcement learning. This framework includes a policy network and an evaluation network. The policy network comprises two networks: a new policy network and an old policy network. Both policy networks have identical structures and parameter settings. This reinforcement learning framework, along with the simulation environment built in step 2 and the neural Kalman model constructed in step 3, together form a time-driven optimization model based on hidden state representation, which is then trained. Specifically:

[0108] First, let's consider the state. Initially, the hidden states with Markov state transition properties were calculated using a neural Kalman model. Hidden state Generate actions using the old policy network .

[0109] Actions generated by the old policy network For multidimensional actions that satisfy a Beta distribution, i.e. Among them, the distribution parameters and Through the neural layers that constitute the policy network and Approximately, as shown in equations (4) and (5). Then... Each variable in The probability distribution generated by the old strategy network is sampled, as shown in equation (6).

[0110] Subsequently, the simulation environment, constructed from historical energy data, is based on the actions... , the current state Transform into the state of the next moment And receive instant rewards Complete an interaction, and Feedback is fed into the neural Kalman model to generate latent variables for the next time step. .

[0111] conduct Each interaction completes one interaction cycle, resulting in... Using a set of interactive data, a new policy network and an evaluation network are trained based on the interactive data, and the policy parameters are updated using gradient descent. With evaluation network parameters The parameter update process is as follows:

[0112] The reward function of the new policy network is defined as shown in equations (7) to (9).

[0113] The evaluation network uses a multilayer perceptron network. Approximation, i.e. The output of the evaluation network is used to calculate the advantage function using equation (9). When training the evaluation network, the reward function of the evaluation network is defined as shown in equations (10) to (12).

[0114] After a parameter update based on the reward functions shown in equations (7) and (10), the updated new policy network parameters are passed to the old policy network one-to-one. After the old policy network is updated, it is used for data collection in the next interaction cycle to generate new interaction data and perform the next parameter update. Real-time reward values ​​are continuously accumulated during the interaction process. When the accumulated instant reward value converges, the training is completed and the optimal parameters are determined. Based on the trained time-driven model, the optimal action can be automatically selected according to the current state collected in real time by the steel energy system as a time-driven scheduling strategy.

[0115] Step 5: Construct an event-driven optimization model based on the linking of state transition matrices:

[0116] Another reinforcement learning framework and simulation environment are used to construct an event-driven optimization model. The construction of this model is essentially the same as that of the time-driven optimization model based on hidden state representation, except that the event-driven optimization model does not include a neural Kalman model. That is, the reinforcement learning framework in the event-driven optimization model is also constructed using the proximal policy optimization method in reinforcement learning. Furthermore, the simulation environment is constructed according to step 2. By linking time-varying state transition matrices, the neural Kalman state representation is combined with the event-driven model to construct an event-driven optimization model based on state transition matrix linking. This model is used to obtain the event-driven state, and the optimized event-driven policy is obtained through time-event dual-scale policy co-computation. Specifically:

[0117] First, the time-driven optimization model based on hidden state representation trained in step 4 is trained at a fixed time period. The rolling computation time-driven scheduling strategy is not executed.

[0118] When a scheduled event is triggered, the event-driven state at the time the previous scheduled event occurred is used as a reference. Based on the time-driven state between the previous scheduling event and the current scheduling event, the event-driven state after the occurrence of the current scheduling event is calculated according to equations (13) and (14). .

[0119] Will Actions generated by the old policy network in the input event-driven optimization model , will obtain Given the event-driven optimization model in a simulation environment, the state is obtained. With instant rewards ;Will Feedback is fed into the time-driven optimization model based on hidden state representation to update the corresponding state. The time-driven optimization model based on hidden state representation continues with a fixed time period. Rolling computation continues until the next scheduling event occurs, enabling collaborative computation of a time- and event-based dual-scale strategy. For events that occur... The secondary scheduling event, based on the collaborative computation of a time-event dual-scale strategy, yields... Group interactive data, using this Train the event-driven optimization model using group interactive data, and update its new policy network parameters, old policy network parameters, and evaluation network parameters. The parameter update process is the same as in step 4.

[0120] To reduce scheduling frequency while maximizing learning effectiveness, the reward of the event-driven optimization model is defined as... ,in , for adjacent events arrive Within range The sum of Adjacent events arrive Instant rewards for event-driven optimization models within a given scope The sum of, and also: when the event has not occurred, For the simulation environment in the event-driven model, based on Calculations show that when the event occurs, In an event-driven model, the simulation environment is determined by the actions taken. Calculated. Accumulated continuously during the interaction process. When the cumulative Training is complete upon convergence, and the optimal parameters for the event-driven model are determined.

[0121] The trained event-driven model based on state transition matrix linking, together with the trained time-driven optimization model based on hidden state representation, achieves time-event dual-scale collaborative computation. This allows the optimal action to be calculated based on the current state collected in real time by the industrial energy system when the corresponding event occurs, serving as an event-driven scheduling strategy.

[0122] To verify the effectiveness of this embodiment, actual data from the energy system of a domestic steel plant from April to May 2022 (collected by the data acquisition and monitoring control system) was selected for the experiment, with a data sampling interval of 1 minute. The data from the first 24 days was used as the training set for model training, and the remaining data was used as the test set. Since a minute-level intraday scheduling strategy needs to be formulated during real-time scheduling, the interaction cycle length was set to 2 hours. Set to 5 minutes.

[0123] To assess the model training performance, using the same training set, we compared four RL methods with this embodiment: PPO-clip (based on PPO), Asynchronous Advantage Actor-Critic (A3C), RL based on Confidence Region Policy Optimization (TRPO), and PPO-TCN (based on PPO combined with a Temporal Convolutional Network). Multiple experiments were conducted for each method, and the resulting learning curves and standard deviations are shown below. Figure 4 As shown in the figure, the three PPO-based methods (PPO-clip, PPO-TCN, and this embodiment) exhibit better learning stability. Compared to the A3C and TRPO models, this embodiment and PPO-TCN achieve larger reward values ​​in the later stages of training, demonstrating the effectiveness of the hidden state representation model in the industrial data learning process.

[0124] Figure 5 Figure (a) shows the time-driven reward values ​​in step 5 of this embodiment. and event-driven reward value The convergence curves, due to the co-training process of time and events, exhibit similar fluctuation trends in both convergence curves. Furthermore, from... Figure 5 As can be seen from (b) and (c), the difference in rewards between the two convergence curves gradually decreases as the number of training iterations increases. At the end of the training process, the event-driven reward even exceeds the time-driven reward. The above results show that the event-driven scheduling method proposed in this embodiment can achieve the overall expected training effect.

[0125] Regarding the scheduling effects of different energy media, Figure 6 Figures (a) and (b) describe the operating results of coke oven gas (COG) holders under time-driven and event-driven strategies, respectively. Because the COG source fluctuates little, the daily operation of the holders is relatively stable. However, the time-driven strategy requires periodic scheduling actions, leading to frequent fluctuations in COG holder positions. Conversely, the event-driven strategy reduces the number of scheduling operations, and its optimized holder operation results better reflect the actual process requirements on-site. Figure 6 Figures (c) and (d) show the operational optimization effect of blast furnace gas (BFG) cabinet capacity under time-of-use pricing. Figure 6 As can be seen in (c), when the electricity price changes from off-peak to peak hours, the optimized BFG capacity shows a downward trend, in order to utilize stored gas to increase the company's self-generated electricity. Conversely, when the electricity price decreases from high to low (e.g., ...), the optimized BFG capacity decreases. Figure 6 As shown in (d), the cabinet capacity will be higher than the original value, achieving greater gas storage. The optimized self-generated power generation varies with the purchase and sale prices as follows: Figure 7As shown, the results indicate that during periods of low electricity prices, the steel energy system purchases electricity from the external grid to reduce costs; during periods of peak electricity prices, the steel energy system utilizes steam turbines and CHP to increase power generation and reduce the amount of electricity purchased during peak periods. Therefore, as... Figure 7 As shown in (a), self-generated electricity decreases when electricity prices transition from peak to off-peak hours; Figure 7 As shown in Figure (b), self-generated electricity increases as electricity prices transition from off-peak to peak hours. These results demonstrate that this embodiment optimizes energy utilization by taking into account changes in electricity prices, thereby helping to reduce electricity costs.

[0126] To verify the effectiveness of this embodiment in reducing overall operating costs, traditional manual experience scheduling and some commonly used optimization methods and state representation RL methods were selected as comparison methods to compare with the method in this embodiment. The comparison methods include: on-site manual experience scheduling (method a), mixed integer linear programming method (method b), model predictive control optimization method (method c), original PPO-based method (method d), PPO-based method combined with temporal convolutional network method (method e), time-driven optimization method based on neural Kalman state representation (method f), asynchronous advantage Actor-Critic method (method g), and confidence region-based policy optimization method (method h).

[0127] Based on the same test set data, this embodiment and methods a to h are used to calculate the optimized operating costs under three scenarios: peak electricity price, valley electricity price, and transition period, respectively, and the mean, minimum, maximum, and standard deviation of the optimized operating costs under each scenario are statistically analyzed. Figure 8 The peak electricity price (e.g.) is given Figure 8 As shown in (a), valley time (as shown in the middle) Figure 8 As shown in (b), transition period (as shown in the middle) Figure 8 As shown in (c), the overall operating cost value (as shown in the figure) is as follows: Figure 8 The box plot shown in (d) is presented in Table 1, which displays the corresponding numerical statistics. Figure 8Table 1 shows that this embodiment outperforms other methods in terms of average cost under different electricity price scenarios. Table 2 further summarizes the total daily operating cost of different methods under the test set. The results in Table 2 show that: method c is inferior to method d and this embodiment in terms of total cost; method b is prone to getting trapped in local optima and performs worse than this embodiment in optimizing energy system operating costs; compared with baseline RL methods like methods d, g, and h, this embodiment establishes the temporal correlation between states to fully leverage the optimization effect of RL, achieving lower operating costs; this embodiment outperforms method e in all three scenarios, indicating that using neural Kalman blotting for state representation is more suitable for industrial data modeling than traditional hidden state representation methods; although the cost improvement of this embodiment is limited compared to method f, it can significantly reduce the costs associated with frequent scheduling. Furthermore, the cost variance of this embodiment is small, indicating that it can adapt well to different energy generation, consumption, and conversion conditions. These results demonstrate the effectiveness of the time-event dual-scale collaborative computation method in this embodiment in establishing the potential dynamic relationship between states before and after an event, improving the convergence and long-term return perception of the RL strategy, thereby bringing higher economic benefits.

[0128] Table 1: Comparison of operating costs under different scenarios

[0129]

[0130] Table 2: Comparison Results of Daily Operating Costs

[0131]

Claims

1. An event-triggered energy system optimization method based on a time-event dual-scale framework, characterized in that, include: Step 1: Construct an optimization problem for the operation of an industrial energy system; Based on the optimization objectives of the industrial energy system and the constraints that need to be considered in actual operation, an optimization problem for the operation of the industrial energy system is constructed, with the optimization variable being the flow rate of the energy medium input to each adjustable device in the industrial energy system; Step 2: Construct reinforcement learning elements including state, action, and reward to build a simulation environment; Based on the industrial energy system optimization problem, reinforcement learning elements are constructed to build a simulation environment. These elements include: states... ,action and instant rewards Among them, state Includes energy data from industrial energy systems; actions Input the adjustment amount of the energy medium flow rate for each adjustable device in the industrial energy system; instant reward. The process consists of the objective and constraints in the optimization problem in step 1; step 3: based on the neural Kalman model, obtain the state representation with explicit time distribution characteristics; step 4: construct a reinforcement learning framework using the proximal policy optimization method, and jointly construct and train a time-driven optimization model based on the hidden state representation with the simulation environment and the neural Kalman model; step 5: construct an event-driven optimization model with another reinforcement learning framework and the simulation environment, and combine it with the neural Kalman model through time-varying state matrix linking to construct and train an event-driven optimization model based on state transition matrix linking, and finally achieve event-triggered optimization scheduling through time-event dual-scale collaborative computation.

2. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 1, characterized in that, Step 3 includes Constructing hidden state representations based on a neural Kalman model, starting from the state Starting with the formed time series, the hidden states with explicit time distribution characteristics, i.e., Markov state transition properties, are calculated using a neural Kalman filter. This forms a hidden state sequence; The neural Kalman model is In the formula, The iterative process of the Kalman filter is described as shown in equation (1). Let be the hidden state error covariance matrix. Hidden state transition matrix For the state observation matrix, The covariance matrix of the observation noise, The hidden state noise covariance matrix; (1) In the formula, Indicates Kalman gain, Represents the identity matrix. Represents the transpose of a matrix; Parameters in the neural Kalman model and Determined through neural network learning; due to the various states in the hidden state sequence. Following the quasi-linear relationship in equation (3), a long short-term memory network is used for computation. and ,and and the neural network weights in LSTM The model parameters, used as parameters in the neural Kalman model, are updated during subsequent learning. (3) Step 4 includes constructing a time-driven optimization model based on hidden state representation: A reinforcement learning framework is constructed using a proximal policy optimization method. This framework includes a policy network and an evaluation network; the policy network comprises two parts: a new policy network and an old policy network. This reinforcement learning framework, along with the simulation environment built in step 2 and the neural Kalman model constructed in step 3, together form a time-driven optimization model based on hidden state representation, which is then trained. Specifically: First, from the state Initially, the hidden states with Markov state transition properties were calculated using a neural Kalman model. Hidden state Generate actions using the old policy network ; Then, the simulation environment adjusts according to the actions. , the current state Transform into the state of the next moment And receive instant rewards To complete an interaction, Feedback is fed into the neural Kalman model to generate latent variables for the next time step. ; conduct After each interaction completes one interaction cycle, the result is obtained. Using interactive data, we define the reward functions for the new policy network and the evaluation network. We then train both networks using the interactive data and update the policy parameters using gradient descent. With evaluation network parameters Among them, strategy parameters Including parameters of the new policy network and the parameters of the neural Kalman model in step 3. and ; After a parameter update, the updated policy network parameters are passed to the old policy network. The updated old policy network then collects data for the next interaction cycle, generating new interaction data and performing the next parameter update. Real-time rewards are continuously accumulated during the interaction process. When the accumulated instant reward value converges, the training is completed and the optimal parameters are determined. Based on the trained time-driven model, the optimal action can be automatically selected according to the current state collected in real time by the industrial energy system as a time-driven scheduling strategy. Step 5 involves constructing an event-driven optimization model based on the linking of state transition matrices: Another reinforcement learning framework and simulation environment are used to construct an event-driven optimization model. The simulation environment is constructed according to step 2, and the reinforcement learning framework has the same structure as the time-driven optimization model based on hidden state representation in step 4. Through time-varying state transition matrix linking, the neural Kalman state representation is combined with the event-driven model to construct an event-driven optimization model based on state transition matrix linking. This model is used to obtain the event-driven state, and the optimized event-driven policy is obtained through time-event dual-scale policy collaborative computation. Specifically: First, the time-driven optimization model based on hidden state representation trained in step 4 is trained at a fixed time period. The rolling computation time-driven scheduling strategy is not executed; When a scheduled event is triggered, the event-driven state at the time the previous scheduled event occurred is used as a reference. By combining the time-driven states between the previous scheduling event and the current scheduling event, the event-driven state after the occurrence of the current scheduling event is calculated. The calculation process is as follows: Assume the previous scheduling event is separated from the current scheduling event by a certain interval. Time period Then, according to the formula, we can obtain: (13) Without disrupting the time-varying relationship of the state variables, when the event is triggered, the equations in equation (13) are multiplied together to obtain... In order to integrate State information of adjacent time periods; because ,make This yields the event-driven state transition relationships, which are the state transition matrix links: (14) Will Actions generated by the old policy network in the input event-driven optimization model , will obtain Given the event-driven optimization model in a simulation environment, the state is obtained. With instant rewards ;Will Feedback is fed into the time-driven optimization model based on hidden state representation to update the corresponding state. The time-driven optimization model based on hidden state representation continues with a fixed time period. Rolling computation continues until the next scheduling event occurs, achieving collaborative computation of a time- and event-based dual-scale strategy; for events that occur... This scheduling event yields... Group interactive data, using this Train the event-driven optimization model using group interactive data, and update its new policy network parameters, old policy network parameters, and evaluation network parameters. The parameter update process is the same as in step 4. The reward of the event-driven optimization model is defined as Among them, time-driven reward value , for adjacent events arrive Within the scope The sum of event-driven reward values Adjacent events arrive Instant rewards for event-driven optimization models within a given scope The sum; continuously accumulated during the interaction process. When the cumulative Training is complete upon convergence, and the optimal parameters of the event-driven model are determined. The trained event-driven model based on state transition matrix linking, together with the trained time-driven optimization model based on hidden state representation, achieves time-event dual-scale collaborative computation. This allows the optimal action to be calculated based on the current state collected in real time by the industrial energy system when the corresponding event occurs, serving as an event-driven scheduling strategy.

3. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 2, characterized in that, In step 1, the optimization objective is to minimize the economic operating cost within a specific time frame, and the constraints include energy balance constraints, quality balance constraints, equipment safe operation constraints, and equipment output limits.

4. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 2, characterized in that, In step 3, and The calculation formula is as shown in equation (2), and its input is... and , For Long Short-Term Memory Networks in Internal hidden variables at time; (2)。 5. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 2, characterized in that, In step 4, the old policy network generates actions. The process is as follows: Actions generated by the old policy network For multidimensional actions that satisfy a Beta distribution, i.e. ,in, This indicates that the input is Below, the old strategy network selects actions. probability distribution, distribution parameters and Through the neural layers that constitute the policy network and Approximate, then Each variable in The sampling is obtained by sampling based on the probability distribution generated by the old policy network. The specific sampling formula is as follows: (6) in, Indicates action Dimensions , This represents the number of probability distributions output by the policy network. This is interference from the Beta distribution.

6. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 5, characterized in that, Distribution parameters and The calculation formula is as follows: (4) (5) in, Activation function , Activation function ; and These are the parameter variables of the old policy network. This represents the maximum range of motion.

7. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 2, characterized in that, In step 4, the parameter update process is as follows: The reward function for the new policy network is defined as follows: (7) (8) in, This represents the probability distribution generated by the network that satisfies the old policy. of The expected value of the network reward under the new strategy is calculated. for The importance ratio of the new and old strategy networks at any given moment. For the pruning mechanism in the near-end policy optimization method, it means that the guarantee is... exist Within the range, Hyperparameters for controlling the degree of shearing This indicates the distribution of the new policy network in the generated data. Select Action The probability of This indicates the distribution generated by the old policy network. Select Action The probability of The dominant function is approximated using the time difference method; The evaluation network uses a multilayer perceptron network. Approximation, i.e. The reward function for evaluating the network is defined as follows: (10) (11) (12) in, It is the target value. Indicates the time step from the current time. Start, to the interaction cycle The Each time step In time step of error, This represents the discount factor.

8. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 7, characterized in that, Advantage function The calculation formula is as follows: (9) in, Indicates the state Following the distribution Take action below The expected cumulative reward that can be obtained can be approximated by the time difference method as follows: , Indicates time difference, This indicates the Monte Carlo estimate.

9. The event-triggered energy system optimization method based on a time-event dual-scale framework according to claim 2, characterized in that, In step 5, if the event does not occur, For the simulation environment in the event-driven model, based on Calculations show that when the event occurs, In an event-driven model, the simulation environment is determined by the actions taken. Calculated.

Citation Information

Patent Citations

  • Deep reinforcement learning method for microgrid energy scheduling under consideration of source load uncertainty

    CN116247648A

  • Comprehensive energy system economic dispatching model method based on deep reinforcement learning

    CN119273066A