Energy system low-carbon scheduling method and device based on deep reinforcement learning
By optimizing the decision network input of the energy system using Riemannian manifold features and swarm stochastic neural differential equation modeling techniques, the problem of slow training and decision-making speed of multi-agent deep reinforcement learning algorithms in energy systems is solved, achieving fast response and global optimization scheduling effects.
Patent Information
- Application Number
- CN202511555045.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing multi-agent deep reinforcement learning algorithms suffer from slow training and decision-making speeds in energy systems, making it difficult to achieve rapid response and global optimization.
By extracting Riemannian manifold features from local state information, the attention weights of local decision networks are evaluated, key decision networks are selected, and a population stochastic neural differential equation modeling technique is used to construct a population distribution of manifold features, optimize the input of decision networks, and combine deep reinforcement learning to train local decision networks, thereby reducing input dimensionality and decision complexity and enhancing decision stability.
It significantly improves the scheduling efficiency and decision-making stability of the energy system, enables rapid response and global optimization, and meets the decision-making needs in complex environments.
Smart Images

Figure CN121503983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power grid dispatching technology, and in particular to a low-carbon dispatching method and device for energy systems based on deep reinforcement learning. Background Technology
[0002] Energy systems have become an important means of improving energy efficiency and promoting the consumption of renewable energy. However, with the expansion of system scale and the large-scale integration of renewable energy, the random fluctuations between sources and loads have increased, making it difficult for traditional scheduling methods to achieve rapid response and global optimization. Current scheduling methods for this type of problem include mathematical programming algorithms, heuristic algorithms, and deep reinforcement learning algorithms. Mathematical programming methods, represented by mixed-integer linear programming, can guarantee the optimality of the solution, but they have weak adaptability to uncertainty and are inefficient when dealing with high-dimensional problems. Heuristic algorithms such as genetic algorithms and particle swarm optimization have global search capabilities, but their convergence speed is slow and the quality of the solutions is unstable. Deep reinforcement learning algorithms, represented by multi-agent deep reinforcement learning, learn scheduling strategies directly from the environment through data-driven approaches, and have strong adaptability to uncertainty, but they suffer from slow training and decision-making speeds when facing high-dimensional problems in energy systems.
[0003] The technical problem addressed in this application is the last scheduling method mentioned above, namely, how to improve the training and decision-making efficiency of energy low-carbon scheduling algorithms based on multi-agent deep reinforcement learning, so as to achieve rapid real-time scheduling of energy systems. Summary of the Invention
[0004] This application provides a low-carbon scheduling method and apparatus for energy systems based on deep reinforcement learning, which can reduce the input data dimensionality and computational burden of a single local decision network, improve solution efficiency, and realize rapid scheduling of energy systems.
[0005] In a first aspect, embodiments of this application provide a low-carbon scheduling method for energy systems based on deep reinforcement learning, comprising:
[0006] The acquisition step involves obtaining environmental state information of the energy system at the target time.
[0007] The environmental status information includes global status information and local status information; the local status information includes the status of renewable energy units, cogeneration units, energy storage units, and load units.
[0008] A mapping function is used to extract the Riemannian manifold features of each unit state in the local state information;
[0009] The attention weights of each local decision network are evaluated based on the characteristics of each Riemannian manifold.
[0010] Each local decision network is selected based on its attention weights to obtain at least one key decision network.
[0011] Calculate the manifold feature distribution and the predicted manifold distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network, and obtain the local decision variables by combining them with the local decision network corresponding to the unit state input; input the global state information, the manifold feature distribution and the predicted manifold distribution at the next time step into the global decision network to obtain the global decision variables;
[0012] Construct a reward function based on minimizing the total cost of low-carbon scheduling in the energy system;
[0013] The local decision networks are trained based on environmental state information, local decision variables, global decision variables, reward functions, and deep reinforcement learning techniques to obtain the trained local decision networks.
[0014] Obtain the current environmental state information of the energy system; based on the current environmental state information and each trained local decision network, obtain the target local decision variables and apply them to the energy system.
[0015] Furthermore, the above calculation of the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network includes:
[0016] A manifold feature population distribution function is constructed based on the drift function and the diffusion function; the empirical distribution of the manifold feature corresponding to each key decision network is calculated and input into the manifold feature population distribution function to obtain the manifold feature population distribution; the manifold feature population distribution is input into the discretized manifold feature population distribution function to obtain the predicted population distribution at the next time step.
[0017] Furthermore, the aforementioned training of each local decision network based on environmental state information, local decision variables, global decision variables, reward functions, and deep reinforcement learning techniques yields the trained local decision networks, including:
[0018] By inputting global decision variables, local decision variables, global state information, and manifold characteristics into the global evaluation network and the target evaluation network, the decision evaluation value and the decision target value are obtained.
[0019] By inputting environmental state information and global decision variables into the reward function, the decision scheduling cost is obtained.
[0020] State transition is performed based on global decision variables, local decision variables, and global state information to obtain local state information and global state information at the next time step.
[0021] Calculate the actual population distribution at the next moment based on the local state information at the next moment;
[0022] The decision scheduling cost, global state information at the target time and the next time, the predicted population distribution and the actual population distribution at the next time, local decision variables, global decision variables and manifold feature population distribution are put into the experience pool;
[0023] The target sample is obtained by uniformly sampling the experience pool;
[0024] The global evaluation network and the target evaluation network are updated based on the target sample, the decision evaluation value, and the decision target value.
[0025] Update the global decision network and each local decision network based on the updated global evaluation network;
[0026] Take the next time step after the target time as the target time, return to the acquisition step, and continue until the target time reaches the preset duration or the number of loops reaches the preset number of iterations, to obtain the trained local decision networks.
[0027] Furthermore, the method also includes:
[0028] The network parameters of the mapping function, the learning rate of the drift function, and the learning rate of the diffusion function are updated based on the target sample.
[0029] Furthermore, the global status information includes electricity load demand, heat load demand, gas load demand, renewable energy output, cumulative carbon emissions of the energy system, electricity price, carbon price, and energy system storage status.
[0030] Furthermore, global decision variables include local decision variables and the power purchased by the energy system from the grid;
[0031] Local decision variables include the scheduling actions of renewable energy units, cogeneration units, energy storage units, and load units. The scheduling action of renewable energy units is the actual power of the renewable energy units; the scheduling action of cogeneration units is the active power output and thermal power output of the cogeneration units; the scheduling action of energy storage units is the charging and discharging active power of the energy storage units; and the scheduling action of load units is the amount of abandoned active power load, abandoned reactive power load, and abandoned thermal load of the load units.
[0032] Furthermore, the reward function constructed above based on minimizing the total cost of low-carbon dispatching of the energy system includes:
[0033] Based on global state information, local state information, global decision variables, fuel unit cost, carbon emission allocation per unit of electricity, cost of curtailing wind and solar power, cost of curtailing active power load, cost of curtailing reactive power load, and cost of curtailing heat load, calculate operating costs, carbon costs, cost of curtailing wind and solar power load, and penalties for exceeding limits.
[0034] Subtracting the operating costs, carbon costs, and costs of curtailing wind, solar, and load from the penalty for exceeding limits yields the reward function.
[0035] Secondly, embodiments of this application provide a low-carbon scheduling device for an energy system based on deep reinforcement learning, comprising:
[0036] The acquisition module is used to acquire environmental state information of the energy system at a target time; the environmental state information includes global state information and local state information; the local state information includes the state of renewable energy units, cogeneration units, energy storage units, and load units.
[0037] The extraction module is used to extract the Riemannian manifold features of each unit state in the local state information using a mapping function;
[0038] The weighting module is used to evaluate the attention weights of each local decision network based on the characteristics of each Riemannian manifold.
[0039] The filtering module is used to filter each local decision network according to each attention weight to obtain at least one key decision network.
[0040] The decision module is used to calculate the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network, and to obtain local decision variables by inputting the unit state into the corresponding local decision network; and to input the global state information, the manifold feature population distribution and the predicted population distribution at the next time step into the global decision network to obtain global decision variables.
[0041] The cost constraint module is used to construct a reward function based on minimizing the total cost of low-carbon scheduling of the energy system.
[0042] The training module is used to train each local decision network based on environmental state information, local decision variables, global decision variables, reward function and deep reinforcement learning techniques, so as to obtain each trained local decision network.
[0043] The application module is used to acquire the current environmental state information of the energy system; based on the current environmental state information and each trained local decision network, the target local decision variables are obtained and applied to the energy system.
[0044] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the steps of the energy system low-carbon scheduling method based on deep reinforcement learning as described in any of the above embodiments.
[0045] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed, implements the steps of the energy system low-carbon scheduling method based on deep reinforcement learning as described in any of the above embodiments.
[0046] In summary, compared with the prior art, the beneficial effects of the technical solution provided in this application include at least the following:
[0047] This application provides a low-carbon scheduling method for energy systems based on deep reinforcement learning. By extracting Riemannian manifold features and analyzing population distribution, the input structure of the decision network is modified. The input of the global decision network in existing technologies is changed from environmental state information to global state information plus manifold feature population distribution and the predicted population distribution at the next time step. Similarly, the input of a single local decision network in existing technologies is changed from local state information to manifold feature population distribution, the predicted population distribution at the next time step, and its corresponding unit state. This significantly reduces the input dimensionality of the decision network and the complexity of decision search, thereby effectively improving algorithm efficiency. It should also be noted that if only the Riemannian manifold features of the key decision network are introduced to update the input of the decision network, the scheduling algorithm will only focus on the unit state corresponding to the key decision network, ignoring other unit states in the local state information. This results in limited representation ability and insufficient decision stability in the complex and ever-changing environment of energy systems. Therefore, this application further introduces population distribution analysis. This technology can analyze the temporal evolution of population features, thereby predicting the distribution characteristics of the next time step and incorporating them into the decision-making process, thus enhancing the decision stability of the scheduling algorithm in complex environments. Attached Figure Description
[0048] Figure 1 A flowchart illustrating a low-carbon scheduling method for energy systems based on deep reinforcement learning, provided as an embodiment of this application.
[0049] Figure 2 This is a structural diagram of a low-carbon scheduling device for an energy system based on deep reinforcement learning, provided as an embodiment of this application. Detailed Implementation
[0050] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0051] Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0052] Please see Figure 1This application provides a low-carbon scheduling method for energy systems based on deep reinforcement learning, including:
[0053] Step S1: Acquisition Step - Acquire the environmental state information of the energy system at the target time.
[0054] The environmental status information includes global status information and local status information.
[0055] Global state information S g This includes electricity load demand, heat load demand, gas load demand, renewable energy output, cumulative carbon emissions of the energy system, electricity price, carbon price, and energy system storage status.
[0056] Specifically, global state information refers to the shared state information obtained by the energy system dispatch center, which consists of 10 vectors:
[0057]
[0058] In the formula, t is the target time, which can also be considered as the time when a decision needs to be made; The predicted electrical load demand at time t; The predicted heat load demand at time t; The predicted gas load demand at time t; The predicted renewable energy output value at time t; Let be the cumulative carbon emissions of the energy system at time t; and The electricity price and carbon price at time t are respectively; SOC t Let t be the energy storage state of the energy system, satisfying SOC. t ∈[0,1].
[0059] Local state information includes multiple unit states: renewable energy unit state Cogeneration unit status Energy storage unit status and load unit status Specifically, the local state information S l The specific definitions are as follows:
[0060]
[0061] In the formula, RE, CHP, SOC, and LOAD are the node sets of the renewable energy unit, hotspot cogeneration unit, energy storage unit, and load unit, respectively. The maximum active power output of node i is predicted at time t; s i,t The value represents the working status of node i at time t; 1 indicates online and 0 indicates offline. Let t-1 be the active power output of node i in the renewable energy unit at the previous time. The maximum heat output of node i in the cogeneration unit at time t; and Let be the power generation efficiency and thermal efficiency of node i in the cogeneration unit at time t, respectively. and Let be the charge / discharge efficiency and energy storage status of the energy storage device at energy storage unit node i at time t, respectively. satisfy and Let S represent the active load, reactive load, and predicted heat demand of load unit node i at time t. In subsequent embodiments, all local state information will be denoted as S. l,i,t .
[0062] Step S2: Use a mapping function to extract the Riemannian manifold features of each unit state in the local state information.
[0063] Taking the definition of local state information given in the above embodiment as an example, it is equivalent to extracting the Riemann flow pattern features corresponding to the renewable energy unit state, cogeneration unit state, energy storage unit state and load unit state respectively.
[0064] Traditional feature extraction methods struggle to accurately capture the structure of local states, resulting in poor attention sampling performance. Therefore, this application selects the Riemannian manifold space as the target space for projection and constructs a mapping function ψ(S l,i ;θ ψ ):S l →M extracts features from local state information to assist attention sampling, thereby improving the accuracy of subsequent key decision network sampling. Specifically, a mapping function ψ(S) is constructed. l,i,t ;θ ψ ), will transfer the local state information S l,i,t Projecting onto the Riemannian manifold space captures its structure. That is:
[0065] e i (t)=ψ(S l,i,t ;θ ψ )
[0066] Among them, e i (t) is S l,i,t Riemannian characteristics in Riemannian manifold space.
[0067] θ ψ The network parameters, which are mapping functions, are trained by minimizing the following loss function:
[0068]
[0069] in, This is a distance-preserving loss term, used to ensure that the distance on the flow pattern is consistent with the original state distance; S l,i ,S l,j λ represents the local state information sampled from the experience pool. c ‖R μνρσ || 2 R is the curvature regularization term, used to ensure the smoothness of the manifold curve; μνρσ For Riemann curvature tensor; λ is a topology-preserving term used to maintain the topological structure of the data. c With λ t These are the curvature coefficient and the topology coefficient, respectively, used to dynamically adjust the weights of various parameters. It is the k-dimensional Bertie number, used to provide feedback on topological features; To provide dynamically equivariant constraints, ensuring consistent dynamic changes before and after embedding. For Christofel's symbol; and For the linear combination of Christofel symbols, it is used to measure the change of the connection itself on the manifold; and For the nonlinear combination term of the Christofel notation, used to reflect the self-interactions caused by manifold bending; g λκ To measure the inverse of a tensor, satisfying g λκ g κμ =δ λμ g μν Let be the Riemannian metric of the mapping function, where A(S) represents the partial derivative of the embedding function ψ with respect to the manifold coordinates e; g ) is derived from global state information S g A parameterized symmetric positive definite matrix, used to characterize the importance of different state orientations; λ r δ is the regularization coefficient, used to control the strength of the regularization term; μν The Kronecker delta function represents the Euclidean metric.
[0070] The above steps can construct a mapping function that accurately extracts the features of local state manifolds, intuitively reflecting the characteristics of local states.
[0071] Step S3: Evaluate the attention weights of each local decision network based on the characteristics of each Riemannian manifold.
[0072] Step S4: Filter each local decision network according to its attention weights to obtain at least one key decision network. Based on the Riemannian manifold features e corresponding to each extracted unit state... i(t), this application further employs a dynamic sampling mechanism to select key local decision networks. Since the states and decisions of each local decision network are coupled, the unit state corresponding to each local decision network can, to a certain extent, characterize the overall state of the current energy system. Therefore, this application, based on e i The magnitude of (t) is used to assess the importance of each local decision network, and the set of key decision networks Δ with a sampling capacity of k is sampled accordingly, i.e.:
[0073]
[0074] st|Γ|=k
[0075] Where Γ is a set of all local decision networks; α i (t) represents the attention weights to be assigned.
[0076]
[0077] Where, α i (t) represents the normalized attention weights of each local decision network at the current time step, reflecting the ability of the unit state of local agent i to represent the states of other units; I is the set of local agents; ξ(S g,t ) = MLP key (S g,t ) and η(e i (t))=MLP value (e i (t)) is a learnable neural network, used to encode the manifold features corresponding to the global state information and the local state information, respectively; <ξ(S g,t ),η(e i (t))> represents the vector inner product; τ is a temperature parameter used to control the concentration of attention distribution. Through the above steps, the key decision network set Δ and its local state information Riemannian manifold feature set N can be obtained. t ={e i (t)|i∈Δ} can effectively represent the current state information (S) g,t ,{S l,i,t}).
[0078] Step S5: Calculate the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network, and input them into the local decision network corresponding to the unit state input to obtain local decision variables; input the global state information, manifold feature population distribution and the predicted population distribution at the next time step into the global decision network to obtain global decision variables.
[0079] In existing technologies, local decision networks π θl,i (·|S g,t ,{Sl,i,t}) is a deep neural network deployed on local agents. In the existing training and execution phases, this network relies on the current environmental state information (S) g,t ,{S l,i,t}), generate the local decision variable A of the corresponding scheduling unit. l,i,t Global decision network π θg (·|S g,t ,{S l,i,t}) is a deep neural network in a global intelligent agent, which calculates based on the currently observed state information (S) g,t ,{S l,i,t Output the global decision vector A. g,t To maximize the value of state-action pairs.
[0080] The global decision variables include local decision variables and the power purchased by the energy system from the grid.
[0081] Global decision variables include local decision variables and the power that the energy system purchases from the grid.
[0082]
[0083] Among them, {A l} is the set of local decision variables. Let t be the power that the energy system purchases from the grid at time t (i.e., the active power exchanged with the external grid). A positive value indicates the purchase of electricity, and a negative value indicates the sale of electricity.
[0084] Local decision variables include renewable energy unit scheduling actions. Cogeneration unit scheduling actions Energy storage unit scheduling actions and load unit scheduling actions The scheduling actions for renewable energy units are based on their actual power output; for combined heat and power (CHP) units, they are based on their active and thermal output; for energy storage units, they are based on their charging and discharging active power; and for load units, they are based on their abandoned active, reactive, and thermal loads. The specific definitions of the scheduling actions for each unit are as follows:
[0085]
[0086] In the formula, For the actual output of renewable energy unit node i at time t, satisfying and Let be the active power output and thermal power output of node i in the cogeneration unit at time t, respectively, satisfying... and Let be the active power of charging and discharging at node i of the energy storage unit at time t, where a positive value represents charging and a negative value represents discharging, satisfying the following conditions: and Let be the active power load, reactive power load, and heat load of the load unit node at time t, respectively, satisfying the following conditions: and To simplify the expression, all local decision variables will be denoted as A from now on. l,i,t .
[0087] However, due to the uncertainties and fluctuations in the source-load relationship, real-world integrated energy systems exhibit complex and variable states during operation, making accurate characterization difficult. The aforementioned screening of key decision networks ignores the states of other units, resulting in limited characterization capabilities and insufficient decision stability in complex and variable environments. To address this issue, this application further proposes a population stochastic neural differential equation modeling technique. This technique continuously models population dynamics, analyzes the temporal evolution of population characteristics, predicts the distribution characteristics of the next time step, and incorporates them into the decision-making process, thereby enhancing the decision stability of the scheduling algorithm in complex environments.
[0088] Specifically, a manifold feature population distribution function is constructed based on the drift function and the diffusion function; the empirical distribution of the Riemann manifold feature corresponding to each key decision network is calculated and input into the manifold feature population distribution function to obtain the manifold feature population distribution; the manifold feature population distribution is input into the discretized manifold feature population distribution function to obtain the predicted population distribution at the next time step.
[0089] First, calculate the empirical distribution function and the Riemannian manifold feature set N of the key decision network sample Δ. t As a set of discrete variables, it is difficult to directly extract time-series features. Therefore, an empirical distribution function is constructed to transform the discrete features into a continuous probability distribution, so as to quantitatively analyze their time-series evolution. Specifically, the empirical distribution function μ is defined. Δ (t) is as follows:
[0090]
[0091] Where δ is the Dirac function.
[0092] The state evolution of an integrated energy system is a stochastic process dominated by state transition equations and influenced by source-load uncertainty fluctuations. The actual evolution process can be viewed as a superposition of a deterministic process and a stochastic process, making it difficult to accurately analyze its temporal characteristics using traditional methods. Therefore, this application employs stochastic neural differential equations to model the population distribution continuously over time, thereby extracting its temporal characteristics. Specifically, the population distribution function ρ, representing the flow pattern at time t, is constructed. t As shown in the following formula:
[0093]
[0094] stρ t =μ Δ (t)
[0095] in, f is the expectation of the population distribution function; θ (s g ,e,ρ t ) is the drift function, which determines the trend of population distribution evolution; the update method of its network parameter θ is given below; e is the distribution function ρ. t A point on; σ φ (s g ,ρ t ) is the diffusion function, which determines the strength and direction of the random perturbation in the population distribution evolution. The update method of its network parameter φ is given below; dW t Let be the infinitesimal increment of Brownian motion, representing unpredictable random noise in the environment, satisfying...
[0096] The population distribution function constructed above reflects the evolution of population distribution characteristics over continuous time. Essentially, it is a continuous function and cannot be directly solved on a computer. Therefore, this application further performs Euler discretization on the above continuous equation to obtain the predicted population distribution at the next time step.
[0097]
[0098] in, Let be a random vector sampled from a standard Gaussian distribution. For the discretized |e i | Dimensional random shock.
[0099] Through the above steps, the temporal evolution characteristics of the manifold feature distribution can be analyzed, and the population distribution at the next time step can be predicted. These steps analyze the spatial distribution information of the states from a temporal perspective, effectively characterizing the evolutionary patterns of state information. Therefore, in this application, the inputs to the global decision network and the local decision network can be redefined as follows:
[0100]
[0101] This clearly demonstrates the reduction in dimensionality of the inputs to both global and local decision networks compared to existing technologies.
[0102] Step S6: Construct a reward function based on minimizing the total cost of low-carbon scheduling of the energy system.
[0103] Specifically, the reward function is the objective function. The goal of low-carbon dispatching of the energy system is to minimize the total cost and maintain stable operation. Therefore, this application calculates the operating cost, carbon cost, wind and solar curtailment cost, and over-limit penalty based on global state information, local state information, global decision variables, fuel unit cost, carbon emission allocation per unit of electricity, cost of curtailing wind and solar power, cost of curtailing active power load, cost of curtailing reactive power load, and cost of curtailing thermal load; then, the over-limit penalty is set... Subtract operating costs carbon cost Costs of curtailing wind, solar, and load Obtain the reward function r(S) g,t ,{S l,i,t},A g,t The formula is defined as follows:
[0104]
[0105] Where, π fuel,i The unit cost of fuel used by the unit; h is the carbon emission allocation per unit of electricity; η i The cost per unit of wind and solar curtailment; η P,i η Q,i With η H,i These represent the costs of curtailed active power load, curtailed reactive power load, and curtailed heat load, respectively; d t d represents a restricted system variable; max and d min These correspond to the upper and lower limits of the variable, respectively.
[0106] Step S7: Train each local decision network based on environmental state information, local decision variables, global decision variables, reward function, and deep reinforcement learning techniques to obtain the trained local decision networks.
[0107] Specifically, the above-mentioned training of each local decision network based on environmental state information, local decision variables, global decision variables, reward function, and deep reinforcement learning techniques yields the trained local decision networks, including:
[0108] Step S71: Input the global decision variables, local decision variables, global state information, and flow pattern characteristic group distribution into the global evaluation network and the target evaluation network to obtain the decision evaluation value and the decision target value.
[0109] Global Evaluation Network Q ω (S g,t ,{S l,i,t},A g,t ) is a deep neural network in a global intelligent agent used to evaluate the value of state-action pairs, obtain decision evaluation values, and feed them back to the global decision network to guide its updates.
[0110] Target Evaluation Network Q ω′ (S g,t ,{S l,i,t},A g,t The target network is the one corresponding to the global evaluation network. The update trend of the global evaluation network is to minimize the mean square error between the decision evaluation value and the decision target value output by the target evaluation network.
[0111] As can be seen from the above embodiments, the current manifold feature population distribution ρ is extracted from the key decision network set Δ. t and the predicted population distribution at the next moment. They analyze the spatial distribution information of states from a temporal perspective, effectively characterizing the evolutionary patterns of state information. Therefore, the input to the global / target evaluation network is also redefined in this application as follows:
[0112] Q ω (S g,t ,ρ t A g,t ,{A l,i,t})
[0113] Q ω′ (S g,t ,ρ t A g,t ,{A l,i,t})
[0114] Step S72: Input the environmental state information and global decision variables into the reward function to obtain the decision scheduling cost.
[0115] Step S73: Based on the global decision variables, local decision variables, and global state information, perform state transition to obtain the local state information and the global state information at the next time step.
[0116] Specifically, first define the local state transition function based on the global state information and local decision variables:
[0117] S l,t+1 =p(g|S g,t A l,t )
[0118] The local state transition function is used to calculate the local state information S of the energy system at the next time step. l,t+1 The next moment
[0119] s i,t , and The local state information at the next time step is obtained by adding random perturbations to the known global state information at time t, and can be calculated using the following formula:
[0120]
[0121] Among them, E i Δt represents the upper limit of energy storage capacity of the energy storage device; Δt represents the time interval between adjacent decisions.
[0122] Then, based on the global state information and global decision variables, a global state transition function is defined:
[0123] S g,t+1 =p(g|S g,t A g,t )
[0124] The global state transition function is used to calculate the global state information S of the energy system at the next moment. g,t+1 The next moment and The global state information at time t is obtained by adding random perturbations to the known global state information at time t. The remaining global state information at the next time step can be calculated using the following formula:
[0125]
[0126] Where, ρ i The carbon emission factor of the fuel used by the unit; μ i This represents the unit fuel consumption of unit i.
[0127] Step S74: Calculate the actual population distribution at the next time step based on the local state information at the next time step. Here, the actual population distribution ρ at the next time step is calculated. t+1 The population distribution function used is the aforementioned manifold characteristic population distribution function, which is equivalent to the actual population distribution ρ calculated at the next time step based on the local state information at the next time step. t+1 It is the predicted population distribution at the next moment. The true value.
[0128] Step S75: Put the decision scheduling cost, global state information at the target time and the next time, predicted population distribution and actual population distribution at the next time, local decision variables, global decision variables and manifold feature population distribution into the experience pool.
[0129] Specifically, the sample structure of the experience pool is also updated accordingly in this application:
[0130]
[0131] Step S76: Uniformly sample the experience pool to obtain the target sample.
[0132] Step S77: Update the global evaluation network and the target evaluation network based on the target sample, the decision evaluation value, and the decision target value; update the global decision network and each local decision network based on the updated global evaluation network.
[0133] Since this application updates the structure of the input data for these networks, the corresponding loss function J of the global evaluation network is... ω (ω), the loss function J of the global decision network θg (θ g Loss function of local decision network The corresponding adjustment is as follows:
[0134]
[0135] The parameters ω of the global evaluation network and ω′ of the target evaluation network are updated using the following formula:
[0136]
[0137] ω′=κω+(1-κ)ω′
[0138] The parameters θ of the global decision network g The update method is as follows:
[0139]
[0140] The parameters θ of the local decision network l,i The update method is as follows:
[0141]
[0142] In the above formula, B represents the target sample uniformly sampled from the experience pool; r t =r(S g,t ,{S l,i,t},A g,t ) represents the decision-making and scheduling cost; γ represents the discount factor; α represents the temperature coefficient, used to balance maximizing cumulative rewards with strategy exploration; λ represents the temperature coefficient. Q κ represents the learning rate of the global evaluation network; κ represents the update parameters of the target evaluation network. θg (θ g Q in ) ω (S g,t ,ρ t A g,t ,{A l,i,t}) represents the updated global evaluation network; λ π is the learning rate for the global / local decision network.
[0143] Step S78: Take the next time after the target time as the target time, return to step S1, and continue until the target time reaches the preset duration or the number of loops reaches the preset number of iterations, to obtain the trained local decision networks.
[0144] Furthermore, the method also includes:
[0145] Step S79: Update the network parameters of the mapping function, the learning rate of the drift function, and the learning rate of the diffusion function based on the target sample.
[0146] The process of updating the network parameters of the mapping function has been described in the above embodiments and will not be repeated here.
[0147] Because the environment of a real energy system is affected by multiple factors, there is still a deviation between the predicted and actual population distribution. Therefore, this application updates the network parameters of the aforementioned drift and diffusion functions based on the gradient descent method, thereby reducing the deviation between the predicted and actual distributions. The optimization objective of the update process can be defined by the following loss function:
[0148]
[0149] Where B is the target sample from the experience pool; D KL KL divergence is used to measure the predicted population distribution at the next time step. Compared with the actual population distribution ρ t+1 The deviation. Using the above formula, parameters θ and φ can be dynamically corrected, that is:
[0150]
[0151] Where σ and χ are the learning rates of the drift function and the diffusion function, respectively.
[0152] Through the above update process, the global decision network continuously approaches the optimal strategy, while the global evaluation network can more accurately assess the value of state-action pairs. Then, the updated global evaluation network is used to guide the parameter updates of each local decision network, making them approach the optimal strategy. Finally, the trained local decision networks are applied to the low-carbon scheduling of the energy system.
[0153] Step S8: Obtain the current environmental state information of the energy system; based on the current environmental state information and each trained local decision network, obtain the target local decision variables and apply them to the energy system.
[0154] Specifically, the processing of current environmental state information is consistent with the generation process described above. It involves calculating the Riemannian manifold characteristics of each unit state, selecting key decision networks, calculating the manifold characteristic group distribution and the predicted group distribution at the next time step based on the Riemannian manifold characteristics of the unit states corresponding to the key decision networks, reconstructing the input data of the local decision networks based on the manifold characteristic group distribution, the predicted group distribution at the next time step, and the unit state and global state information, and inputting it into each trained local decision network. The output of each local decision network is used as the target local decision variable and applied to the energy system.
[0155] From a theoretical innovation perspective, this application proposes for the first time a distributed manifold sampling and swarm stochastic neural differential equation modeling technique, which is used to improve the low-carbon scheduling algorithm for integrated energy systems based on multi-agent deep reinforcement learning. This improves the algorithm's training and decision-making efficiency while enhancing its decision-making stability in complex operating environments, meeting the needs of integrated energy systems for rapid response and stable operation. From a socio-economic perspective, this application can efficiently and accurately solve low-carbon scheduling problems for integrated energy systems, minimizing system costs and carbon emissions while ensuring stable system operation, thus promoting the development of integrated energy systems.
[0156] Please see Figure 2 This application provides a low-carbon scheduling device for energy systems based on deep reinforcement learning, comprising:
[0157] The acquisition module 101 is used to acquire the environmental state information of the energy system at the target time; wherein, the environmental state information includes global state information and local state information; the local state information includes the state of renewable energy units, the state of combined heat and power units, the state of energy storage units and the state of load units.
[0158] Extraction module 102 is used to extract the Riemann manifold features of each unit state in the local state information using a mapping function.
[0159] Weight module 103 is used to evaluate the attention weights of each local decision network based on the characteristics of each Riemannian manifold.
[0160] The filtering module 104 is used to filter each local decision network according to each attention weight to obtain at least one key decision network.
[0161] The decision module 105 is used to calculate the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemann manifold features corresponding to each key decision network, and to obtain local decision variables by inputting the unit state into the corresponding local decision network; and to input the global state information, the manifold feature population distribution and the predicted population distribution at the next time step into the global decision network to obtain global decision variables.
[0162] Cost constraint module 106 is used to construct a reward function based on minimizing the total cost of low-carbon scheduling of the energy system.
[0163] Training module 107 is used to train each local decision network based on environmental state information, local decision variables, global decision variables, reward function and deep reinforcement learning techniques, so as to obtain each trained local decision network.
[0164] Application module 108 is used to obtain the current environmental state information of the energy system; based on the current environmental state information and each trained local decision network, the target local decision variables are obtained and applied to the energy system.
[0165] The specific limitations of the energy system low-carbon scheduling device based on deep reinforcement learning provided in this embodiment can be found in the embodiment of the energy system low-carbon scheduling method based on deep reinforcement learning described above, and will not be repeated here. Each module in the aforementioned energy system low-carbon scheduling device based on deep reinforcement learning can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0166] This application provides a computer device that may include a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it causes the processor to perform the steps of a deep reinforcement learning-based low-carbon scheduling method for energy systems, as described in any of the above embodiments. The working process, details, and technical effects of the computer device provided in this embodiment can be found in the above embodiments regarding a deep reinforcement learning-based low-carbon scheduling method for energy systems, and will not be repeated here.
[0167] This application provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the steps of a low-carbon scheduling method for an energy system based on deep reinforcement learning, as described in any of the above embodiments. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0168] The working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment can be found in the embodiment above regarding a low-carbon scheduling method for energy systems based on deep reinforcement learning, and will not be repeated here.
[0169] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0171] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A low-carbon scheduling method for energy systems based on deep reinforcement learning, characterized in that, include: The acquisition step involves obtaining environmental state information of the energy system at the target time. The environmental status information includes global status information and local status information; the local status information includes the status of renewable energy units, cogeneration units, energy storage units, and load units. The Riemann manifold features of each unit state in the local state information are extracted using a mapping function; The attention weights of each local decision network are evaluated based on the characteristics of each Riemannian manifold. Each of the local decision networks is filtered according to the attention weights to obtain at least one key decision network; Calculate the manifold feature distribution and the predicted manifold distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network, and obtain the local decision variables by combining them with the local decision network corresponding to the unit state input; input the global state information, the manifold feature distribution and the predicted manifold distribution at the next time step into the global decision network to obtain the global decision variables; A reward function is constructed based on minimizing the total cost of low-carbon scheduling of the energy system. Based on the environmental state information, the local decision variables, the global decision variables, the reward function, and deep reinforcement learning techniques, each of the local decision networks is trained to obtain the trained local decision networks. Obtain the current environmental state information of the energy system; based on the current environmental state information and each trained local decision network, obtain the target local decision variables and apply them to the energy system.
2. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 1, characterized in that, The step of calculating the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemann manifold features corresponding to each key decision network includes: constructing a manifold feature population distribution function based on the drift function and the diffusion function; calculating the empirical distribution based on the Riemann manifold features corresponding to each key decision network and inputting it into the manifold feature population distribution function to obtain the manifold feature population distribution; and inputting the manifold feature population distribution into the discretized manifold feature population distribution function to obtain the predicted population distribution at the next time step.
3. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 2, characterized in that, The process of training each of the local decision networks based on the environmental state information, the local decision variables, the global decision variables, the reward function, and deep reinforcement learning techniques to obtain the trained local decision networks includes: The global decision variables, the local decision variables, the global state information, and the manifold characteristic group distribution are input into the global evaluation network and the target evaluation network to obtain the decision evaluation value and the decision target value. The environmental state information and the global decision variables are input into the reward function to obtain the decision scheduling cost; Based on the global decision variables, the local decision variables, and the global state information, a state transition is performed to obtain the local state information and the global state information at the next time step. Calculate the actual population distribution at the next moment based on the local state information at the next moment; The decision scheduling cost, the global state information at the target time and the next time, the predicted population distribution and the actual population distribution at the next time, the local decision variables, the global decision variables, and the manifold feature population distribution are placed into the experience pool; The experience pool is uniformly sampled to obtain the target sample; The global evaluation network and the target evaluation network are updated based on the target sample, the decision evaluation value, and the decision target value. The global decision network and each of the local decision networks are updated based on the updated global evaluation network. Take the next moment after the target moment as the target moment, return to the acquisition step, and continue until the target moment reaches a preset duration or the number of loops reaches a preset number of iterations, to obtain the trained local decision networks.
4. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 3, characterized in that, Also includes: The network parameters of the mapping function, the learning rate of the drift function, and the learning rate of the diffusion function are updated based on the target sample.
5. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 3, characterized in that, The global status information includes electricity load demand, heat load demand, gas load demand, renewable energy output, cumulative carbon emissions of the energy system, electricity price, carbon price, and energy system energy storage status.
6. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 5, characterized in that, The global decision variables include local decision variables and the power purchased by the energy system from the grid; the local decision variables include renewable energy unit scheduling actions, cogeneration unit scheduling actions, energy storage unit scheduling actions, and load unit scheduling actions. The renewable energy unit scheduling action is the actual power of the renewable energy unit; The scheduling action of the cogeneration unit refers to the active power output and thermal output of the cogeneration unit. The energy storage unit scheduling action is the charging and discharging active power of the energy storage unit; The load unit scheduling action is the amount of active power load abandoned, reactive power load abandoned, and heat load abandoned by the load unit.
7. The low-carbon scheduling method for energy systems based on deep reinforcement learning according to claim 6, characterized in that, The construction of the reward function based on minimizing the total cost of low-carbon scheduling of the energy system includes: Based on the global state information, the local state information, the global decision variables, the unit cost of fuel, the carbon emission allocation per unit of electricity, the unit cost of curtailing wind and solar power, the cost of curtailing active power load, the cost of curtailing reactive power load, and the cost of curtailing thermal load, calculate the operating cost, carbon cost, cost of curtailing wind and solar power load, and penalty for exceeding limits; subtract the operating cost, carbon cost, and cost of curtailing wind and solar power load from the penalty for exceeding limits to obtain the reward function.
8. A low-carbon scheduling device for energy systems based on deep reinforcement learning, characterized in that, include: The acquisition module is used to acquire environmental state information of the energy system at a target time; wherein, the environmental state information includes global state information and local state information; the local state information includes the state of renewable energy units, the state of combined heat and power units, the state of energy storage units, and the state of load units; The extraction module is used to extract the Riemann manifold features of each unit state in the local state information using a mapping function; The weighting module is used to evaluate the attention weights of each local decision network based on the characteristics of each Riemannian manifold. The filtering module is used to filter each local decision network according to each attention weight to obtain at least one key decision network. The decision module is used to calculate the manifold feature population distribution and the predicted population distribution at the next time step based on the Riemannian manifold features corresponding to each key decision network, and to obtain local decision variables by inputting the unit state into the corresponding local decision network; and to input the global state information, the manifold feature population distribution and the predicted population distribution at the next time step into the global decision network to obtain global decision variables. The cost constraint module is used to construct a reward function based on minimizing the total cost of low-carbon scheduling of the energy system; The training module is used to train each of the local decision networks based on the environmental state information, the local decision variables, the global decision variables, the reward function, and deep reinforcement learning techniques, so as to obtain each trained local decision network. The application module is used to acquire the current environmental state information of the energy system; obtain the target local decision variables based on the current environmental state information and each trained local decision network, and apply them to the energy system.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the energy system low-carbon scheduling method based on deep reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed, it implements the steps of the energy system low-carbon scheduling method based on deep reinforcement learning as described in any one of claims 1 to 7.
Citation Information
Cited By
DCS automatic control system for natural gas purification station technological process
CN121995893A