A low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system
By combining the third-generation non-dominated sorting genetic algorithm with deep reinforcement learning, the problem of insufficient adaptability of genetic strategies in the electricity-hydrogen-heat integrated energy system is solved, generating an efficient and uniform Pareto solution set, and optimizing the low-carbon economic operation of the system.
Patent Information
- Application Number
- CN202511349629.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-22
AI Technical Summary
In existing optimization methods for low-carbon economic operation of integrated electric-hydrogen-heat energy systems, traditional algorithms struggle to adaptively adjust genetic strategies, resulting in low efficiency when exploring the Pareto front and difficulty in generating uniformly distributed non-dominated solution sets, thus failing to effectively balance operating costs and carbon emission targets.
We employ a third-generation non-dominated sorting genetic algorithm combined with deep reinforcement learning. By constructing a Markov decision process and a dual-deep Q-network, we achieve adaptive selection of genetic operations, generate a uniformly distributed Pareto solution set, use a multi-level reward mechanism to guide the genetic operation strategy, and combine a reference point association selection strategy to optimize population evolution.
It significantly improves the convergence speed of the algorithm and the quality of the Pareto solution set, reduces system operating costs and optimizes carbon emissions, and generates better decision-making schemes.
Smart Images

Figure CN120850811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of intelligent energy systems and artificial intelligence, and specifically relates to a low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system. BACKGROUND
[0002] As a new type of energy system, the electricity-hydrogen-heat integrated energy system can realize the collaborative optimization of multiple energy forms by integrating electrolytic hydrogen production, hydrogen fuel cell power generation, electric heat conversion and multi-type energy storage systems, and provides a new technical path for the low-carbon economic operation of energy systems.
[0003] However, the optimal operation of the electricity-hydrogen-heat integrated energy system faces many challenges. First, the system contains multiple energy forms and energy conversion equipment, and there is a complex coupling relationship between the subsystems, resulting in a highly nonlinear characteristic of the overall system. Second, there is usually a contradiction between reducing operating costs and reducing carbon emissions, and a reasonable balance needs to be sought between the two. In addition, system operation also needs to consider the start-stop state of various devices, climbing constraints, and the charging and discharging characteristics of the energy storage system and other constraint conditions.
[0004] For multi-objective optimization problems, existing solution methods mainly include traditional mathematical programming methods and evolutionary algorithm methods. Traditional mathematical programming methods such as weighted sum method, constraint method, etc., although have high calculation accuracy, can only obtain a single optimal solution and are difficult to fully reflect the trade-off relationship between multiple objectives. Evolutionary algorithms such as non-dominated sorting genetic algorithm and multi-objective particle swarm optimization algorithm can generate a series of non-dominated solutions in one run, but usually use fixed crossover and mutation strategies in the genetic operation process, which cannot adaptively adjust the operation parameters according to the population characteristics in the optimization process, and is prone to cause the algorithm to fall into local optimum or slow convergence speed when exploring the complex Pareto front. Existing reinforcement learning methods have been applied to the solution of single-objective optimization problems, but their application in the field of multi-objective optimization is not sufficient enough. In particular, how to organically combine reinforcement learning with evolutionary algorithms, and use the adaptive decision-making ability of reinforcement learning to guide the selection of genetic operations in evolutionary algorithms, so as to improve the efficiency and quality of multi-objective optimization, is a direction worthy of further research.
[0005] In summary, in the problem of low-carbon economic dual-objective operation optimization of the electricity-hydrogen-heat integrated energy system, how to design an algorithm that can adaptively adjust the genetic strategy, efficiently explore the Pareto front and generate a uniformly distributed non-dominated solution set, has become a technical problem to be solved.
[0006] SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides a low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system. The present application fully considers the complex coupling relationship between electricity, heat and hydrogen energy in the electricity-hydrogen-heat integrated energy system, as well as the trade-off between the two mutually contradictory objectives of operation cost and carbon emission, on this basis, solves the technical problem that the genetic operation strategy of the existing multi-objective optimization algorithm is fixed and cannot be adaptively adjusted according to the population characteristics, and makes up for the deficiencies of the traditional method in that the efficiency is low when exploring the Pareto frontier and it is difficult to generate a uniformly distributed non-dominated solution set, thereby providing a comprehensive and efficient optimization scheme for system operation decision-making.
[0008] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:
[0009] The present application provides a low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system, comprising:
[0010] Step 1: establishing a low-carbon economic dual-objective operation optimization problem model for an electricity-hydrogen-heat integrated energy system;
[0011] Step 2: constructing a dual-objective operation optimization problem solving framework with the third-generation non-dominated sorting genetic algorithm as the core;
[0012] Step 3: modeling the population evolution decision problem in the genetic algorithm as a Markov decision process, and constructing an interaction mechanism between the deep reinforcement learning agent and the population evolution environment;
[0013] Step 4: training the agent using the double deep Q network algorithm, and outputting genetic actions according to the environment state;
[0014] Step 5: obtaining new offspring by combining the selection strategy associated with the reference point and the genetic actions output by the agent;
[0015] Step 6: iteratively evolving the above-mentioned step 4 and step 5 until the maximum number of iterations is reached.
[0016] Further, the operation cost and carbon emission minimization problem expression of the electricity-hydrogen-heat integrated energy system is as follows:
[0017] (1),
[0018] (2),
[0019] (3),
[0020] (4),
[0021] (5),
[0022] (6),
[0023] (7),
[0024] (8),
[0025] (9),
[0026] (10),
[0027] (11),
[0028] (12),
[0029] (13),
[0030] (14),
[0031] (15),
[0032] (16),
[0033] (17),
[0034] (18),
[0035] (19),
[0036] (20),
[0037] (21),
[0038] (22),
[0039] (23),
[0040] (24),
[0041] (25),
[0042] (26),
[0043] (27),
[0044] (28),
[0045] (29),
[0046] (30),
[0047] (31),
[0048] (32),
[0049] wherein: represents the operation cost in T time period, represents the carbon emission cost in T time period, T represents the total number of optimization time slots, represents the photovoltaic curtailment cost in t time slot, represents the operation cost of electrolyzer and fuel cell, and respectively represent the hydrogen purchase cost and electricity purchase cost, represents the electricity purchase power from the grid in t time slot, represents the hydrogen purchase amount from the hydrogen market in t time slot, respectively represent the grid electricity carbon emission factor and market hydrogen carbon emission factor; represents the photovoltaic curtailment cost coefficient, represents the photovoltaic curtailment power in t time slot; and respectively represent the operation cost coefficient, start-up cost coefficient and shutdown cost coefficient of equipment x, and respectively represent the electrolyzer and fuel cell, and respectively represent the operation state, start-up state and shutdown state of equipment x, and respectively represent the operation state, start-up state and shutdown state of electrolyzer, and respectively represent the operation state, start-up state and shutdown state of fuel cell; respectively represent the electricity price in t time slot and hydrogen market price; represents the state of charge of electric energy storage in t time slot, and respectively represent the charging power and discharging power of electric energy storage in t time slot, represents the time slot length, and respectively represent the charging efficiency and discharging efficiency of the electrical energy storage; represents the maximum state of charge of the electrical energy storage; represents the maximum charging and discharging power of the electrical energy storage; represents the energy level of the thermal storage system at time slot t, and respectively represent the charging power and discharging power of the thermal energy storage at time slot t, and respectively represent the charging efficiency and discharging efficiency of the thermal energy storage; represents the maximum charging and discharging power of the thermal energy storage; represents the maximum energy level of the thermal storage system; represents the hydrogen storage amount of the hydrogen tank at time slot t, represents the input power of the electrolyzer at time slot t, represents the hydrogen input amount of the fuel cell at time slot t; represents the maximum hydrogen storage amount of the hydrogen tank; represents the maximum input power of the electrolyzer; represents the electrical power output by the fuel cell at time slot t, represents the maximum output power of the fuel cell, and respectively represent the conversion efficiencies of the fuel cell from hydrogen to electricity and from hydrogen to heat; represents the maximum hydrogen purchase amount from the hydrogen market; represents the input power of the electric boiler at time slot t; represents the photovoltaic power generation at time slot t, represents the electrical load demand at time slot t, represents the thermal load demand at time slot t, represents the electrolyzer efficiency; represents the electric boiler efficiency; represents the maximum input power of the electric boiler, where the decision variable is , , , , , , , , , , and .
[0050] Further, the third generation non-dominated sorting genetic algorithm core framework is as follows:
[0051] Step 2-1: initialize the first generation population with a population size of , , is a candidate solution, , , , , , , , , , , , where is the electricity purchased from the grid for T time slots, is the hydrogen purchased from the hydrogen market for T time slots, is the electrolyzer input power for T time slots, is the fuel cell hydrogen input for T time slots, is the electric boiler input power for T time slots, is the electric energy storage charging power for T time slots, is the electric energy storage discharging power for T time slots, is the thermal energy storage charging power for T time slots, is the thermal energy storage discharging power for T time slots, is the operating state of the electrolyzer for T time slots, is the operating state of the fuel cell for T time slots, is the PV curtailment power for T time slots.
[0052] Step 2-2: Perform crossover operation on randomly selected parents to generate offspring, the crossover operation expression is as follows:
[0053] (33),
[0054] where: is the crossover coefficient, is the offspring;
[0055] Step 2-3: Perform mutation on offspring individuals , the mutation operation expression is as follows:
[0056] (34),
[0057] where: is the mutation strength control factor, is a Gaussian distribution with mean 0 and variance ;
[0058] Step 2-4: Merge the parent population with the offspring population into , and then stratify into for fast non-dominated sorting:
[0059] (35),
[0060] (36),
[0061] where: is the first layer Pareto front, denotes the (k+1)th non-dominated front; For two solutions X and Y, X dominates Y is denoted as if and only if: , ,
[0062] where and denote the objective function values of the solution X corresponding to the operation cost and carbon emission, respectively; and denote the objective function values of the solution Y corresponding to the operation cost and carbon emission, respectively.
[0063] Step 2-5: Generate reference points with uniform distribution Calculate the distance from the normalized individual objective value to each reference point:
[0064] (37),
[0065] where: A is the number of reference points; and denote the normalized objective values related to the operation cost and carbon emission, respectively;
[0066] Step 2-6: Select reference points The nearest solution is selected to complement the next generation population;
[0067] Step 2-7: Repeat steps 2-2 to 2-6 until the maximum number of iterations G is reached.
[0068] Further, on the basis of the traditional genetic algorithm framework, reinforcement learning is introduced to intelligently select genetic operations. The population evolution process is regarded as a Markov decision process, and its state space contains information such as the current iteration progress, the objective function values of the parent individuals, and the dominance relationship, fully reflecting the characteristics of population evolution. The action space is designed as a variety of mutation and crossover operations with different parameters, allowing the agent to choose different genetic strategies according to the evolution stage. The reward mechanism comprehensively considers whether the offspring dominates the parent, the improvement degree of the objective value, and the distance from the existing solution set, guiding the agent to learn more effective genetic operation selection strategies, where the environment state of the Markov decision process is:
[0069] (38),
[0070] wherein: is the algebra to which the current algorithm iteration has arrived; is the algebra required by the total algorithm iteration; , , , are the operating cost target value of parent one, the carbon emission target value of parent one, the operating cost target value of parent two, the carbon emission target value of parent two, respectively; , are the operating cost target value normalization factor, the carbon emission target value normalization factor, respectively; is the dominance relation index, when dominates , takes 1, dominates , takes 0; the action space of multiple genetic strategies is:
[0071] (39),
[0072] wherein: is the mutation genetic strategy controlled by different mutation intensity factors; is the mixed linear a crossover genetic strategy controlled by different crossover coefficients a; each operation has different effects for different evolution stages and population characteristics; the multi-level reward mechanism considering the dominance relation and the improvement rate is as follows:
[0073] (40)
[0074] wherein: is the i-th target value of the offspring; , are the i-th target value of parent one and parent two, respectively; is the improvement rate of the offspring relative to the parent, is the minimum Euclidean distance; the improvement rate item quantifies the optimization degree of the offspring relative to the parent; is the adjustment parameter, wherein controls the strength of the dominance reward and the improvement reward, controls the amplitude of the similarity punishment, controls the decay rate of the similarity punishment.
[0075] Further, to realize this intelligent decision-making process, a double deep Q network is constructed, including an online Q network and a target Q network, both of which are multi-layer neural networks, wherein: the online Q network is a mapping from the environment state to the Q value, using Parameterization is performed, aligning the number of input layer neurons to the dimension of the environment state and the number of output layer neurons to the dimension of the action space. The network uses Rectified Linear Units (ReLU) as the activation function for the hidden layers. A target Q-network is used to stabilize the learning process; its structure is identical to that of the online Q-network. It is parameterized, and its parameters are periodically updated from an online network.
[0076] Furthermore, the training process of the dual-deep Q-network is mainly implemented through two mechanisms: experience replay and target network update. The experience replay mechanism utilizes the transfer samples obtained from the agent's interaction with the environment. Stored in the playback buffer pool, where This represents the state of the algorithm at generation g. This represents the action selected in the g-th generation. Indicates the selection action in the g-th generation. The rewards received This indicates the next state after the transition.
[0077] During the action selection process, the agent adopts actions based on its current state. - Greedy strategy for decision-making: based on probability Random exploration, with a probability of 1- Choose the action with the highest Q value. This strategy balances exploration and utilization, ensuring sample diversity.
[0078] During training, a batch of samples is randomly sampled for learning. This mechanism breaks the correlation between samples, improving learning efficiency and stability. Network parameter updates employ the Temporal Difference (TD) algorithm, and its target Q-value is calculated as follows:
[0079] (41),
[0080] in: This is a discount factor used to balance the importance of immediate and future rewards. For the target network, It represents the set of all possible actions at the next moment. The loss function of an online network is defined as the mean square value of the TD error, calculated as follows:
[0081] (42),
[0082] in: For online networks, These are the network parameters. Parameter optimization uses the stochastic gradient descent algorithm, calculated by retrieving the loss function with respect to the parameters. Update the gradient ,in For learning rate, denotes the gradient of the loss function with respect to the parameters . The parameters of the target network are then copied from the online network by a soft update manner: where is the update rate. This soft update mechanism updates the parameters of the target network slowly, ensuring the stability of the training process and avoiding the dramatic fluctuations of the value function estimates.
[0083] Further, the reinforcement learning is combined with the genetic algorithm to construct a complete interactive optimization framework:
[0084] Step 6-1: initialize the population , the population size is , and initialize the parameters of the double deep Q network and , and set the experience replay pool D;
[0085] Step 6-2: select the parent individual pair , and construct the environment state of the Markov decision process;
[0086] Step 6-3: the double deep Q network agent selects the action according to the ε-greedy strategy, and executes the corresponding mutation or crossover operator to generate offspring;
[0087] Step 6-4: evaluate the fitness of the offspring, calculate the reward value , and store the transition sample in the experience replay pool D;
[0088] Step 6-5: sample a batch of data from D, and update the online network parameters of the double deep Q network by the time difference algorithm , and update the target network synchronously every fixed number of steps;
[0089] Step 6-6: merge the parent and offspring populations, perform non-dominated sorting and reference point association, and select the next generation population;
[0090] Step 6-7: repeat steps 6-2 to 6-6 until the maximum number of iterations G is reached.
[0091] Overall, the present application includes three main modules from the system architecture: a data acquisition and processing module, an intelligent optimization decision module, and a system collaborative control module. The data acquisition and processing module is responsible for providing necessary input information for optimization decision; the intelligent optimization decision module includes a multi-objective optimization submodule and a reinforcement learning submodule, which are responsible for finding the Pareto optimal solution and adaptively adjusting the optimization strategy, respectively; the system collaborative control module is responsible for receiving the optimization decision and executing, maintaining the balanced operation of the multi-energy flow system.
[0092] Beneficial effects: A low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system, compared with prior art:
[0093] (1) Combined with double-depth Q network and third-generation non-dominated sorting genetic algorithm, adaptive selection of genetic operation is realized, and the limitations of fixed strategy of traditional algorithm are overcome.
[0094] (2) The designed multi-level reward mechanism enables the agent to adaptively select the optimal strategy according to the population evolution stage, and significantly improves the algorithm convergence speed.
[0095] (3) By combining reference point association and intelligent genetic operation, a more optimal and uniformly distributed Pareto solution set is generated. BRIEF DESCRIPTION OF DRAWINGS
[0096] Figure 1 is a flow chart of a low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system provided by the present application;
[0097] Figure 2 is a Pareto curve comparison chart of the method and other methods;
[0098] Figure 3 is the economic performance of each scheme of the method and other methods under the condition that the carbon emission constraint in the Pareto front is 136,500 kg. DETAILED DESCRIPTION
[0099] In order to clearly explain the technical scheme and innovative advantages of the present application, the specific implementation scheme of the present application will be described in conjunction with the drawings. It should be particularly pointed out that the present embodiment is only used for explanatory description, and does not constitute a limitation on the claims of the present application.
[0100] Example 1
[0101] As shown in Figure 1 , the present application provides a flow chart of a low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system, which comprises the following steps:
[0102] 1. Establishing a model of minimizing the operation cost and carbon emission of the electricity-hydrogen-heat integrated energy system, the expression is as follows:
[0103] (1),
[0104] (2),
[0105] (3),
[0106] (4),
[0107] (5),
[0108] (6),
[0109] (7),
[0110] (8),
[0111] (9),
[0112] (10),
[0113] (11),
[0114] (12),
[0115] (13),
[0116] (14),
[0117] (15),
[0118] (16),
[0119] (17),
[0120] (18),
[0121] (19),
[0122] (20),
[0123] (21),
[0124] (22),
[0125] (23),
[0126] (24),
[0127] (25),
[0128] (26),
[0129] (27),
[0130] (28),
[0131] (29),
[0132] (30),
[0133] (31),
[0134] (32),
[0135] wherein: represents the operation cost in T time period, represents the carbon emission cost in T time period, T represents the total number of optimization time slots, represents the photovoltaic curtailment cost in t time slot, represents the operation cost of electrolyzer and fuel cell, and respectively represent the hydrogen purchase cost and electricity purchase cost, represents the electricity purchased from the grid in t time slot, represents the hydrogen purchased from the hydrogen market in t time slot, respectively represent the grid electricity carbon emission factor and market hydrogen carbon emission factor; represents the photovoltaic curtailment cost coefficient, represents the photovoltaic curtailment power in t time slot; and respectively represent the operation cost coefficient, start-up cost coefficient and shutdown cost coefficient of equipment x, and respectively represent the electrolyzer and fuel cell, and respectively represent the operation state, start-up state and shutdown state of equipment x, and respectively represent the operation state, start-up state and shutdown state of electrolyzer, and respectively represent the operation state, start-up state and shutdown state of fuel cell; respectively represent the electricity price in t time slot and hydrogen market price; represents the state of charge of electric energy storage in t time slot, and respectively represent the charging power and discharging power of electric energy storage in t time slot, denotes the time slot length, and denote the electric energy storage charging efficiency and discharging efficiency, respectively; denotes the maximum state of charge of the electric energy storage; denotes the maximum charging and discharging power of the electric energy storage; denotes the energy level of the thermal storage system at time slot t, and denote the charging and discharging power of the thermal energy storage at time slot t, respectively, and denote the charging and discharging efficiency of the thermal energy storage, respectively; denotes the maximum charging and discharging power of the thermal energy storage; denotes the maximum energy level of the thermal storage system; denotes the hydrogen storage amount of the hydrogen tank at time slot t, denotes the input power of the electrolyzer at time slot t, denotes the hydrogen input amount of the fuel cell at time slot t; denotes the maximum hydrogen storage amount of the hydrogen tank; denotes the maximum input power of the electrolyzer; denotes the electric power output by the fuel cell at time slot t, denotes the maximum output power of the fuel cell, and denote the hydrogen-to-electricity and hydrogen-to-heat conversion efficiencies of the fuel cell, respectively; denotes the maximum hydrogen purchase amount from the hydrogen market; denotes the input power of the electric boiler at time slot t; denotes the photovoltaic power generation at time slot t, denotes the electric load demand at time slot t, denotes the thermal load demand at time slot t, denotes the electrolyzer efficiency; denotes the electric boiler efficiency; denotes the maximum input power of the electric boiler, wherein the decision variable is , , , , , , , , , , and .
[0136] 2. Constructing a dual-objective operation optimization problem solving framework with a third-generation non-dominated sorting genetic algorithm as the core;
[0137] Further, the third-generation non-dominated sorting genetic algorithm core framework is as follows:
[0138] Step 2-1: Initialize the population size as the first generation population , as a candidate solution, , , , , , , , , , , , where is the electricity purchase power from the grid in T time slots, is the hydrogen purchase amount from the hydrogen market in T time slots, is the electrolyzer input power in T time slots, is the fuel cell hydrogen input amount in T time slots, is the electric boiler input power in T time slots, is the electric energy storage charging power in T time slots, is the electric energy storage discharging power in T time slots, is the thermal energy storage charging power in T time slots, is the thermal energy storage discharging power in T time slots, is the operating state of the electrolyzer in T time slots, is the operating state of the fuel cell in T time slots, is the photovoltaic curtailment power in T time slots;
[0139] Step 2-2: Perform a crossover operation on the randomly selected parent to generate offspring, and the crossover operation expression is as follows:
[0140] (33),
[0141] where: is the crossover coefficient, is the offspring;
[0142] Step 2-3: Perform mutation on the offspring individual , and the mutation operation expression is as follows:
[0143] (34),
[0144] where: is the mutation strength control factor, is a Gaussian distribution with a mean of 0 and a variance of ; and
[0145] Step 2-4: Replace the parent population with offspring population merged into and then stratified into Fast non-dominated sorting is performed:
[0146] (35),
[0147] (36),
[0148] where: is the first layer of Pareto front, denotes the k+1th non-dominated front; For two solutions X and Y, X dominates Y is denoted as if and only if: , ,
[0149] wherein and respectively represent the operating cost and carbon emission objective function values corresponding to the solution X of the optimization problem; and respectively represent the operating cost and carbon emission objective function values corresponding to the solution Y of the optimization problem.
[0150] Step 2-5: Generate reference points with uniform distribution Calculate the distance from the normalized individual objective value to each reference point:
[0151] (37),
[0152] wherein A is the number of reference points; and respectively represent the normalized objective values related to the operating cost and carbon emission;
[0153] Step 2-6: Select reference points The nearest solution is selected to complement the next generation population;
[0154] Step 2-7: Repeat steps 2-2 to 2-6 until the maximum number of iterations G is reached.
[0155] 3. Model the population evolution decision problem in genetic algorithm as a Markov decision process;
[0156] In the embodiment of the present application, the environment state of the Markov decision process is:
[0157] (38),
[0158] wherein: is the number of iterations reached by the current algorithm; the number of iterations of the total algorithm; , , , are the running cost target value of parent one, the carbon emission target value of parent one, the running cost target value of parent two, and the carbon emission target value of parent two, respectively; , are the running cost target value normalization factor and the carbon emission target value normalization factor, respectively; is the dominance relation index, which takes 1 when dominates and 0 when dominates ; the action space of multiple genetic strategies is:
[0159] (39),
[0160] wherein: is a mutation genetic strategy controlled by different mutation intensity factors; is a mixed linear α crossover genetic strategy controlled by different crossover coefficients α; each operation has a differentiated effect on different evolutionary stages and population characteristics; the multi-level reward mechanism considering the dominance relation and the improvement rate is as follows:
[0161] (40)
[0162] wherein: is the i-th target value of the offspring; , are the i-th target values of parent one and parent two, respectively; is the improvement rate of the offspring relative to the parents, is the minimum Euclidean distance; the improvement rate term quantifies the optimization degree of the offspring relative to the parents; is a regulation parameter, wherein controls the strength of the dominance reward and the improvement reward, controls the amplitude of the similarity penalty, controls the decay rate of the similarity penalty.
[0163] 4. An interactive mechanism between the agent and the optimization environment is established, and the agent regards the multi-objective optimization process of the electric-hydrogen-thermal integrated energy system as an interactive evolutionary environment. In the evolutionary process of each generation of population, the agent, as a genetic strategy decision maker, selects an action from the action space containing multiple genetic strategies After the environment receives the action of the agent, it performs the corresponding genetic operation to generate offspring individuals, and calculates the operation cost and carbon emissions of the offspring individuals through the integrated energy system model. , forming a complete interactive feedback.
[0164] 5. The agent learns the optimal genetic strategy by experience replay and network parameter update, realizing the collaborative optimization of reinforcement learning and genetic algorithm;
[0165] In the embodiment of the application, the double deep Q network structure includes an online Q network and a target Q network, both of which are multi-layer neural networks, wherein: the online Q network is a mapping from the environment state to the Q value, parameterized using , the number of input layer neurons is aligned with the dimension of the environment state, the number of output layer neurons is aligned with the dimension of the action space, and the network uses a rectified linear unit (ReLU) as a hidden layer activation function; the target Q network is used to stabilize the learning process, and has the same structure as the online Q network, parameterized using , and its parameters are regularly obtained from the online network update.
[0166] In the embodiment of the application, the double deep Q network training process is mainly realized through two mechanisms of experience replay and target network update. The transition samples obtained by the agent interacting with the environment are stored in the replay buffer through the experience replay mechanism, wherein represents the state of the algorithm iteration to the gth generation, represents the action selected in the gth generation, represents the reward obtained by the action selected in the gth generation, represents the next state after transition. In the action selection process, the agent uses
[0167] - greedy strategy for decision making: with a probability random exploration, with a probability of 1- selecting the action with the maximum Q value , this strategy balances exploration and utilization, ensuring the diversity of samples. During training, a batch of samples are randomly sampled for learning, which breaks the correlation between samples and improves the learning efficiency and stability. The network parameter update uses the time difference (TD) algorithm, and the target Q value calculation formula is as follows:
[0168]
[0169] (41), wherein:
[0170] is a discount factor, balancing the importance of immediate and future rewards, is the target network, is the set of all possible actions at the next time step. The loss function of the online network is defined as the mean squared value of the TD error, calculated as follows:
[0171] (42),
[0172] where: is the online network, is the network parameter. Parameter optimization uses the stochastic gradient descent algorithm, which updates the parameters by calculating the gradient of the loss function with respect to the parameters , where is the learning rate, represents the gradient of the loss function with respect to the parameters . The parameters of the target network are then copied from the online network by soft update: , where is the update rate.
[0173] 6. Reinforcement learning is combined with genetic algorithm to build a complete interactive optimization framework;
[0174] Step 6-1: Initialize the population , the population size is , and initialize the parameters of the double deep Q network and , and set the experience replay pool D;
[0175] Step 6-2: Select parent individual pair , construct the environment state of Markov decision process ;
[0176] Step 6-3: The double deep Q network agent selects action according to the ε-greedy strategy, and executes the corresponding mutation or crossover operator to generate offspring;
[0177] Step 6-4: Evaluate the fitness of the offspring, calculate the reward value , and store the transition sample in the experience replay pool D;
[0178] Step 6-5: Sample batch data from D, update the online network parameters of the double deep Q network by the time difference algorithm, and update the target network every fixed number of steps;
[0179] Step 6-6: Merge the parent and offspring populations, perform non-dominated sorting and reference point association, and select the next generation population;
[0180] Step 6-7: Repeat steps 6-2 to 6-6 until the maximum number of iterations G is reached.
[0181] To demonstrate the effectiveness of the method of the present application, four comparative schemes are introduced.
[0182] Scheme one adopts a weighted sum method combined with a convex optimization solver, which converts the multi-objective optimization problem into a single-objective problem by assigning different weights to the operating cost and carbon emission target. The algorithm constructs a linear mixed integer programming model and uses the commercial solver (CPLEX) to solve it by calling the convex optimization Python package (CVXPY), ensuring the optimality of the single-objective problem. After completing all weight combination optimizations, non-dominated solutions are extracted from the solution set to form the Pareto front.
[0183] Scheme two adopts a multi-objective particle swarm optimization algorithm, which extends particle swarm optimization to the multi-objective field. Each particle represents an energy system scheduling scheme, and the velocity and position are updated by the individual optimal and global optimal guide. The algorithm maintains a capacity-limited external archive of non-dominated solutions and uses crowding distance to maintain population diversity.
[0184] Scheme three adopts the second generation of non-dominated sorting genetic algorithm, which uses fast non-dominated sorting to classify individuals into different front levels. For individuals in the same front, the crowding distance is used to maintain diversity. The algorithm uses tournament selection, and individual selection is based on the comprehensive evaluation of non-dominated level and crowding distance.
[0185] Scheme four adopts the third generation of non-dominated sorting genetic algorithm, which introduces a reference point mechanism to maintain population diversity and solve the limitations of the second generation of non-dominated sorting genetic algorithm in high-dimensional objective space. The algorithm maintains the same fast non-dominated sorting method, but uses a reference point-based selection strategy instead of crowding distance to obtain a more evenly distributed non-dominated solution set.
[0186] Figure 2 and Figure 3 The performance comparison results of different optimization schemes in the electric-hydrogen-thermal integrated energy system after 600 iterations to complete convergence are shown. As shown in Figure 2 , the degree of curve shift towards the lower left reflects the performance of the algorithm, and the closer the curve is to the lower left corner, the better the corresponding solution set can achieve lower operating cost and less carbon emissions. The solution set produced by the scheme of the present application presents better distribution uniformity, providing decision makers with a rich selection of schemes. Figure 3 Further quantitative analysis of the economic performance of each scheme under the same carbon emissions shows that the method of the present application achieves the optimal cost control effect compared to all comparative schemes, reducing operating costs by 3.97%-24.81% and improving the economic operation efficiency of the system.Figure 2 The data points in Figure 8 represent the Pareto-optimal solutions found, with the partial solutions being closer together and appearing to overlap on the graph.
Claims
1. A method for low-carbon economic operation optimization of an electricity-hydrogen-heat integrated energy system, characterized in that, Comprising the following steps: Step 1: Establishing an electricity-hydrogen-heat integrated energy system low-carbon economy double target operation optimization problem model; Step 2: Building a double target operation optimization problem solving framework with the third generation non-dominated sorting genetic algorithm as the core; Step 3: Modeling the population evolution decision problem in genetic algorithm as a Markov decision process, and constructing an interaction mechanism between the deep reinforcement learning agent and the population evolution environment; Step 4: The agent uses a double deep Q network algorithm for training, and outputs genetic actions according to the environment state; Step 5: Obtain new offspring by combining the selection strategy associated with the reference point and the genetic actions output by the agent; Step 6: Iterative evolution of steps 4 and 5 above until the maximum iteration number is reached. 2.The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 1, characterized in that, The operation cost and carbon emission minimization problem of the electricity-hydrogen-heat integrated energy system in step 1 is expressed as follows: (1), (2), (3), (4), (5), (6), (7), (8), (9), (10), (11), (12), (13), (14), (15), (16), (17), (18), (19), (20), (21), (22), (23), (24), (25), (26), (27), (28), (29), (30), (31), (32), wherein: represents the operation cost in T time period, represents the carbon emission cost in T time period, T represents the total number of optimization time slots, represents the PV curtailment cost in t time slot, represents the operation cost of electrolyzer and fuel cell, and respectively represent the hydrogen purchase cost and electricity purchase cost, represents the electricity purchase power from the grid in t time slot, represents the hydrogen purchase amount from the hydrogen market in t time slot, respectively represent the grid electricity carbon emission factor and market hydrogen carbon emission factor; represents the PV curtailment cost coefficient, represents the PV curtailment power in t time slot; and respectively represent the operation cost coefficient, start-up cost coefficient and shut-down cost coefficient of equipment x, and respectively represent the electrolyzer and fuel cell, and respectively represent the operation state, start-up state and shut-down state of equipment x, and respectively represent the operation state, start-up state and shut-down state of electrolyzer, and respectively represent the operation state, start-up state and shut-down state of fuel cell; respectively represent the electricity price in t time slot and the hydrogen market price; represents the state of charge of electric energy storage in t time slot, and respectively represent the charging power and discharging power of electric energy storage in t time slot, represents the time slot length, and respectively represent the charging efficiency and discharging efficiency of electric energy storage; represents the maximum state of charge of electric energy storage; represents the maximum charging and discharging power of electric energy storage; represents the energy level of thermal energy storage system in t time slot, and respectively represent the charging power and discharging power of thermal energy storage in t time slot, and respectively represent the charging efficiency and discharging efficiency of thermal energy storage; Pmax,th represents the maximum heat charging and discharging power of the thermal storage; Pmax,es represents the maximum energy level of the thermal storage system; Ht represents the hydrogen storage amount of the hydrogen tank at time t, Pelec,t represents the input power of the electrolyzer at time t, Hfc,t represents the hydrogen input amount of the fuel cell at time t; Hmax represents the maximum hydrogen storage amount of the hydrogen tank; Pmax,elec represents the maximum input power of the electrolyzer; Pfc,t represents the electric power output by the fuel cell at time t, Pmax,fc represents the maximum output power of the fuel cell, and ηh2-fc and ηh2-fc,th represent the efficiencies of the hydrogen-to-electricity and hydrogen-to-heat conversions of the fuel cell, respectively; Hmax,market represents the maximum hydrogen purchase amount from the hydrogen market; Pelec,t represents the input power of the electric boiler at time t; Ppv,t represents the photovoltaic power generation amount at time t, Deltat represents the electric load demand at time t, Deltath represents the thermal load demand at time t, ηelec represents the electrolyzer efficiency; ηelec represents the electric boiler efficiency; Pmax,elec represents the maximum input power of the electric boiler, wherein the decision variable is , , , , , , , , , , and .
3. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 2, characterized in that, The core framework of the third generation non-dominated sorting genetic algorithm in step 2 is as follows: Step 2-1: Initialize a population of size N for the first generation P0 where P is the grid electricity purchase power from the grid for T time slots,P is the hydrogen purchase amount from the hydrogen market for T time slots,P is the electrolyzer input power for T time slots,P is the fuel cell hydrogen input amount for T time slots,P is the electric boiler input power for T time slots,P is the electric energy storage charging power for T time slots,P is the electric energy storage discharging power for T time slots,P is the thermal energy storage charging power for T time slots,P is the thermal energy storage discharging power for T time slots,P is the operating state of the electrolyzer for T time slots,P is the operating state of the fuel cell for T time slots,P is the photovoltaic curtailment power for T time slots; Step 2-2: Selecting a random parent The crossover operation is performed to generate offspring, which is expressed as follows: (33), wherein: is a cross coefficient, is a progeny; Step 2-3: Selection of offspring individuals Perform mutation, mutation operation expression as follows: (34), wherein: is a mutation strength control factor, is a Gaussian distribution with mean 0 and variance . Step 2-4: Merge parent population with offspring population into and stratify into perform fast non-dominated sorting: (35), (36), wherein: is the first layer Pareto front, denotes the k+1th non-dominated front; For two solutions X and Y, X dominates Y is denoted as if and only if: , wherein and denote the objective function values of the two objectives, i.e. the operating cost and carbon emission, respectively, corresponding to the solution X of the optimization problem; and denote the objective function values of the two objectives, i.e. the operating cost and carbon emission, respectively, corresponding to the solution Y of the optimization problem; Step 2-5: Generating evenly distributed reference points Calculate the distance from the normalized individual target value to each reference point: (37), wherein: A is the number of reference points; and respectively represent normalized target values related to operating costs and carbon emissions. Step 2-6: Select reference points The closest solution to make up the next generation population; Step 2-7: Repeat steps 2-2 to 2-6 until the maximum number of iterations G is reached.
4. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 3, characterized in that, The environment state of the Markov decision process in step 3 is: (38), wherein: is the algebra to which the current algorithm iteration has arrived; is the algebra that the total algorithm iteration requires; , , , are the operating cost target value of the parent one, the carbon emission target value of the parent one, the operating cost target value of the parent two, the carbon emission target value of the parent two, respectively; , are the operating cost target value normalization factor, the carbon emission target value normalization factor, respectively; is the dominance relation index, when dominates , takes 1, dominates , takes 0; The action space containing multiple genetic strategies is: (39), wherein: a mutation genetic strategy controlled by different mutation intensity factors; a hybrid linear a-crossover genetic strategy controlled by different crossover coefficients a; The reward function is designed as follows: (40), wherein: is the i-th objective value of the offspring; , are the i-th objective values of parent one and parent two, respectively; is the improvement rate of the offspring relative to the parent, is the minimum Euclidean distance; the improvement rate term quantifies the degree of optimization of the offspring relative to the parent; is a tuning parameter, wherein controls the strength of the dominance reward and the improvement reward, controls the magnitude of the similarity penalty, controls the decay rate of the similarity penalty.
5. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 4, characterized in that, The interaction mechanism between the deep reinforcement learning agent and the population evolution environment in step 3 is designed as follows: The agent regards the multi-objective optimization process of the integrated energy system as an interactive evolutionary environment. In the evolutionary process of each generation, the agent, as a genetic strategy decision maker, observes the state , selects an action from the action space containing multiple genetic strategies , and the environment receives the action of the agent, performs the corresponding genetic operation to generate offspring individuals, calculates their operation cost and carbon emissions through the integrated energy system model, and then calculates the reward value according to the improvement of the offspring relative to the parent , forming a complete interactive feedback.
6. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 5, characterized in that, The double deep Q network structure in step 4 comprises an online Q network and a target Q network, both of which are multi-layer neural networks, wherein: the online Q network is a mapping from an environment state to a Q value, and is parameterized using an input layer neuron number aligned with an environment state dimension and an output layer neuron number aligned with an action space dimension, and the network uses a rectified linear unit (ReLU) as a hidden layer activation function; the target Q network is used for stabilizing a learning process, and has the same structure as the online Q network, and is parameterized using an input layer neuron number aligned with an environment state dimension and an output layer neuron number aligned with an action space dimension, and the network uses a rectified linear unit (ReLU) as a hidden layer activation function; the target Q network is used for stabilizing a learning process, and has the same structure as the online Q network, and is parameterized using 7. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 6, characterized in that, The double deep Q network training process of step 4 is realized through two mechanisms of experience replay and target network update, the experience replay mechanism obtains transition samples by the interaction of the agent and the environment are stored in the replay buffer pool, wherein represents the state of the algorithm iteration to the gth generation, represents the action selected in the gth generation, represents the action selected in the gth generation the reward obtained, represents the next state after transition, and during training, a batch of samples are randomly sampled for learning, and the network parameter update adopts the time difference (TD) algorithm, and the calculation formula of the target Q value is as follows: (41), where: is a discount factor balancing the importance of immediate and future rewards, is the target network, is the set of all possible actions at the next time step, the loss function of the online network is defined as the mean squared value of the TD error, which is calculated as follows: (42), in: For online networks, For the network parameters, stochastic gradient descent is used for parameter optimization, by calculating the loss function with respect to the parameters. Update the gradient ,in For learning rate, This indicates the loss function with respect to the parameters. The gradient, the parameters of the target network Then it replicates from the online network via a soft update: ,in This refers to the update rate.
8. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 7, characterized in that, The genetic action output by the step 5 intelligent agent obtains a new offspring through the combination of deep reinforcement learning and genetic algorithm Select appropriate genetic operations, and adopt a greedy strategy for action selection: with a probability Random exploration, with a probability of 1- Select the action with the maximum Q value After selecting the action, generate offspring individuals according to the corresponding genetic operator according to the operation type, and calculate the reward value .
9. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 8, characterized in that, The process of the electricity-hydrogen-heat integrated energy system low-carbon economy operation optimization method in step 6 is executed according to the following steps: Step 6-1: Initialize population , population size is , while initializing the double deep Q network parameters and , set the experience replay pool D; Step 6-2: Selecting pairs of parent individuals , the environment state of the Markov decision process ; Step 6-3: The double deep Q network agent selects an action according to an ε-greedy policy , and executes a corresponding mutation or crossover operator to generate offspring; Step 6-4: Evaluate the fitness of the offspring, calculate the reward value The transition sample is stored in the experience replay pool D; Step 6-5: Sample batch data from D, update online network parameters of double deep Q network through time difference algorithm , update target network every fixed step ; Step 6-6: Merge the parent and child populations, perform non-dominated sorting and reference point association, and select the next generation population; Step 6-7: Repeat steps 6-2 to 6-6 until the maximum number of iterations G is reached.
10. The low-carbon economic operation optimization method for an electricity-hydrogen-heat integrated energy system according to claim 1, characterized in that, The method includes a data acquisition and processing module, an intelligent optimization decision module, and a system collaborative control module; The data acquisition and processing module is responsible for collecting real-time operation data and prediction information of the electricity-hydrogen-heat integrated energy system, including photovoltaic power generation prediction, load demand prediction, and energy price information, and performing data preprocessing to provide necessary input information for optimization decision; The intelligent optimization decision module includes a multi-objective optimization submodule and a reinforcement learning submodule; among them, the multi-objective optimization submodule is responsible for finding the Pareto optimal solution between cost and carbon emissions; the reinforcement learning submodule is responsible for adaptively adjusting the optimization strategy; The system collaborative control module is responsible for receiving optimization decisions and executing them to maintain power balance, heat balance, and hydrogen energy balance, and realize the collaborative operation of the multi-energy flow system.
Citation Information
Patent Citations
Integrated energy system multi-objective optimization operation method based on improved particle swarm optimization
CN115718982A
Optimized operation method for electricity-gas comprehensive energy system
CN116070739A