An Adaptive Multi-Objective Evolutionary Method for Optimizing the Capacity of Off-Grid Wind-Solar Hydrogen Production and Storage Systems
Patent Information
- Application Number
- CN202411620214.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-11-14
AI Technical Summary
离网模式下的风光发电储能系统涉及的单元较多,模型也更加复杂,传统的基于种群的算法通常采用单一搜索策略,容易陷入局部最优,难以快速找到全局最优解
[0089] Compared with existing technologies, this invention, by adaptively selecting appropriate search strategies for individuals at different stages of evolution, can more quickly search for a set of Pareto optimal solutions for users, enabling them to formulate corresponding energy storage system capacity configuration schemes based on their preferences. The strategy pool designed in this invention introduces gradient information based on several classic differential evolution search strategies with different characteristics and simulated binary crossover operators. Using gradients to guide the algorithm's search reduces the randomness and blindness of the search to a certain extent, enhancing the algorithm's search efficiency. The policy adaptive network designed in this invention can determine whether an individual tends to converge or explore based on its different positions in the decision space and target space, reflecting this probabilistically in search strategies with different characteristics. This method solves the problem of search strategy uniformity, has good portability, and can be applied to other multimodal optimization problems. The policy adaptive network designed in this invention does not directly use the strategy with the highest probability for searching, but instead uses... Greedy strategies and roulette wheel methods increase the exploratory nature of the adaptive network in its early stages, while balancing exploration and utilization in the mid-to-late stages helps discover better strategies and reduces the risk of the network getting trapped in local optima. The gradient update of this policy network does not depend on other metrics; the reward for policy updates is determined solely by the Pareto dominance relation, making this adaptive network highly portable and applicable to other multi-objective evolutionary optimization methods.
Smart Images

Figure CN119579347B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a multi-objective optimization method, specifically an adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind and solar hydrogen production and storage systems, belonging to the field of energy storage system capacity optimization technology. Background Technology
[0002] With the deepening of the global energy revolution and the establishment of my country's "dual-carbon" goals, the proportion of renewable energy sources such as wind and solar power integrated into the power grid is constantly increasing. Building a new power system centered on renewable energy has become an important direction for the modern energy system. By the end of 2023, China's installed wind and solar power capacity reached 440 million kilowatts and 610 million kilowatts respectively, ranking first in the world. However, wind and solar power generation is characterized by significant fluctuations and intermittency, making power quality unstable and difficult to reliably supply power to loads. To ensure the stable operation of the power system, wind and solar power curtailment occurs frequently, resulting in a significant waste of renewable energy and substantial economic losses.
[0003] In addressing these challenges, hydrogen energy stands out as a clean and efficient energy source, becoming a crucial pathway to achieving the "dual carbon" goals. Hydrogen production through renewable energy-based water electrolysis (green electricity hydrogen production) not only significantly reduces environmental pollution from traditional hydrogen production (grey hydrogen) but also efficiently converts surplus electricity from wind and solar power into hydrogen, thereby improving the utilization rate of renewable energy. By absorbing surplus electricity, green electricity hydrogen production avoids wind and solar curtailment, achieving deep integration of renewable energy and full utilization of resources.
[0004] Off-grid energy storage systems play a crucial role in the green electricity-to-hydrogen process. These systems effectively address the intermittency and volatility of wind and solar power generation, ensuring power supply stability through energy storage dispatch. Furthermore, surplus electricity can be used for hydrogen production, forming a highly efficient electricity-hydrogen complementary energy system. The independence of off-grid energy storage systems makes them particularly suitable for remote areas and scenarios not covered by the main grid, significantly improving energy efficiency and reducing dependence on the main grid.
[0005] Despite the numerous advantages of off-grid energy storage systems, capacity optimization becomes more complex due to the close interrelationships between various power generation, hydrogen production, energy storage, and hydrogen storage units in off-grid mode. Both excessively large and insufficient storage capacity can negatively impact the system's economics and performance. If the storage capacity is too small, it will be unable to effectively absorb renewable energy or meet the load demand of the electrolyzer, leading to wind and solar curtailment, increased load shedding penalties, and decreased hydrogen production. Conversely, excessive capacity will increase system investment and maintenance costs, thus affecting economics. Because multiple conflicting objectives need to be considered, the optimization process must employ multi-objective optimization methods to simultaneously consider investment, operating costs, wind and solar curtailment penalties, and maximizing output. Off-grid wind and solar power energy storage systems involve numerous units and more complex models. Traditional population-based algorithms typically employ a single search strategy, which is prone to getting trapped in local optima and struggles to quickly find the global optimum. Summary of the Invention
[0006] The purpose of this invention is to provide an adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind and solar hydrogen production and storage systems. By adaptively selecting appropriate search strategies for individuals at different stages of evolution, a set of Pareto optimal solutions can be found for users more quickly, so that users can formulate corresponding energy storage system capacity configuration schemes according to their preferences. This can reduce the randomness and blindness of the algorithm search and enhance the search efficiency of the algorithm.
[0007] To achieve the above objectives, this invention provides an adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen production and storage systems, comprising the following steps:
[0008] S1: Construct the objective function; based on the energy storage system model, minimize the unit hydrogen production cost, wind and solar curtailment penalties, and load shedding penalties;
[0009] S2: Construct the search space; use the capacity of each component of the wind and solar power generation unit, energy storage unit, and hydrogen and oxygen storage unit as decision variables, and set constraints that conform to the actual situation of the system;
[0010] S3: Initialize the population; based on the individuals in the initial population constructed in step S2, make them satisfy the constraints;
[0011] S4: Adaptive selection of search strategy; the policy adaptive network selects an appropriate search strategy based on the individual's decision variables and objective function value, and evolves a new subpopulation; the population iteration process includes constructing a policy pool, selection method, and policy adaptive network;
[0012] S5: Update the population and policy adaptation network; merge parent and child populations to calculate the non-dominated level and crowding of all individuals, determine the parameters of the reward update policy adaptation network based on the non-dominated level, and select a new population.
[0013] The energy storage system model of this invention includes a power generation unit, an energy storage unit, a hydrogen and oxygen storage unit, and an electrolyzer. The power generation unit consists of a wind turbine and a photovoltaic generator, converting wind and solar energy into electrical energy to power the electrolyzer for hydrogen and oxygen production, which is then stored in the hydrogen and oxygen storage unit. When there is a surplus of generated electricity, the system checks if the energy storage unit is full. If the energy storage unit has remaining capacity, it begins storing electricity using a battery and supercapacitor. If the energy storage unit is full, energy waste occurs, resulting in wind and solar power curtailment. The electrolyzer typically operates at its rated power, producing hydrogen and oxygen through water electrolysis. This hydrogen and oxygen are stored in hydrogen and oxygen storage tanks, respectively. If the combined discharge of the power generation unit and the energy storage unit cannot meet the rated power of the electrolyzer, the electrolyzer will operate at reduced load to maintain system stability. This results in a decrease in hydrogen and oxygen production, affecting the system's economic efficiency.
[0014] The objective function constructed in step S1 of this invention is as follows:
[0015]
[0016] In the formula: The annual unit cost of hydrogen production;
[0017] Annual investment cost;
[0018] Annual maintenance cost;
[0019] The mass of hydrogen per year;
[0020] Punishment for abandoning wind and light for the year;
[0021] Annual load reduction penalty;
[0022] In formula (1)
[0023]
[0024] In the formula: This refers to the depreciation rate;
[0025] This refers to the service life;
[0026] , , , , , , These refer to the capacities of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively.
[0027] , , , , , , These are the investment costs per unit capacity of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively.
[0028] In formula (1),
[0029] In the formula: , , , , , The annual maintenance costs per unit capacity of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks are respectively.
[0030] In formula (1),
[0031] In the formula: Hydrogen production rate per unit capacity per unit time per unit power of the electrolyzer;
[0032] In formula (1),
[0033] in
[0034] In the formula: This refers to the penalty coefficient for wind and solar power curtailment.
[0035] This refers to the power that is curtailed from wind and solar power.
[0036] Hourly wind power generation capacity;
[0037] Hourly photovoltaic power generation;
[0038] This refers to the rated power of the electrolytic cell;
[0039] This refers to the hourly discharge power of the battery.
[0040] This refers to the hourly discharge power of the supercapacitor.
[0041] In formula (1),
[0042]
[0043] In the formula: This is the load shedding penalty factor;
[0044] This represents the hourly operating power of the electrolytic cell.
[0045] In step S2 of this invention, the decision variables are the capacities of the wind turbine, photovoltaic generator, battery, supercapacitor, electrolyzer, hydrogen storage tank, and oxygen storage tank, and the power balance constraint is:
[0046]
[0047] In the formula: The charging power of the battery;
[0048] The charging power of the supercapacitor;
[0049] Range constraints: All variables mentioned above are non-negative and have upper and lower bound constraints.
[0050] The initial population in step S3 of this invention Depend on individual Composition, each individual have Variables ,individual Each variable The initialization method is as follows:
[0051]
[0052] In the formula: and This represents the upper and lower bounds of the variable;
[0053] A random number between 0 and 1 that follows a uniform distribution.
[0054] In step S4 of this invention, the strategy pool is constructed based on existing efficient search strategies to design a new gradient-guided search strategy. The strategy pool consists of formulas (11), (12), (13), (14), and (15), and the specific formulas are as follows:
[0055]
[0056] In the formula: , and The parent solution is randomly selected;
[0057] It is the global optimal solution in the current population, and in multi-objective optimization, it is the solution in the first non-dominated layer;
[0058] and These are predefined parameters;
[0059] These are random numbers sampled from [0,1] that have a uniform distribution;
[0060]
[0061] In the formula: These are uniformly distributed random values sampled in [0,1].
[0062] It is a predefined parameter;
[0063]
[0064]
[0065]
[0066] The policy adaptive network is a multi-layer fully connected network. Its input is the positional state of each individual in the decision space and the target space, and its output is the probability of choosing different search strategies.
[0067] The specific steps of step S5 of the present invention are as follows:
[0068] S51: After merging the parent and child populations, it is necessary to calculate the non-dominated level and crowding degree of all individuals, and update the new population based on this. In the capacity configuration optimization of energy storage systems, since multiple objectives are involved, the new population cannot be selected solely based on fitness values as in single-objective optimization. Therefore, it is necessary to calculate the dominance relationship between all solutions and sort the solutions by non-dominated level according to the dominance relationship, selecting the solutions with higher non-dominated levels to enter the new population. In order to ensure the diversity of the population, it is also necessary to evaluate the crowding degree of the solutions in the objective space, retain those solutions with relatively sparse distribution, and select some solutions from dense regions to join the new population.
[0069] The definition of dominance relation is as follows: Given two solutions and If on all targets, At least with Equally good, and on at least one objective, Superior Then the solution Dominant Solution Non-dominated level: Solutions that are not dominated by other solutions are called the first non-dominated level (level 1). After removing the first level, solutions that are not dominated by the remaining solutions are found, forming the second non-dominated level (level 2). The formula for calculating crowding is as follows:
[0070]
[0071] In the formula: and These are solutions In the target The objective value of the adjacent solution above;
[0072] and The goal The maximum and minimum values of all solutions;
[0073] S52: Determine the reward for different search strategies based on the dominance relationship, and update the parameters of the policy adaptive network. The main idea for updating the parameters is as follows: If a strategy is good, then the expected value of the state for all states S should be large. Therefore, the objective function of the policy adaptive network is:
[0074]
[0075] In the formula: The location of the decision space and the target space;
[0076] For a certain search strategy;
[0077] These are the parameters of the network;
[0078] The parameters in the policy adaptive network are: The individual is in a state Next, take action The probability of;
[0079] Indicates the state The following strategy The reward received;
[0080] The parameters of the adaptive network are updated by adjusting the strategy. , making The larger the value, the stronger the policy adaptive network becomes. Gradient ascent can be used to solve this maximization problem.
[0081] Let the parameters of the current policy adaptive network be... So, the new parameters Update via the following methods:
[0082]
[0083] In the formula: It is the learning rate;
[0084] For policy gradient;
[0085] Rewards are the key aspect of gradient updates. Taking Equation 8 as an example, the reward design after the search strategy is as follows:
[0086]
[0087] In the formula: Indicates a dominant relationship. express It outperforms at least one objective function ;
[0088] The relationship between the rewards is as follows: .
[0089] Compared with existing technologies, this invention, by adaptively selecting appropriate search strategies for individuals at different stages of evolution, can more quickly search for a set of Pareto optimal solutions for users, enabling them to formulate corresponding energy storage system capacity configuration schemes based on their preferences. The strategy pool designed in this invention introduces gradient information based on several classic differential evolution search strategies with different characteristics and simulated binary crossover operators. Using gradients to guide the algorithm's search reduces the randomness and blindness of the search to a certain extent, enhancing the algorithm's search efficiency. The policy adaptive network designed in this invention can determine whether an individual tends to converge or explore based on its different positions in the decision space and target space, reflecting this probabilistically in search strategies with different characteristics. This method solves the problem of search strategy uniformity, has good portability, and can be applied to other multimodal optimization problems. The policy adaptive network designed in this invention does not directly use the strategy with the highest probability for searching, but instead uses... Greedy strategies and roulette wheel methods increase the exploratory nature of the adaptive network in its early stages, while balancing exploration and utilization in the mid-to-late stages helps discover better strategies and reduces the risk of the network getting trapped in local optima. The gradient update of this policy network does not depend on other metrics; the reward for policy updates is determined solely by the Pareto dominance relation, making this adaptive network highly portable and applicable to other multi-objective evolutionary optimization methods. Attached Figure Description
[0090] Figure 1 This is an operational block diagram of the wind and solar power generation hydrogen production and storage system of the present invention;
[0091] Figure 2 This is the overall flowchart of the adaptive multi-objective evolutionary algorithm in this invention;
[0092] Figure 3 This is a flowchart of the population cyclic iteration of the adaptive multi-objective evolutionary algorithm in this invention;
[0093] Figure 4 This is a schematic diagram of the gradient-guided search strategy in this invention;
[0094] Figure 5 A comparison chart of the HV index between the evolution method of this invention and the classic NSGAII algorithm;
[0095] Figure 6 This is the Pareto optimal solution set searched in the embodiments of the present invention. Detailed Implementation
[0096] The invention will now be further described with reference to the accompanying drawings.
[0097] like Figures 1-4 As shown, an adaptive multi-objective evolutionary method for optimizing the capacity of an off-grid wind-solar hydrogen production and storage system is presented. The energy storage system model of this invention includes structures such as a power generation unit, an energy storage unit, a hydrogen and oxygen storage unit, and an electrolyzer. The power generation unit consists of a wind turbine and a photovoltaic generator, which convert wind and solar energy into electrical energy to supply the electrolyzer for hydrogen and oxygen production, which is then stored in the hydrogen and oxygen storage unit. When there is a surplus of generated electricity, the system checks whether the energy storage unit is full. If the energy storage unit still has capacity, it begins storing electricity using a storage unit composed of batteries and supercapacitors. If the energy storage unit is full, energy waste occurs, resulting in wind and solar power curtailment. The electrolyzer typically operates at its rated power, producing hydrogen and oxygen through water electrolysis. This hydrogen and oxygen are stored in hydrogen and oxygen storage tanks, respectively. If the combined discharge of the power generation unit and the energy storage unit cannot meet the rated power of the electrolyzer, the electrolyzer will operate at off-load to maintain system stability. This will cause a decrease in hydrogen and oxygen production, affecting the system's economic efficiency.
[0098] The adaptive multi-objective evolutionary algorithm of the present invention includes the following steps:
[0099] S1: Constructing the objective function; Based on the energy storage system model, the unit hydrogen production cost, wind and solar curtailment penalties, and load shedding penalties are minimized as the objectives. Since the hydrogen and oxygen produced by water electrolysis are proportional, this invention only considers the hydrogen cost. The formula for the objective function is as follows:
[0100] The objective function constructed in step S1 of this invention is as follows:
[0101]
[0102] In the formula: The annual unit cost of hydrogen production;
[0103] Annual investment cost;
[0104] Annual maintenance cost;
[0105] The mass of hydrogen per year;
[0106] Punishment for abandoning wind and light for the year;
[0107] Annual load reduction penalty;
[0108] In formula (1)
[0109]
[0110] In the formula: This refers to the depreciation rate;
[0111] This refers to the service life;
[0112] , , , , , , These refer to the capacities of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively.
[0113] , , , , , , These are the investment costs per unit capacity of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively.
[0114] In formula (1),
[0115] In the formula: , , , , , The annual maintenance cost per unit capacity of wind turbine generators, photovoltaic generators, batteries, electrolyzers, hydrogen storage tanks, and oxygen storage tanks are respectively.
[0116] In formula (1),
[0117] In the formula: Hydrogen production rate per unit capacity per unit time per unit power of the electrolyzer;
[0118] In formula (1),
[0119] in
[0120] In the formula: This refers to the penalty coefficient for wind and solar power curtailment.
[0121] This refers to the power that is curtailed from wind and solar power.
[0122] Hourly wind power generation capacity;
[0123] Hourly photovoltaic power generation;
[0124] This refers to the rated power of the electrolytic cell;
[0125] This refers to the hourly discharge power of the battery.
[0126] This refers to the hourly discharge power of the supercapacitor.
[0127] In formula (1),
[0128]
[0129] In the formula: This is the load shedding penalty factor;
[0130] This represents the hourly operating power of the electrolytic cell.
[0131] S2: Construct the search space; use the capacity of each component of the wind and solar power generation unit, energy storage unit, and hydrogen and oxygen storage unit as decision variables, and set constraints that conform to the actual situation of the system;
[0132] In step S2 of this invention, the decision variables are the capacities of the wind turbine, photovoltaic generator, battery, supercapacitor, electrolyzer, hydrogen storage tank, and oxygen storage tank, and the power balance constraint is:
[0133]
[0134] In the formula: The charging power of the battery;
[0135] The charging power of the supercapacitor;
[0136] Range constraints: All variables mentioned above are non-negative and have upper and lower bound constraints.
[0137] S3: Initialize the population; based on the individuals in the initial population constructed in step S2, make them satisfy the constraints;
[0138] Initial population Depend on individual Composition, each individual have Individual variables Each variable The initialization method is as follows:
[0139]
[0140] In the formula: and This represents the upper and lower bounds of the variable;
[0141] A random number between 0 and 1 that follows a uniform distribution.
[0142] S4: Adaptive selection of search strategy; the policy adaptive network selects an appropriate search strategy based on the individual's decision variables and objective function value, and evolves a new subpopulation; the population iteration process includes constructing a policy pool, selection method, and policy adaptive network;
[0143] In step S4 of this invention, the strategy pool is constructed based on existing efficient search strategies to design a new gradient-guided search strategy. The strategy pool consists of formulas (11), (12), (13), (14), and (15), and the specific formulas are as follows:
[0144]
[0145] In the formula: , and The parent solution is randomly selected;
[0146] It is the global optimal solution in the current population, and in multi-objective optimization, it is the solution in the first non-dominated layer;
[0147] and These are predefined parameters;
[0148] These are random numbers sampled from [0,1] that have a uniform distribution;
[0149]
[0150] In the formula: These are uniformly distributed random values sampled in [0,1].
[0151] It is a predefined parameter;
[0152]
[0153]
[0154]
[0155] The policy adaptive network described is a multi-layered, fully connected network. Its input consists of the positional states of each individual in the decision space and the target space, and its output is the probability of choosing different search strategies. Because it considers both the positional states in the decision space and the target space, the designed policy adaptive network is more effective in predicting policy probabilities. Figure 3 As shown, after the policy adaptive network calculates the probability of each policy, it will... The greedy and roulette wheel methods select a specific search strategy for execution, and the specific steps are as follows: This invention will use... The probability of randomly selecting a search strategy, in order to The probability is calculated based on the policy probability of the adaptive network, and the search strategy is selected using a roulette wheel method. In the early stages of network updates, this invention will use a slightly larger probability. The search strategy is randomly selected initially, and later it is selected with a higher probability according to the calculated strategy probability.
[0156] S5: Update the population and policy adaptation network; merge parent and child populations to calculate the non-dominated level and crowding of all individuals, determine the parameters of the reward update policy adaptation network based on the non-dominated level, and select a new population.
[0157] The specific steps of step S5 are as follows:
[0158] S51: After merging the parent and child populations, it is necessary to calculate the non-dominated level and crowding degree of all individuals, and update the new population based on this. In the capacity configuration optimization of energy storage systems, since multiple objectives are involved, the new population cannot be selected solely based on fitness values as in single-objective optimization. Therefore, it is necessary to calculate the dominance relationship between all solutions and sort the solutions by non-dominated level according to the dominance relationship, selecting the solutions with higher non-dominated levels to enter the new population. In order to ensure the diversity of the population, it is also necessary to evaluate the crowding degree of the solutions in the objective space, retain those solutions with relatively sparse distribution, and select some solutions from dense regions to join the new population.
[0159] The definition of dominance relation is as follows: Given two solutions and If on all targets, At least with Equally good, and on at least one objective, Superior Then the solution Dominant Solution Non-dominated level: Solutions that are not dominated by other solutions are called the first non-dominated level (level 1). After removing the first level, solutions that are not dominated by the remaining solutions are found, forming the second non-dominated level (level 2). The formula for calculating crowding is as follows:
[0160]
[0161] In the formula: and These are solutions In the target The objective value of the adjacent solution above;
[0162] and The goal The maximum and minimum values of all solutions;
[0163] S52: Determine the reward for different search strategies based on the dominance relationship, and update the parameters of the policy adaptive network. The main idea for updating the parameters is as follows: If a strategy is good, then the expected value of the state for all states S should be large. Therefore, the objective function of the policy adaptive network is:
[0164]
[0165] In the formula: The location of the decision space and the target space;
[0166] For a certain search strategy;
[0167] These are the parameters of the network;
[0168] The parameters in the policy adaptive network are: The individual is in a state Next, take action The probability of;
[0169] Indicates the state The following strategy The reward received;
[0170] The parameters of the adaptive network are updated by adjusting the strategy. , making The larger the value, the stronger the policy adaptive network becomes. Gradient ascent can be used to solve this maximization problem.
[0171] Let the parameters of the current policy adaptive network be... So, the new parameters Update via the following methods:
[0172]
[0173] In the formula: It is the learning rate;
[0174] For policy gradient;
[0175] Rewards are the key aspect of gradient updates. Taking Equation 8 as an example, the reward design after the search strategy is as follows:
[0176]
[0177] In the formula: Indicates a dominant relationship. express It outperforms at least one objective function ;
[0178] The relationship between the rewards is as follows: .
[0179] The policy pool of this invention designs a new gradient-guided search strategy based on existing efficient search strategies. The specific construction process is as follows: Exploration and convergence operations typically accompany the entire search process of the algorithm. However, as the population iterates, the algorithm should shift from tending to explore the entire search space to tending to converge to the optimal solution or a local optimum. Therefore, an excellent policy pool should contain various search strategies from exploration to convergence. Over the past two decades, differential evolution algorithms and their variants have repeatedly achieved excellent rankings in the IEEE CEC single-objective optimization competition. Below are some differential evolution search strategies with different characteristics:
[0180]
[0181]
[0182]
[0183] In the formula: It is a child's solution;
[0184] It is the dimension of the decision variables;
[0185] , and The parent solution is randomly selected;
[0186] It is the global optimal solution in the current population, and in multi-objective optimization, it is the solution in the first non-dominated layer;
[0187] and These are predefined parameters;
[0188] It is a random number with a uniform distribution sampled in [0,1].
[0189] Formula (9) exhibits strong randomness, which is beneficial for global search and is generally suitable for the initial stage of the search, increasing the diversity of the population. Formula (10) converges quickly but may lead to insufficient diversity, making it suitable for local convergence in the later stages. Formula (11) combines the advantages of global search and local convergence, making it suitable for complex multimodal optimization problems. Furthermore, simulated binary crossover is the most popular crossover operator for continuous optimization in genetic algorithms and decomposition-based multi-objective evolutionary algorithms.
[0190]
[0191] In the formula: These are uniformly distributed random values sampled from [0,1]. It is a predefined parameter.
[0192] Although the above formula is widely used, some invalid searches may still occur during the iteration process. For example... Figure 4 As shown in the left figure, with Two different sets of basic solutions and The possible distributions of the solutions generated by the difference operation are all in On the left side. However, as can be seen from the entire function, the potential region is... On the right side, this phenomenon is caused by an incorrect direction of the difference, which severely affects search efficiency, especially during local searches. To solve this problem, we introduce a gradient to guide the search direction, in the context of... Figure 4 Under the same conditions on the left, the solution generated by introducing the gradient will be distributed in the more promising region on the right. The modified formula is shown below.
[0193]
[0194]
[0195]
[0196] Formula (13) is derived from Formula (9). Formula (9) is too random and blind, and is only suitable for expanding the diversity of the population in the early stage of the search. Formula (13) reduces the randomness to a certain extent after introducing gradient information. Formula (14) is derived from Formula (10). Formula (10) uses difference search near the optimal solution and has a fast convergence speed, but it is easy to get trapped in local optima and is not suitable for multi-objective optimization. Therefore, this invention makes a slight modification so that the basic solution refers to the position of the optimal solution and its own gradient information, so that it is not easy to get trapped in local optima. Formula (15) is derived from Formula (12). While retaining its global search capability, it adds the ability to converge to the optimal solution. Formula (11) contains two difference vectors, which is more complicated and difficult to add gradient information. Moreover, it takes into account both global search and local convergence, so it is not changed.
[0197] In summary, the strategy pool of the present invention will consist of formulas (11), (12), (13), (14) and (15).
[0198] This invention, based on existing differential evolution operators and simulated binary crossover operators, designs a gradient-guided search strategy to reduce the randomness of the algorithm's search. Furthermore, this invention designs a policy adaptive network to solve the problem of single search strategy. This network can adaptively select appropriate search strategies according to different optimization problems and different individual states, effectively enhancing the algorithm's search efficiency and exhibiting good portability.
[0199] An embodiment of the present invention is given.
[0200] Using actual data on wind speed and solar irradiance in a specific region, the optimization objective is to minimize the annual unit cost of hydrogen production, the annual wind and solar curtailment penalty, and the annual load shedding penalty. The population size is set to 50, the number of iterations to 50, and the capacity of the electrolyzer is fixed. The capacity configurations of the wind turbine, photovoltaic generator, battery, supercapacitor, hydrogen storage tank, and oxygen storage tank are optimized. For example... Figure 5 As shown, the adaptive evolutionary method with gradient-guided search strategy pool proposed in this invention is compared with the classic NSGAII. The HV index is a commonly used indicator for analyzing the performance of multi-objective algorithms; the larger the HV index value, the better the convergence and distribution of the method used. From Figure 5 As can be seen, the method of this invention has excellent search capabilities, quickly converging to the vicinity of the optimal solution and becoming relatively stable in the later stages of iteration. In contrast, the HV metric of the NSGAII algorithm oscillates continuously. This is because a population size of 50 is relatively small for NSGAII; although it generates many non-dominated solutions, it fails to approach the optimal solution. This demonstrates that the method proposed in this invention has excellent search efficiency and convergence capabilities. Figure 6As shown, the Pareto optimal solution set searched by the method of this invention, if the optimal solution selected with a 1:1:1 preference for the three objectives, has the following minimum optimization results: annual unit hydrogen production cost of 29, annual wind and solar curtailment penalty of 205, and annual load shedding penalty of 252. The corresponding hydrogen cost is 29 yuan / kg, and the curtailment rate (curtailed wind and solar power divided by total power generation) and power shortage rate (load shedding power divided by total power generation) are both less than 1%. If users have different preferences and do not need to repeatedly solve this problem, they can directly calculate the optimal solution set by weighted summation and sorting according to their respective preferences to find the optimal solution under the corresponding preference.
[0201] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. An adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen energy storage system, characterized in that, Includes the following steps: S1: Construct the objective function; based on the energy storage system model, minimize the unit hydrogen production cost, wind and solar curtailment penalties, and load shedding penalties; S2: Construct the search space; use the capacity of each component of the wind and solar power generation unit, energy storage unit, and hydrogen and oxygen storage unit as decision variables, and set constraints that conform to the actual situation of the system; S3: Initialize the population; based on the individuals in the initial population constructed in step S2, make them satisfy the constraints; S4: Adaptive search strategy selection; the policy adaptive network selects an appropriate search strategy based on the individual's decision variables and objective function value, and evolves new subpopulations; The population iteration process includes building a policy pool, selecting a method, and a policy adaptation network; S5: Update the population and policy adaptation network; merge the parent and child populations to calculate the non-dominated level and crowding of all individuals, determine the parameters of the reward update policy adaptation network based on the non-dominated level, and select a new population; The policy pool consists of the following formula, the specific formula is as follows: wherein: , and are randomly selected parent solutions; is the global optimum solution in the current population; and are predefined parameters; is a random number with uniform distribution sampled in [0, 1]; wherein: is a uniformly distributed random value sampled in [0, 1]; It is a predefined parameter; The policy adaptive network is a multi-layer fully connected network. Its input is the positional state of each individual in the decision space and the target space, and its output is the probability of choosing different search strategies.
2. The adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen production and storage systems according to claim 1, characterized in that, The energy storage system model includes a power generation unit, an energy storage unit, a hydrogen and oxygen storage unit, and an electrolyzer. The power generation unit consists of a wind turbine and a photovoltaic generator. The power generation unit converts wind and solar energy into electrical energy to supply the electrolyzer for hydrogen and oxygen production, which is then stored in the hydrogen and oxygen storage unit.
3. The adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen production and storage systems according to claim 2, characterized in that, The objective function constructed in step S1 is as follows: In the formula: The annual unit cost of hydrogen production; Annual investment cost; Annual maintenance cost; The mass of hydrogen per year; Punishment for abandoning wind and light for the year; Annual load reduction penalty; in: In the formula: This refers to the depreciation rate; This refers to the service life; , , , , , , These refer to the capacities of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively. , , , , , , These are the investment costs per unit capacity of wind turbines, photovoltaic generators, batteries, supercapacitors, electrolyzers, hydrogen storage tanks, and oxygen storage tanks, respectively. in: In the formula: , , , , , The annual maintenance cost per unit capacity of wind turbine generators, photovoltaic generators, batteries, electrolyzers, hydrogen storage tanks, and oxygen storage tanks are respectively. in: In the formula: Hydrogen production rate per unit capacity per unit time per unit power of the electrolyzer; in, in In the formula: This refers to the penalty coefficient for wind and solar power curtailment. This refers to the power that is curtailed from wind and solar power. Hourly wind power generation capacity; Hourly photovoltaic power generation; This refers to the rated power of the electrolytic cell; This refers to the hourly discharge power of the battery. This refers to the hourly discharge power of the supercapacitor. in, In the formula: This is the load shedding penalty factor; This represents the hourly operating power of the electrolytic cell.
4. The adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen production and storage systems according to claim 2, characterized in that, In step S2, the decision variables are the capacities of the wind turbine, photovoltaic generator, battery, supercapacitor, electrolyzer, hydrogen storage tank, and oxygen storage tank, and the power balance constraint is: In the formula: The charging power of the battery; The charging power of the supercapacitor; Range constraints: All variables mentioned above are non-negative and have upper and lower bound constraints.
5. An adaptive multi-objective evolutionary method for optimizing the capacity of an off-grid wind-solar hydrogen production and storage system according to claim 2, characterized in that, Initial population in step S3 Depend on individual Composition, each individual have Individual variables Each variable The initialization method is as follows: In the formula: and This represents the upper and lower bounds of the variable; A random number between 0 and 1 that follows a uniform distribution.
6. The adaptive multi-objective evolutionary method for optimizing the capacity of off-grid wind-solar hydrogen production and storage systems according to claim 2, characterized in that, The specific steps of step S5 are as follows: S51: For multi-objective optimization of energy storage capacity, calculate and merge the non-dominated level and crowding degree of all individuals in the parent and child populations; select the best individuals according to the non-dominated level, combine the crowding degree to maintain population diversity, and screen out a new population. The formula for calculating congestion is as follows: In the formula: and These are solutions In the target The objective value of the adjacent solution above; and The goal The maximum and minimum values of all solutions; S52: Determine the reward for different search strategies based on the dominance relationship, and update the parameters of the strategy adaptive network. The objective function of the strategy adaptive network is: In the formula: The location of the decision space and the target space; For a certain search strategy; These are the parameters of the network; The parameters in the policy adaptive network are: The individual is in a state Next, take action The probability of; Indicates the state The following strategy The reward received; Solving the objective function causes the gradient to rise; Let the parameters of the current policy adaptive network be... So, the new parameters Update via the following methods: In the formula: It is the learning rate; For policy gradient; The reward design after the search strategy is executed is as follows: In the formula: Indicates a dominant relationship. express It outperforms at least one objective function ; The relationship between the rewards is as follows: .
Citation Information
Patent Citations
Global optimization method based on strategy adaptability differential evolution
CN105678401A
Reinforced multi-target firework algorithm for newly added urban energy emergency station
CN114219314A