A coordination control method, device and equipment based on a light hydrogen storage micro-grid system and a medium
By combining action prediction networks with reward functions, a multi-objective balanced control strategy is generated, which solves the problem of synergistic optimization of safety performance, operational performance and energy utilization in photovoltaic-hydrogen storage microgrid systems, and achieves stable and efficient operation of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FOSHAN XIANHU LAB
- Filing Date
- 2026-02-03
- Publication Date
- 2026-06-16
AI Technical Summary
Traditional photovoltaic-hydrogen storage microgrids lack a globally unified coordination and decision-making mechanism, which leads to an inability to balance system safety performance, operational performance and energy utilization, resulting in safety hazards and energy waste.
An action prediction network combined with a reward function is used for training to generate an optimal control strategy that balances multiple objectives. By extracting multi-dimensional features, global coordinated control of photovoltaic, energy storage, and hydrogen production systems is achieved, and the target control action with the highest balance between safety performance, operational performance, and energy utilization is output.
It has achieved stable, efficient and safe operation of the photovoltaic-hydrogen storage microgrid system throughout the entire operating cycle, avoiding system instability and energy waste caused by single-objective optimization in traditional control.
Smart Images

Figure CN122225550A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of photovoltaic-hydrogen storage microgrid control technology, and in particular to a coordinated control method, device, equipment and medium based on a photovoltaic-hydrogen storage microgrid system. Background Technology
[0002] Coordinated control of a photovoltaic-hydrogen storage microgrid generally refers to the overall regulation and control of the operating status of the photovoltaic power generation system, energy storage system, and hydrogen production system within the microgrid, enabling each unit to respond collaboratively to the system's operational needs. In a photovoltaic-hydrogen storage microgrid system, photovoltaic power generation is intermittent and fluctuates, the energy storage system is subject to state of charge constraints, and the hydrogen production system is limited by hydrogen storage pressure state constraints. The operating states of these three systems are interconnected and mutually influential, requiring coordinated control to achieve stable overall system operation.
[0003] In traditional technologies, the coordinated control of photovoltaic-hydrogen storage microgrids typically employs independent control of individual modules. This means that the photovoltaic, energy storage, and hydrogen production systems each have their own independent control modules, formulating control strategies based on their own local operational data, lacking a globally unified coordination and decision-making mechanism. Furthermore, traditional control methods focus on single-objective optimization, failing to balance system safety, operation, and energy utilization. For example, prioritizing energy efficiency can lead to safety hazards such as overcharging of energy storage and DC bus voltage exceeding limits, shortening equipment lifespan. Conversely, prioritizing system safety and stability can excessively limit photovoltaic power output or reduce hydrogen production, resulting in low energy utilization and wasting renewable energy. Therefore, traditional technologies lack a multi-objective balance assessment mechanism between safety performance, operational performance, and energy utilization, failing to guarantee the stable, safe, and efficient operation of the entire system. Summary of the Invention
[0004] The main objective of this application is to propose a coordinated control method, device, equipment, and medium based on a photovoltaic-hydrogen storage microgrid system. Using a reward function as a guided optimal action selection mechanism, the action prediction network can iteratively converge under the constraint of action labels that maximize the balance between safety, operation, and energy utilization throughout the entire operation cycle. In the application phase, by extracting multi-dimensional features from the current operating data, it accurately outputs the target control actions that maximize the balance between safety performance, operational performance, and energy utilization, ensuring that the photovoltaic-hydrogen storage microgrid operates stably and efficiently in an optimal balance state throughout its entire operating cycle.
[0005] To achieve the above objectives, one aspect of this application proposes a coordinated control method based on a photovoltaic-hydrogen storage microgrid system, the method comprising:
[0006] Acquire the current operating data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operating data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; The current operating data is input into a preset action prediction network, which extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics from the current operating data. Then, based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics, the network outputs the target control action corresponding to the current operating data. The target control action includes target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system, respectively. Based on the target control action, the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system are controlled respectively; The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action; the candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample; wherein, the reward function is used to evaluate the balance between the safety performance, operational performance and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action.
[0007] Furthermore, in some embodiments, the power characteristics include: photovoltaic power characteristics, energy storage power characteristics, and hydrogen production system power characteristics; the current characteristics include: photovoltaic current characteristics, energy storage current characteristics, and hydrogen production current characteristics; the target operating state control commands include: operating mode control commands, power control commands, and working state control commands. The step of outputting the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics includes: Based on the characteristics of bus voltage, photovoltaic power, and photovoltaic current, photovoltaic operating condition characteristics are generated. Based on the characteristics of bus voltage, energy storage power, energy storage current, and state of charge, the operating condition characteristics of energy storage are generated. Based on the characteristics of bus voltage, power of hydrogen production system, hydrogen production current, and hydrogen energy storage, the operating conditions of hydrogen production are generated. The photovoltaic operating condition characteristics, the energy storage operating condition characteristics, and the hydrogen production operating condition characteristics are fused together to output system operating condition characteristics that characterize the global operating condition of the photovoltaic-hydrogen storage microgrid system. The system operating condition characteristics, photovoltaic operating condition characteristics, energy storage operating condition characteristics, and hydrogen production operating condition characteristics are weighted and fused, and operating mode control commands for photovoltaic power generation system, power control commands for energy storage system, and working state control commands for hydrogen production system are generated under DC voltage constraints, energy storage state of charge constraints, power balance constraints, and equipment action constraints. The operating mode control command, the power control command, and the working state control command are output as target control actions.
[0008] Furthermore, in some embodiments, it also includes: When the target control action is detected to have been executed, the response operation data of the photovoltaic hydrogen storage microgrid system in the next time period is acquired; Substitute the response operation data into the reward function to calculate the target reward function value corresponding to the target control action; The current operating data, target control actions, operating data of the photovoltaic-hydrogen storage microgrid system for the next time period, and objective function values are used as the target action interaction trajectory and added to the strategy experience pool corresponding to the action prediction network; wherein, the strategy experience pool is used to update the network parameters of the action prediction network.
[0009] Furthermore, in some embodiments, it also includes: When the number of action interaction trajectories in the strategy experience pool exceeds a preset trajectory number threshold, several action interaction trajectories in the strategy experience pool are used as a trajectory set; wherein, the trajectory set includes: several first action interaction trajectories; the first action interaction trajectory includes: first running data of a first time period, a first control action corresponding to the first running data, a first reward function value corresponding to the first control action, and first response running data corresponding to the next time period corresponding to the first time period; For each first action interaction trajectory, the expected benefit value of the first operating data and the expected benefit value of the first response operating data are determined through the evaluation network corresponding to the action prediction network; wherein, the evaluation network is used to evaluate the role of the operating data in a time period on the operating benefit of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, the first reward function value corresponding to the first control action, the expected reward value of the first running data, and the expected reward value corresponding to the first response running data are substituted into the action advantage function to determine the action advantage function value of the first control action; wherein, the action advantage function is used to evaluate the role of the first control action in the operational benefits of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, a single-sample pruning strategy objective term is calculated based on the selection probability ratio corresponding to the first control action, the action dominance function value, and a preset probability ratio pruning interval. This objective term characterizes the contribution of the first action interaction trajectory to the optimization of the action prediction network. The single-sample pruning strategy objective term indicates the degree of optimization contribution of the first action interaction trajectory when the network parameters of the action prediction network are updated. The selection probability ratio characterizes the ratio of the probability of selecting the first control action under the current network parameters of the action prediction network, given the same first running data, to the probability of selecting the first control action when the network parameters of the action prediction network were not updated previously. Based on the single-sample pruning strategy objective item corresponding to each first action interaction trajectory and the total number of action interaction trajectories in the trajectory set, the mean value of the strategy optimization objective function corresponding to the trajectory set is calculated. The network parameters of the current action prediction network are updated with the goal of maximizing the mean of the objective function of the strategy.
[0010] Furthermore, in some embodiments, the target term of the single-sample pruning strategy is calculated according to the following formula: ; in, Indicates in First control action at all times The target item of the single-sample pruning strategy, Indicates in First control action at all times The ratio of the probability of selection, Indicates in First control action at all times The action advantage function value, This represents the clipping function. Indicates will The value of is restricted to The value is retrieved from within. For the preset selection probability range, This represents the preset hyperparameters.
[0011] Furthermore, in some embodiments, the process of constructing the reward function includes: Based on the real-time voltage data and rated voltage data of the DC bus, a voltage stability bonus data item is constructed; Based on the state of charge data of energy storage batteries and multiple different charge warning level thresholds, a reward data item for safe operation of charged batteries is constructed. Based on the pressure status data of the hydrogen storage tank and multiple different pressure warning level thresholds, construct reward data items for the safe operation of the hydrogen storage system. Based on the switching data of photovoltaic power generation system operation mode, the switching data of energy storage system power operation level, and the switching data of hydrogen production system working status, an action switching penalty data item is constructed. Based on the operating power data of the hydrogen production system, the operating power data of the photovoltaic power generation system, the operating power data of the energy storage system, the upper limit power of the photovoltaic system, the photovoltaic efficiency, the converter efficiency, and the electrolyzer efficiency of the hydrogen production system, a system operating efficiency bonus data item is constructed. Based on the standard deviation data of DC voltage, the standard deviation data of hydrogen production power, and the rated value of hydrogen production power, a system operation stability bonus data item is constructed. Based on the current data of the photovoltaic system, the rated current of the photovoltaic system, the current data of the electrolyzer of the hydrogen production system, the rated current of the electrolyzer, the hydrogen production power, and the rated hydrogen production power, construct equipment life protection incentive data items; Based on the preset hydrogen production reward weighting coefficient, actual hydrogen production data, theoretical hydrogen production data, hydrogen production power, and rated hydrogen production power, a hydrogen production reward data item is constructed. Based on the voltage stability reward data item, the first weight value corresponding to the voltage stability reward data item, the charged safe operation reward data item, the second weight value corresponding to the charged safe operation reward data item, the hydrogen storage system safe operation reward data item, and the third weight value corresponding to the hydrogen storage system safe operation reward data item, a first data item for evaluating the safety performance of the photovoltaic hydrogen storage microgrid system is constructed. Based on the action switching penalty data item, the fourth weight value corresponding to the action switching penalty data item, the system operation efficiency reward data item, the fifth weight value corresponding to the system operation efficiency reward data item, the system operation stability reward data item, and the sixth weight value corresponding to the system operation stability reward data item, a second data item is constructed to evaluate the operation performance of the photovoltaic hydrogen storage microgrid system. Based on the equipment life protection reward data item, the seventh weight value corresponding to the equipment life protection reward data item, the hydrogen production reward data item, and the eighth weight value corresponding to the hydrogen production reward data item, a third data item is constructed to evaluate the energy utilization rate of the photovoltaic-hydrogen storage microgrid system. The reward function is constructed based on the first data item, the second data item, and the third data item.
[0012] Furthermore, in some embodiments, the reward function includes: ; ; ; ; ; ; ; ; ; ; in, This represents the reward function value corresponding to the reward function expression. The data item is awarded as a reward for voltage stability. This is a data item for rewarding safe operation of charged batteries. Data items awarded for the safe operation of hydrogen storage systems. To switch the penalty data item for the action, Data items are awarded to improve system operating efficiency. Data items awarded for system operational stability This is a data item for equipment lifespan protection rewards. This is a data item for hydrogen production rewards. As the first weight value, This is the second weight value. As the third weight value, It is the fourth weight value. It is the fifth weight value. It is the sixth weight value. It is the seventh weight value. It is the eighth weight value. This is the real-time voltage of the DC bus. This is the rated voltage of the DC bus. This refers to the state of charge of the energy storage battery. The pressure status of the hydrogen storage tank; For the first The penalty value for each switching action. For indicator functions, when hour, A value of 1 indicates that a switching action occurred between the current time t and the previous time. The corresponding penalty value is valid; when hour, A value of 0 indicates that there was no switching action between the current time t and the previous time. The corresponding penalty value is invalid; , , ; The overall operating efficiency of the photovoltaic-hydrogen storage microgrid system, For hydrogen production capacity, This represents the upper limit of the photovoltaic system's power. This represents the actual power of the photovoltaic system. For the power of the storage battery, For photovoltaic efficiency, For converter efficiency, The efficiency of the electrolyzer in the hydrogen production system is given by the coefficient 0.95, which represents the system's loss factor. The standard deviation of DC voltage, The standard deviation of hydrogen production power. This is the rated value for hydrogen production capacity. For the current of the photovoltaic system, This refers to the rated current of the photovoltaic system. This refers to the current in the electrolyzer of the hydrogen production system. This is the rated current of the electrolytic cell. This represents the actual production of hydrogen. This represents the theoretical hydrogen production.
[0013] To achieve the above objectives, another aspect of this application proposes a coordinated control device based on a photovoltaic-hydrogen storage microgrid system, the device comprising: The operation data acquisition module is used to acquire the current operation data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operation data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; The control action generation module is used to input the current operating data into a preset action prediction network, so that the action prediction network extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics based on the current operating data, and then outputs the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics; wherein, the target control action includes: target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system respectively; The system control module is used to control the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system according to the target control action. The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action; the candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample; wherein, the reward function is used to evaluate the balance between the safety performance, operational performance and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action.
[0014] Another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned coordinated control method for a photovoltaic-hydrogen storage microgrid system.
[0015] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned coordinated control method for a photovoltaic-hydrogen storage microgrid system.
[0016] The embodiments of this application include at least the following beneficial effects: This application provides a coordinated control method, device, equipment, and medium based on a photovoltaic-hydrogen storage microgrid system. Through the constructed action prediction network training and application logic, this invention can achieve optimal coordinated control of the safety performance, operational performance, and energy utilization balance of the photovoltaic-hydrogen storage microgrid system. Specifically, in the application process, the current time period's operational data of the photovoltaic-hydrogen storage microgrid is first collected, such as global operational data like power, current, DC bus voltage, energy storage state of charge, and hydrogen storage pressure state. Using a preset action prediction network as the global coordinated decision-making mechanism, replacing the dispersed control modules of each unit, the global operational data can be used as input to uniformly extract multi-dimensional features such as photovoltaic power, energy storage charge, hydrogen production and storage, and bus voltage. This achieves global perception and feature fusion of the overall system operational status and synchronously outputs target operational status control commands corresponding to the three major units: photovoltaic, energy storage, and hydrogen production, instead of each unit generating its own control commands. This solves the problem of traditional control lacking a global coordinated mechanism. Furthermore, during the training process of the action prediction network, this invention employs a reward function as a guided optimal action selection mechanism. Specifically, for each set of historical operational data samples, the action with the largest reward function value is selected from multiple candidate control actions as the actual control action. This ensures that each control action used for training is the optimal choice in terms of safety performance, operational performance, and energy utilization efficiency after multi-objective comprehensive evaluation, rather than an action that only satisfies a single objective. Since the action with the largest reward function value is used as the core input label for model training, it ensures that the action prediction network can complete iterative convergence under the constraint of the action label with the highest balance between safety, operation, and energy utilization throughout the entire process. This allows the action prediction network to learn and solidify the optimal control action under different operating conditions. Thus, in the practical application stage, the action prediction network can quickly extract multi-dimensional features of the current operational data and accurately output the target control action that best satisfies the balance between safety, operation, and energy utilization efficiency. This solves the problem that traditional technologies cannot achieve a multi-objective balance evaluation mechanism between safety performance, operational performance, and energy utilization efficiency. Through this invention, the photovoltaic-hydrogen storage microgrid can be guaranteed to operate stably and efficiently in an optimal balance state throughout its entire operating cycle. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a coordinated control method for a photovoltaic-hydrogen storage microgrid system provided in an embodiment of this application. Figure 2 This is a schematic diagram of the architecture of the photovoltaic hydrogen storage microgrid system provided in the embodiments of this application; Figure 3 This is a closed-loop control flowchart of the intelligent coordination controller provided in the embodiments of this application; Figure 4This is a flowchart of the training process of the DRL controller provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of a coordinated control device based on a photovoltaic hydrogen storage microgrid system provided in an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0019] In the field of coordinated control of photovoltaic-hydrogen storage microgrids, the focus is typically on regulating the operational status of the photovoltaic power generation system, energy storage system, and hydrogen production system. However, traditional technologies generally employ a modular, independent control model, with each of the three core systems—photovoltaic power generation, energy storage, and hydrogen production—equipped with its own independent control module, lacking a globally unified coordinating decision-making body. Each control module formulates its control strategy based solely on its own local operational data, without incorporating the overall system operational status and constraints from other units. This results in independent generation of control actions by each unit, leading to a lack of coordination and conflicting control actions. Furthermore, it fails to achieve coordinated responses from the three units to the system's operational needs, thus compromising the overall operational stability of the system.
[0020] Traditional control methods focus on single-objective optimization, prioritizing only one performance metric and lacking a comprehensive consideration and balance of system safety, operational performance, and energy efficiency. They fail to establish a multi-objective balance assessment system for these three aspects and lack effective means to quantify their balance, thus hindering the synergistic optimization of these three core performance metrics. Due to limitations of independent control architecture and a single optimization objective, traditional control strategies are often fixed designs, failing to fully consider the intermittency and volatility of photovoltaic power generation, as well as the dynamic changes in the state of charge of energy storage and the pressure state of hydrogen energy storage. When system operating conditions change, the individual control modules struggle to respond quickly and adjust their control strategies, failing to adapt to complex and ever-changing operating scenarios. This further exacerbates system instability and leads to significant fluctuations in energy efficiency, making it impossible to maintain consistently high-efficiency operation.
[0021] In view of this, this invention provides a coordinated control method, device, equipment, and medium based on a photovoltaic-hydrogen storage microgrid system. This invention uses a pre-set action prediction network as the globally unique coordinated decision-making entity, replacing the dispersed and independent control modules of each unit, to achieve globally unified decision-making and command output. Furthermore, this invention establishes a multi-objective balance evaluation and optimal action selection mechanism based on a reward function that can balance safety performance, operational performance, and energy utilization. Through iterative training of the model constrained by the reward function, the action prediction network learns and solidifies the optimal control strategy for multi-objective balance, achieving optimal coordinated control of system balance under all operating conditions.
[0022] Figure 1 This is an optional flowchart of a coordinated control method for a photovoltaic-hydrogen storage microgrid system provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S1 to S3: Step S1: Obtain the current operating data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operating data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; Step S2: Input the current operating data into a preset action prediction network, so that the action prediction network extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics based on the current operating data, and then outputs the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics; wherein, the target control action includes: target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system respectively; The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action. The candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample. The reward function is used to evaluate the balance between the safety performance, operational performance, and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action. Step S3: Control the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system according to the target control action.
[0023] As shown in step S2 of this embodiment, a multi-objective balanced optimal control action screening criterion is constructed through the reward function. During the training process of the action prediction network, for each set of historical operating data samples, all candidate control actions are traversed. The comprehensive balance of system safety performance, operating performance and energy utilization after the execution of each candidate action is quantitatively evaluated through the reward function. The candidate control action with the largest reward function value is selected as the actual control action corresponding to the sample. This ensures that the training labels input to the network model are all control actions with the optimal balance after multi-dimensional comprehensive evaluation. This allows the action prediction network to complete iterative training under the constraint of the action label with the optimal balance, thereby possessing the ability to deeply solidify the optimal control strategy for multiple operating conditions.
[0024] Since the training process of the action prediction network takes the actual control action with the largest reward function value as the output target, through continuous iterative learning, the action prediction network can gradually master the control logic that achieves the optimal balance of the three core performances of the system under different operating conditions, and solidify these optimal control strategies into the network model. Then, the finally converged action prediction network can autonomously identify the control action requirements under different operating scenarios without manual intervention to adjust and optimize the target, and form a multi-objective balance control capability that adapts to all operating conditions.
[0025] Furthermore, as illustrated in steps S1 to S3 of this embodiment, in practical applications, the action prediction network can accurately output the target control action with optimal balance, ensuring optimal operation of the system throughout its entire lifecycle. Specifically, firstly, global operating data of the photovoltaic-hydrogen storage microgrid for the current time period is collected and input into the pre-trained action prediction network. The network can then quickly extract multi-dimensional features such as power, current, bus voltage, energy storage state of charge, and hydrogen production and storage saturation. Based on the optimal control strategy solidified during the training phase, it directly outputs the corresponding target control action. This target control action can maximize the balance between system safety performance, operating performance, and energy utilization, thus solving the shortcomings of traditional technologies that focus on single-objective optimization and lack multi-objective balance evaluation. This achieves stable, safe, and efficient collaborative control of the photovoltaic-hydrogen storage microgrid system throughout its entire operating cycle.
[0026] Therefore, this invention, through the optimal action selection and model training mechanism guided by the reward function, enables the action prediction network to learn the optimal balancing strategy during the training phase and output the optimal balancing action during the application phase, thereby improving the overall operation quality and reliability of the photovoltaic-hydrogen storage microgrid system and ensuring the stable, safe and efficient operation of the entire system.
[0027] For step S1, in some embodiments, when acquiring the current operating data of the photovoltaic-hydrogen storage microgrid system in the current time period, the current power data includes the real-time power of the photovoltaic system, the real-time power of the energy storage system, and the real-time hydrogen production power; the current data includes the real-time current of the photovoltaic system, the real-time current of the energy storage system, and the real-time current of the hydrogen production system.
[0028] It is understood that this invention can collect operating parameters such as DC bus voltage, hydrogen energy storage pressure status, state of charge, current of each unit, and power data of each unit in real time, thereby constructing the state information of the photovoltaic-hydrogen storage microgrid system and providing multi-dimensional state information for the subsequent action prediction network's action decision-making.
[0029] For step S2, in a preferred embodiment, the power characteristics include: photovoltaic power characteristics, energy storage power characteristics, and hydrogen production system power characteristics; the current characteristics include: photovoltaic current characteristics, energy storage current characteristics, and hydrogen production current characteristics; the target operating state control command includes: operating mode control command, power control command, and working state control command. The step of outputting the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics includes: Based on the characteristics of bus voltage, photovoltaic power, and photovoltaic current, photovoltaic operating condition characteristics are generated. Based on the characteristics of bus voltage, energy storage power, energy storage current, and state of charge, the operating condition characteristics of energy storage are generated. Based on the characteristics of bus voltage, power of hydrogen production system, hydrogen production current, and hydrogen energy storage, the operating conditions of hydrogen production are generated. The photovoltaic operating condition characteristics, the energy storage operating condition characteristics, and the hydrogen production operating condition characteristics are fused together to output system operating condition characteristics that characterize the global operating condition of the photovoltaic-hydrogen storage microgrid system. The system operating condition characteristics, photovoltaic operating condition characteristics, energy storage operating condition characteristics, and hydrogen production operating condition characteristics are weighted and fused, and operating mode control commands for photovoltaic power generation system, power control commands for energy storage system, and working state control commands for hydrogen production system are generated under DC voltage constraints, energy storage state of charge constraints, power balance constraints, and equipment action constraints. The operating mode control command, the power control command, and the working state control command are output as target control actions.
[0030] In illustrative terms, in the application of motion prediction networks, this embodiment of the invention decomposes the operating data into power and current characteristics corresponding to three major units: photovoltaic, energy storage, and hydrogen production. Combined with bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics, it achieves perception of the operating status of each unit and targeted feature capture.
[0031] To address the differences in operating characteristics among the three major units—photovoltaics, energy storage, and hydrogen production—corresponding dimensional features are integrated to generate exclusive operating condition features. For example, the photovoltaic operating condition integrates voltage, power, and current features, while the energy storage operating condition superimposes state of charge features, thus achieving accurate feature identification of the operating scenarios of each unit.
[0032] By integrating the operating characteristics of photovoltaics, energy storage, and hydrogen production, global operating characteristics that characterize the overall operating status of the system are generated. This enables the action prediction network to perceive the energy supply and demand balance, equipment coordination relationships, and potential conflicts from a global system perspective. This ensures that the output control commands not only meet the operating needs of individual units but also satisfy the overall system coordination optimization goals, thus avoiding conflicts in the control actions of different units.
[0033] In this embodiment of the invention, DC voltage constraints, energy storage state of charge constraints, power balance constraints, and equipment action constraints are also incorporated into the instruction generation process. This avoids situations where the pursuit of multiple objectives leads to the breaking of equipment physical limits or system safety boundaries, ensuring that the generated control instructions are engineering feasible.
[0034] Therefore, the embodiments of the present invention generate multiple different control commands by weighted fusion of global operating condition characteristics and operating condition characteristics of each unit, combined with multiple constraints. This enables the output target control action to accurately match the balance requirements between safety performance, operational performance and energy utilization. Without manual intervention to adjust and optimize the target, it can autonomously achieve dynamic trade-offs between multiple targets, thus solving the drawbacks of traditional single-target optimization.
[0035] In some embodiments, the coordination control method of the present invention is applied to, for example... Figure 2 In the illustrated photovoltaic-hydrogen storage microgrid system architecture, the photovoltaic-hydrogen storage microgrid system is the physical control object of this invention, comprising photovoltaic power generation units, an energy storage battery system, a water electrolysis hydrogen production system, and a shared DC bus. Indicatively, each unit is connected to the DC bus via an independent DC / DC converter to achieve power coupling and energy interaction. Specifically, the photovoltaic power generation unit includes a photovoltaic array and a photovoltaic converter; the energy storage system includes a lithium-ion battery and an energy storage converter; the water electrolysis hydrogen production system includes an electrolyzer, a hydrogen production power source, and a hydrogen storage tank; and the shared DC bus is connected to the same bus via independent DC / DC converters to achieve power coupling.
[0036] like Figure 2 As shown, the intelligent coordination controller is the decision-making unit and action control unit of this invention, and is composed of a state observer and a DRL controller.
[0037] Here, DRL controller refers to Deep Reinforcement Learning Controller. The state observer is used to collect real-time operating parameters such as DC bus voltage (UDC), photovoltaic power, state of charge (SBC), and current of each unit, forming a system state vector to provide information for the DRL controller's decision-making. The DRL controller receives the system state vector output by the state observer and outputs an action vector. This action vector contains three dimensions of discrete control commands: operating mode control commands for the photovoltaic converter (0 for MPPT, 1 for CVC mode); power control commands for the hydrogen production power source (0 for low power, 1 for medium power, 2 for high power); and operating state control commands for the energy storage converter (0 for standby, 1 for charging, 2 for discharging).
[0038] Indicatively, this invention does not require the establishment of precise mathematical models for photovoltaic-hydrogen storage microgrids, such as nonlinear equations for photovoltaic power output fluctuations and time-varying characteristic models for hydrogen production loads. Through continuous interaction between the intelligent coordinating controller and the system, the system autonomously learns control laws. Compared with traditional control methods based on static optimization models, this invention can effectively cope with uncertain operating conditions such as random fluctuations in photovoltaic power output and sudden changes in hydrogen production loads, and avoid control failures caused by deviations between model assumptions and reality.
[0039] For step S3, the operating mode control command, the power control command, and the working state control command output in step S2 are directly converted into executable operations for each unit, without the need for additional complex intermediate conversion steps. For example, the mode switching of the photovoltaic converter can directly respond to 0 / 1 commands through the hardware interface, and the energy storage power control can directly track the power command through dual closed-loop regulation. Illustratively, when the target control action output in step S2 is: Photovoltaic power generation system operation mode control command: 1, indicating CVC (Constant Voltage Control mode); Energy storage system power control command: 0, indicating standby; Hydrogen production system operating status control command: 0, indicating low power level operation. The aforementioned target control actions can be used to control the photovoltaic converter to maintain CVC mode operation, providing voltage support for the energy storage discharge and hydrogen production system, thus preventing system shutdown; control the energy storage converter to enter standby mode, avoiding battery capacity degradation caused by unrestrained discharge; and, based on low-power level control commands, control the hydrogen production power supply to reduce the electrolyzer power to 20kW for low-power operation. Therefore, through coordinated matching of control actions of various units, such as stable voltage supply, energy storage standby, and low-power hydrogen production, frequent equipment start-ups and shutdowns or mode switching can be avoided, reducing mechanical and electrical stress and ensuring the continuous and reliable operation of the photovoltaic-hydrogen storage microgrid in low-energy scenarios.
[0040] In some embodiments, the coordination control method of the present invention uses the execution flow of an intelligent coordination controller as the execution carrier, such as... Figure 3 The closed-loop control flowchart of the intelligent coordination controller shown is illustrated below. The action decision-making process and feedback update process of the intelligent coordination controller in this embodiment of the application are as follows: State awareness is performed by a state observer, and system operating parameters are collected to construct a state vector, providing a data foundation for subsequent decision-making. Upon entering the decision-making phase, the DRL controller receives the state vector S. t Under multiple safety boundaries including voltage constraints, energy storage constraints, power balance constraints, and action constraints, the optimal action is calculated through an algorithm. Enter the action execution step and select the optimal action a. t The commands are converted into specific control instructions and sent to the converters in photovoltaic, energy storage, and hydrogen production systems to drive the equipment to adjust its operating status. After executing the instruction, the system enters a new state. Simultaneously, through the system feedback loop, the system's operational data, i.e., the operational state, is used to calculate the immediate reward r based on the multi-objective reward function. t To quantify the quality of actions.
[0041] Finally, the strategy is updated by storing experience data such as status, actions, and rewards into the experience pool. Once a certain amount is accumulated, the DRL controller parameters are optimized to achieve iterative strategy upgrades and form a closed loop.
[0042] In the aforementioned closed-loop process, the system's operating parameters across all dimensions can be collected by the state observer. The DRL controller calculates the optimal action under multiple safety constraints and transforms it into directly executable control commands. It combines a multi-objective reward function to quantify the merits of actions and optimizes controller parameters by accumulating data through an experience pool. Simultaneously, it employs a reward function-guided optimal action selection and action prediction network training mechanism. This ensures that the control strategy always remains within the safe and feasible domain, avoiding equipment damage and system instability. It also achieves multi-objective dynamic balance, avoiding the one-sidedness of traditional single-objective optimization. Furthermore, it allows the control strategy to autonomously iterate and upgrade to adapt to complex operating conditions, reducing frequent equipment switching and wear and tear, and significantly improving the accuracy, robustness, stability, and overall performance of the system.
[0043] In a preferred embodiment, the action prediction network of the present invention also corresponds to a policy experience pool. The policy experience pool can be used to optimize the parameters of the action prediction network and control the action selection strategy. The basic unit stored in the pool is the action interaction trajectory. Each trajectory completely records a set of closed-loop interaction information from system state perception, control action execution to environmental response feedback. Specifically, it includes: system operation data of the current time period, control action output based on the current state, system response operation data of the next period after the action is executed, and target reward function value for quantitative evaluation of the effect of this action execution.
[0044] Indicatively, the policy experience pool continuously receives new trajectory data throughout the training process of the action prediction network, and also continues to receive new trajectory data during the application phase of the action prediction network. Indicatively, when the number of trajectories in the policy experience pool reaches a preset threshold, a batch of trajectory sets can be extracted. The evaluation network calculates the value of two adjacent states and the action advantage function value of each control action, thereby constructing a single-sample pruning policy objective term to obtain the mean of the policy optimization objective function. This allows for the synchronous updating of the parameters of both the action prediction network and the evaluation network. Therefore: When the target control action is detected to have been executed, the response operation data of the photovoltaic hydrogen storage microgrid system in the next time period is acquired; Substitute the response operation data into the reward function to calculate the target reward function value corresponding to the target control action; The current operating data, target control actions, operating data of the photovoltaic-hydrogen storage microgrid system for the next time period, and objective function values are used as the target action interaction trajectory and added to the strategy experience pool corresponding to the action prediction network; wherein, the strategy experience pool is used to update the network parameters of the action prediction network.
[0045] It is understood that, in the embodiments of the present invention, the training and updating of the action prediction network can be based on a large number of diverse system operating condition samples to determine which actions are optimal in which states, which actions will cause the system to deviate from the safe operating range, and how to achieve a balance among multiple objectives. By accumulating sufficient trajectories in the policy experience pool, the action prediction network can learn the control action selection rules under different operating conditions.
[0046] In a preferred embodiment, the parameter update strategy of the action prediction network of the present invention is derived from batch sampling of action interaction trajectories in the experience pool and calculation of the mean of the strategy optimization objective function, then: When the number of action interaction trajectories in the strategy experience pool exceeds a preset trajectory number threshold, several action interaction trajectories in the strategy experience pool are used as a trajectory set; wherein, the trajectory set includes: several first action interaction trajectories; the first action interaction trajectory includes: first running data of a first time period, a first control action corresponding to the first running data, a first reward function value corresponding to the first control action, and first response running data corresponding to the next time period corresponding to the first time period; For each first action interaction trajectory, the expected benefit value of the first operating data and the expected benefit value of the first response operating data are determined through the evaluation network corresponding to the action prediction network; wherein, the evaluation network is used to evaluate the role of the operating data in a time period on the operating benefit of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, the first reward function value corresponding to the first control action, the expected reward value of the first running data, and the expected reward value corresponding to the first response running data are substituted into the action advantage function to determine the action advantage function value of the first control action; wherein, the action advantage function is used to evaluate the role of the first control action in the operational benefits of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, a single-sample pruning strategy objective term is calculated based on the selection probability ratio corresponding to the first control action, the action dominance function value, and a preset probability ratio pruning interval. This objective term characterizes the contribution of the first action interaction trajectory to the optimization of the action prediction network. The single-sample pruning strategy objective term indicates the degree of optimization contribution of the first action interaction trajectory when the network parameters of the action prediction network are updated. The selection probability ratio characterizes the ratio of the probability of selecting the first control action under the current network parameters of the action prediction network, given the same first running data, to the probability of selecting the first control action when the network parameters of the action prediction network were not updated previously. Based on the single-sample pruning strategy objective item corresponding to each first action interaction trajectory and the total number of action interaction trajectories in the trajectory set, the mean value of the strategy optimization objective function corresponding to the trajectory set is calculated. The network parameters of the current action prediction network are updated with the goal of maximizing the mean of the objective function of the strategy.
[0047] Understandably, this invention constrains the selection probability ratio by setting a preset probability ratio pruning interval, which can prevent drastic changes in the strategy during a single parameter update. For example, when the selection probability of the current network for the same action differs significantly from that of the previous network, the pruning mechanism limits the optimization contribution weight of that sample, avoiding the model learning extreme or unstable control strategies due to aggressive updates, ensuring smooth convergence during training, and preventing the problem of performance fluctuations during training leading to abrupt changes in control actions during application.
[0048] This invention combines the action advantage function value (used to quantify the actual contribution of control actions to the system's operational benefits) with the selection probability ratio (used to characterize the consistency between the old and new strategies). Single-sample pruning of strategy objective terms can accurately distinguish the optimization value of different action interaction trajectories. Samples with high advantage values and strong consistency between the old and new strategies are assigned higher optimization weights, while samples with low advantage values or excessively different strategies are appropriately suppressed. This allows the model to focus on high-value, high-reliability trajectory samples during batch sampling training and updates, avoiding interference from invalid data and improving the efficiency of the model in learning multi-objective balance control laws.
[0049] By introducing the probability ratio, the current policy's utilization of high-quality actions is preserved, while the reasonable pruning interval leaves room for policy exploration, preventing the model from getting trapped in local optima. This allows the action prediction network to not only solidify the optimal balance strategy under known conditions but also adapt to unseen complex conditions, further improving the model's generalization and robustness across all conditions.
[0050] Furthermore, by calculating the mean of the objective terms of all single-sample pruning strategies in the trajectory set, the mean of the strategy optimization objective function is obtained, which can offset the random bias of individual samples (such as the fluctuation of accidental reward values under extreme conditions). This mean can more objectively reflect the overall optimization direction contained in the batch of samples, ensuring that parameter updates are based on group optimality rather than individual randomness, making the control strategy learned by the model more universal, avoiding the model from being biased towards a single objective due to individual abnormal samples, and thus always learning and updating around the core objective of optimal balance between safety, operation and energy utilization.
[0051] As an illustration, the target term of the single-sample pruning strategy can be calculated using the following formula: ; in, Indicates in First control action at all times The target item of the single-sample pruning strategy, Indicates in First control action at all times The ratio of the probability of selection, Indicates in First control action at all times The action advantage function value, This represents the clipping function. Indicates will The value of is restricted to The value is retrieved from within. For the preset selection probability range, This represents the preset hyperparameters.
[0052] Indicatively, the mean of the policy optimization objective function is the average conservative profit improvement value corresponding to the entire trajectory set. It is the final optimization target for updating the parameters of the action prediction network. It can measure the overall optimization degree of the new policy compared to the old policy under the current batch set. The larger the value, the better the policy.
[0053] It is understood that, in each batch sampling and network parameter update of the embodiments of the present invention, the single-sample optimization contribution of each sample in the trajectory set under the safety pruning constraint can be calculated, and the contributions of all samples can be averaged to obtain an objective function value representing the overall strategy optimization level. Then, the parameters of the action prediction network are adjusted to maximize this value, so that the action prediction network learns better coordinated control actions of the photovoltaic-hydrogen storage microgrid. Subsequently, the corresponding control actions can be generated more accurately based on each operating data (i.e., the operating status of the photovoltaic-hydrogen storage microgrid system in each time period).
[0054] In illustrative terms, embodiments of the present invention can store action interaction trajectory data into a shared policy experience pool. When the policy experience pool data accumulates to a preset threshold, an advantage estimation algorithm is used to calculate the advantage value of each action, and the network parameters are updated through an objective function with a pruning mechanism. Based on the updated network parameters, the action prediction network can better make decisions on control actions in subsequent time periods. Thus, through the above-mentioned continuous experience sampling and parameter updates, online self-learning and adaptive optimization of the action selection strategy are achieved.
[0055] In a preferred embodiment, the action prediction network of the present invention serves as a deep reinforcement learning (DRL) controller, and its network structure comprises two functionally independent but cooperative deep neural networks: an action network (Actor) and a critique network (Critic). For action networks, they can receive state vectors. It outputs the probability distribution of each discrete action. Its hidden layer uses the Tanh activation function, which, as the hyperbolic tangent function, is a commonly used nonlinear activation function in deep learning. It is sensitive to dynamic changes in the system and is beneficial for achieving smooth control. The action network includes: Input layer: Number of neurons and state vector The dimensions are consistent, receiving full-dimensional operating status including DC bus voltage UDC, battery state of charge (SOC), hydrogen storage tank pressure, photovoltaic output power, current of each unit and historical operation; Hidden layers: Two fully connected hidden layers are used, with 128 and 64 neurons respectively. The Tanh activation function is used, with an output range of [...]. [1,1] It has advantages such as central symmetry, smooth differentiability and gradient stability, and can be highly sensitive to millisecond-level voltage fluctuations or power ramp rate changes commonly found in photovoltaic hydrogen storage systems, which is conducive to achieving smooth adjustment of electrolyzer power and precise switching of energy storage charging and discharging.
[0056] Output layer: It has 2 neurons (corresponding to the shape parameters α and β of the Beta distribution) and is used to describe the probability distribution of the discrete action space (photovoltaic converter mode, hydrogen production power level, energy storage operation status). The Beta distribution is a continuous probability distribution defined on the interval (0,1), and its core function is to describe the probability of probability.
[0057] For the evaluation network, it can receive the same state vector. Output the long-term value of this state. Estimate: Input layer: Consistent with the action network, it receives the state vector. This ensures a consistent approach to extracting state features, preventing policy-value mismatches caused by differences in representation. Hidden layers: The structure is the same as the hidden layers of the action network.
[0058] Output layer: Only one neuron outputs a single-valued state value function V(st). This value function comprehensively reflects the advantages and disadvantages of the current system operating state in terms of hydrogen production potential, system stability margin, and equipment safety margin. It can guide the DRL controller from short-sighted operation to long-term collaboration.
[0059] This invention can use an improved proximal policy optimization (PPO) algorithm to train a deep reinforcement learning DRL controller to achieve efficient collaborative control between photovoltaic, energy storage and hydrogen production units.
[0060] like Figure 4As shown in the training flowchart of the DRL controller, in this embodiment of the invention, at startup, the parameters of Actor (action prediction network) and Critic (evaluation network) are first initialized, and the distributed interactive environment is prepared. The DRL controller acts as an intelligent agent in a microgrid environment, performing control actions and collecting state data. ,action ,award Next state Closed-loop interaction data is stored in the strategy experience pool; Once the data volume in the strategy experience pool reaches the target, data is extracted from the strategy experience pool, and the advantage value of each action is calculated using the GAE (Generalized Advantage Estimation) algorithm, which serves as the core signal for strategy optimization.
[0061] Based on the advantage value and the probability ratio of the old and new policies, a pruned PPO objective function is constructed. The Actor network parameters are updated by maximizing the mean of the objective function, while the Critic network is updated simultaneously to improve the value estimation accuracy.
[0062] Check if the training has converged or reached the required number of iterations. If it has not converged, return to the sampling stage and continue the loop until a stable coordinated control strategy is obtained.
[0063] In embodiments of the present invention, unlike the existing PPO algorithm training process, the present invention calculates the reward through a reward function. Then, it can be combined with the GAE advantage estimation algorithm, which can balance multiple optimization objectives such as safety performance, operation performance and energy utilization rate of photovoltaic-hydrogen storage microgrid system when calculating the contribution of each action, so as not to sacrifice the overall performance due to over-optimization of a single objective.
[0064] As an illustration, this invention also introduces an importance-weighted sampling mechanism for the samples in the strategy experience pool, based on the reward... The absolute value of the deviation from the historical average reward is used to assign sampling weights. Samples with larger deviations (i.e., more complex or abnormal operating conditions) are sampled first, allowing the DRL controller to focus on learning control strategies in high-risk or high-value scenarios. Finally, the system determines whether the training termination condition is met. If the maximum number of iterations is reached, or key performance indicators such as voltage stability and hydrogen production efficiency are met, training ends and the final strategy is output. Otherwise, parallel sampling and optimization continue until the action prediction network can obtain a stable, efficient, and robust coordinated control strategy.
[0065] In illustrative terms, embodiments of the present invention can periodically and randomly sample a batch of experience data from the experience pool for network updates. To more accurately evaluate the merits of actions, the present invention employs the Generalized Advantage Estimation (GAE) method to calculate the advantage function Ât, whose expression is: ; ; In the formula, γ is the discount factor, and λ is the parameter of the generalized dominance estimation function. Indicates the state The expected value of taking all possible actions This refers to timing difference errors. Therefore, this invention can fuse multi-step timing difference errors. This effectively reduces the variance of advantage estimation and improves the accuracy of policy gradient.
[0066] Furthermore, the action network parameters can be updated using an objective function with a pruning mechanism. This function ensures the stability of the training process by limiting the magnitude of each update, preventing drastic policy oscillations that could lead to system instability. Simultaneously, the evaluation network is updated to better estimate state values.
[0067] Based on the aforementioned advantages, the PPO algorithm updates the action network using an objective function with a pruning mechanism, defined as follows: ; The objective function of the above-mentioned target pruning mechanism can construct a conservative policy update region through the dual mechanism of minimization and pruning. This avoids dangerous behaviors such as DC bus voltage collapse, frequent start-up and shutdown of electrolytic cells, or deep over-discharge of batteries caused by aggressive adjustments, thereby ensuring the safety and convergence of the training process.
[0068] Using this objective function as the optimization target, the gradient ascent method is employed to update the parameters θ of the action network. The learning rate can be set to 0.0003, with 10 iterations per training round. Simultaneously, the evaluation is updated synchronously by minimizing the mean squared error of the value function. When the cumulative update steps of the action network reach a preset threshold (e.g., 200 steps), the current policy parameters are copied to the old policy network for use in the next round of objective function calculation with a pruning mechanism.
[0069] As an illustration, an action prediction network that has been trained and can autonomously learn and update itself in real time based on the trajectory of the policy experience pool can perform the following process in practical applications: At each control moment or time cycle, state-aware operations are performed to comprehensively collect the current operating data of the photovoltaic-hydrogen storage microgrid system. The collected operating parameters include DC bus voltage, photovoltaic power, energy storage system state of charge, hydrogen storage system pressure, and real-time current and power of each photovoltaic, energy storage, and hydrogen production unit, ultimately constructing a 9-dimensional state vector. .
[0070] The obtained state vector The input is fed into the action prediction network (i.e., the DRL controller) to enable the action prediction network to process the state vector. The system undergoes in-depth analysis of its multi-dimensional characteristics, and, combined with engineering safety requirements such as voltage constraints, energy storage constraints, and power balance constraints, calculates the optimal action that achieves the best long-term system benefit. This action The commands are three-dimensional discrete control commands, specifically including photovoltaic converter operating mode commands (MPV: 0 for MPPT mode, 1 for CVC mode), hydrogen production power supply power level commands (MEL: 0 for low power, 1 for medium power, 2 for high power), and energy storage converter operating status commands (MB: 0 for standby, 1 for charging, 2 for discharging). Output the optimal action This is converted into specific control commands that each unit can directly execute, and then sent to the photovoltaic converter, hydrogen production converter, and energy storage converter to drive each device to adjust its operating status according to the commands. For example, when the action... When MPV=0, MEL=2 and MB=1, the photovoltaic converter will be controlled to operate in maximum power point tracking mode, the hydrogen production power supply will operate at 100% rated power, and the energy storage converter will start charging mode, thereby realizing real-time control of the system and ensuring the system's rapid response under dynamic operating conditions. In performing the action After that, the system enters a new operating state. The present invention synchronously initiates the system feedback process, based on And a pre-designed multi-objective reward function to comprehensively evaluate actions. The impact on the system is calculated by this action. The corresponding instant reward r t It quantifies the actions A balance between safety performance, operational performance, and energy efficiency; The empirical data generated in this control cycle ( , ,r t , The data is stored in a shared policy experience pool. Once the experience data accumulates to a preset threshold, the policy update process is initiated. Then, the generalized advantage estimation (GAE) method is used to calculate the advantage value of each state-action pair. Combined with an objective function with a pruning mechanism, the parameters of the action prediction network and the evaluation network are periodically updated to ensure that policy optimization always moves in the direction of maximizing the expected cumulative reward in the future.
[0071] Meanwhile, through multi-environment parallel sampling and importance-weighted sampling mechanisms, the action prediction network can focus on learning control strategies under complex or abnormal working conditions; thus, the action prediction network of this invention has online self-learning capabilities, and can continuously accumulate historical experience and optimize control strategies to ensure the long-term effectiveness and robustness of the control method.
[0072] In a preferred embodiment, the process of constructing the reward function includes: Based on the real-time voltage data and rated voltage data of the DC bus, a voltage stability bonus data item is constructed; Based on the state of charge data of energy storage batteries and multiple different charge warning level thresholds, a reward data item for safe operation of charged batteries is constructed. Based on the pressure status data of the hydrogen storage tank and multiple different pressure warning level thresholds, construct reward data items for the safe operation of the hydrogen storage system. Based on the switching data of photovoltaic power generation system operation mode, the switching data of energy storage system power operation level, and the switching data of hydrogen production system working status, an action switching penalty data item is constructed. Based on the operating power data of the hydrogen production system, the operating power data of the photovoltaic power generation system, the operating power data of the energy storage system, the upper limit power of the photovoltaic system, the photovoltaic efficiency, the converter efficiency, and the electrolyzer efficiency of the hydrogen production system, a system operating efficiency bonus data item is constructed. Based on the standard deviation data of DC voltage, the standard deviation data of hydrogen production power, and the rated value of hydrogen production power, a system operation stability bonus data item is constructed. Based on the current data of the photovoltaic system, the rated current of the photovoltaic system, the current data of the electrolyzer of the hydrogen production system, the rated current of the electrolyzer, the hydrogen production power, and the rated hydrogen production power, construct equipment life protection incentive data items; Based on the preset hydrogen production reward weighting coefficient, actual hydrogen production data, theoretical hydrogen production data, hydrogen production power, and rated hydrogen production power, a hydrogen production reward data item is constructed. Based on the voltage stability reward data item, the first weight value corresponding to the voltage stability reward data item, the charged safe operation reward data item, the second weight value corresponding to the charged safe operation reward data item, the hydrogen storage system safe operation reward data item, and the third weight value corresponding to the hydrogen storage system safe operation reward data item, a first data item for evaluating the safety performance of the photovoltaic hydrogen storage microgrid system is constructed. Based on the action switching penalty data item, the fourth weight value corresponding to the action switching penalty data item, the system operation efficiency reward data item, the fifth weight value corresponding to the system operation efficiency reward data item, the system operation stability reward data item, and the sixth weight value corresponding to the system operation stability reward data item, a second data item is constructed to evaluate the operation performance of the photovoltaic hydrogen storage microgrid system. Based on the equipment life protection reward data item, the seventh weight value corresponding to the equipment life protection reward data item, the hydrogen production reward data item, and the eighth weight value corresponding to the hydrogen production reward data item, a third data item is constructed to evaluate the energy utilization rate of the photovoltaic-hydrogen storage microgrid system. The reward function is constructed based on the first data item, the second data item, and the third data item.
[0073] Specifically, the reward function includes: ; ; ; ; ; ; ; ; ; ; in, This represents the reward function value corresponding to the reward function expression. The data item is awarded as a reward for voltage stability. This is a data item for rewarding safe operation of charged batteries. Data items awarded for the safe operation of hydrogen storage systems. To switch the penalty data item for the action, Data items are awarded to improve system operating efficiency. Data items awarded for system operational stability This is a data item for equipment lifespan protection rewards. This is a data item for hydrogen production rewards. As the first weight value, This is the second weight value. As the third weight value, It is the fourth weight value. It is the fifth weight value. It is the sixth weight value. It is the seventh weight value. It is the eighth weight value. This is the real-time voltage of the DC bus. This is the rated voltage of the DC bus. This refers to the state of charge of the energy storage battery. The pressure status of the hydrogen storage tank; For the first The penalty value for each switching action. For indicator functions, when hour, A value of 1 indicates that a switching action occurred between the current time t and the previous time. The corresponding penalty value is valid; when hour, A value of 0 indicates that there was no switching action between the current time t and the previous time. The corresponding penalty value is invalid; , , ; The overall operating efficiency of the photovoltaic-hydrogen storage microgrid system, For hydrogen production capacity, This represents the upper limit of the photovoltaic system's power. This represents the actual power of the photovoltaic system. For the power of the storage battery, For photovoltaic efficiency, For converter efficiency, The efficiency of the electrolyzer in the hydrogen production system is given by the coefficient 0.95, which represents the system's loss factor. The standard deviation of DC voltage, The standard deviation of hydrogen production power. This is the rated value for hydrogen production capacity. For the current of the photovoltaic system, This refers to the rated current of the photovoltaic system. This refers to the current in the electrolyzer of the hydrogen production system. This is the rated current of the electrolytic cell. This represents the actual production of hydrogen. This represents the theoretical hydrogen production.
[0074] In illustrative terms, the response operation data mentioned in the embodiments of the present invention, or the response operation data samples used when training the action prediction network, include: real-time voltage of the DC bus, state of charge of the energy storage battery, pressure state of the hydrogen storage tank, switching action of the photovoltaic power generation system operating mode, switching action of the energy storage system power operation level, switching action of the hydrogen production system operating state, operating power of the hydrogen production system, operating power of the photovoltaic power generation system, operating power of the energy storage system, photovoltaic efficiency, converter efficiency, electrolyzer efficiency of the hydrogen production system, standard deviation of DC voltage, standard deviation of hydrogen production power, current of the photovoltaic system, electrolyzer current of the hydrogen production system, hydrogen production power, actual hydrogen production, and hydrogen production power.
[0075] It is not difficult to see that by substituting the response operation data into the reward function, the target reward function value corresponding to a certain control action can be calculated. Therefore, by substituting the specific voltage, current, and power data in the response operation data into the reward function, the corresponding reward function value can be obtained.
[0076] Indicatively, in the reward function above, each sub-item corresponds to objectives such as voltage stability, battery safety, hydrogen storage safety, smooth operation, system efficiency, stable operation, equipment protection, and hydrogen production. Weighting coefficients. to Used to balance the priorities of different objectives; for This function employs a piecewise design. When the voltage deviation exceeds ±10% of the rated value, a large negative reward is given to penalize severe over-limit behavior; when the deviation is within ±5% to ±10%, a moderate negative reward is given to encourage rapid recovery; when the deviation is less than ±2%, a positive reward is given to incentivize the system to maintain high-precision and stable operation; otherwise, there is no reward. Through this nonlinear reward mechanism, the system can prioritize avoiding extreme abnormal states and pursue higher stability within the normal range, thereby achieving fine-grained control of the DC voltage.
[0077] for , The function value is used to quantify the safety level and operational quality of the energy storage battery's State of Charge (SOC). To maintain the DC bus voltage near its rated value and ensure stable system operation, a piecewise quadratic penalty function is employed, imposing heavy penalties for severe voltage deviations and providing positive rewards when the voltage is highly stable to encourage precise control; a dead-zone design avoids oversensitivity to minor fluctuations. The function employs a piecewise nonlinear design: when SOC > 0.9, it is considered dangerous overcharge, and a strong negative penalty is applied to prevent battery thermal runaway; when SOC > 0.8, it is considered a warning overcharge, and a moderate negative reward is given to prompt the system to reduce charging power; when SOC < 0.1, it is considered dangerous over-discharge, and a strong negative penalty is applied to prevent permanent battery damage; when SOC < 0.2, it is considered a warning over-discharge, and a moderate negative reward is given to encourage the system to start charging; when 0.4 ≤ SOC ≤ 0.6, it is in the optimal operating range, and a positive reward is given to incentivize the system to maintain within this range; other cases receive zero reward. Through this mechanism, the system can optimize its long-term cycle life and operating efficiency while ensuring battery safety.
[0078] for , The function value is used to quantify the safety level and operational quality of the hydrogen storage system's pressure state. To maintain the hydrogen storage system within a safe pressure range, prevent accelerated equipment aging due to prolonged high or low pressure operation, and ensure its health status remains within a reasonable range, this... The function employs a piecewise nonlinear design, when A pressure >0.9 is considered an overpressure hazard, and a stronger negative penalty is applied to prevent structural damage to the hydrogen storage tank; when When the value is greater than 0.8, it is considered a high-pressure warning, and a moderate negative reward is given, prompting the system to reduce hydrogen production capacity; when When the pressure is less than 0.1, it is considered a dangerous undervoltage condition, and a strong negative penalty is applied to prevent insufficient hydrogen supply from causing a shutdown; when When the pressure is less than 0.2, it is considered a low-pressure warning, and a moderate negative reward is given to prompt the system to start gas replenishment; when 0.3 ≤ When the value is ≤0.7, it is in the comfort zone and a positive reward is given, and the incentive system maintains this range; otherwise, there is no reward.
[0079] for To reduce frequent device start-ups, shutdowns, and mode switching, by Negative rewards can be applied to situations where actions change within adjacent control cycles, thereby reducing mechanical and electrical stress and extending equipment lifespan. Specifically, applying negative rewards can suppress unnecessary action jumps, thus improving the smoothness of the control strategy and the reliability of equipment operation, making it suitable for multi-source collaborative control scenarios.
[0080] for In order to maximize the utilization rate of renewable energy and improve the overall efficiency of the system, through Taking into account the photovoltaic power generation efficiency, converter conversion efficiency, and electrolyzer hydrogen production efficiency, positive incentives are given to high-efficiency operation. By normalizing the function, the hydrogen production ratio and power matching degree can be combined, guiding the network to prioritize high-efficiency power conversion and distribution, thereby improving the overall economy and sustainability of the system. for To suppress short-term, drastic fluctuations in power and voltage and encourage stable system operation, this invention... The function can penalize situations where the rate of change of power per unit time exceeds a threshold. The function guides the controller to achieve stable hydrogen energy output while ensuring voltage stability by setting dual threshold conditions, thereby improving the overall system operation quality and equipment lifespan.
[0081] for To prevent critical equipment (such as electrolytic cells and converters) from entering extreme operating conditions, negative rewards are given when current, temperature, or pressure approaches safety limits to mitigate risk; specifically, By setting multiple levels of penalties and positive incentives, the function effectively prevents equipment overload, overheating, and lifespan degradation, thereby improving system reliability and operational economy.
[0082] for Under the premise of meeting safety and stability constraints, through The function provides a positive reward for actual hydrogen production, incentivizing the system to maximize green hydrogen output within the efficient and safe operating range. Specifically, The function, by coupling hydrogen production efficiency with power utilization, guides the controller to pursue high energy conversion efficiency while maximizing hydrogen production, thereby improving the system's economy and sustainability.
[0083] In illustrative terms, the eight sub-items constructed in this embodiment of the invention comprehensively cover the core dimensions of safety performance, operational performance, and energy utilization efficiency of the photovoltaic-hydrogen storage microgrid, avoiding the one-sidedness of traditional reward functions that only focus on a single objective. By weighting each sub-item with weight values from the first to the eighth, and integrating them hierarchically according to safety performance (first data item), operational performance (second data item), and energy utilization efficiency (third data item), a flexible trade-off between multiple objectives is achieved, adapting to different needs such as prioritizing safety, hydrogen production, or energy efficiency.
[0084] Each sub-item adopts a piecewise nonlinear design or quantization calculation logic. For example, the voltage stability reward item is graded and rewarded according to the deviation amplitude, the charge safety item is differentiated and incentivized according to the SOC range, and the action switching item is frequently switched with targeted penalties. This allows the reward function to accurately quantify the actual effect of different control actions. As a result, the action prediction network can clearly perceive which actions will lead to safety risks or which actions can improve the overall performance during training, and thus learn a refined optimal control strategy autonomously.
[0085] The reward function of this invention can accurately quantify the comprehensive balance of each candidate control action, select the training label with the best balance of safety, operation and energy efficiency for the action prediction network, and finally learn and solidify the optimal control logic under all working conditions, ensuring that the control action output in the application stage takes into account multi-dimensional performance, and realizes the stable and efficient operation of the system throughout the entire cycle.
[0086] Please see Figure 5 This application also provides a coordinated control device based on a photovoltaic-hydrogen storage microgrid system, which can realize the above-mentioned coordinated control method based on a photovoltaic-hydrogen storage microgrid system. The device includes: The operation data acquisition module is used to acquire the current operation data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operation data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; The control action generation module is used to input the current operating data into a preset action prediction network, so that the action prediction network extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics based on the current operating data, and then outputs the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics; wherein, the target control action includes: target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system respectively; The system control module is used to control the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system according to the target control action. The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action; the candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample; wherein, the reward function is used to evaluate the balance between the safety performance, operational performance and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action.
[0087] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0088] It should be noted that the device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0089] Those skilled in the art will clearly understand that, for convenience and simplicity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0090] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned coordinated control method for a photovoltaic-hydrogen storage microgrid system. This electronic device can include any smart terminal such as a tablet computer or an in-vehicle computer.
[0091] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0092] Please see Figure 6 , Figure 6 This illustrates the hardware structure of an electronic device according to another embodiment, the electronic device comprising: The processor can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to achieve the technical solutions provided in the embodiments of this application. The memory can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory can store the operating system and other applications. When the technical solutions provided in the embodiments of this application are implemented through software or firmware, the relevant program code is stored in the memory and called and executed by the processor. Input / output interfaces are used to implement information input and output; The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). A bus is used to transfer information between various components of a device, such as processors, memory, input / output interfaces, and communication interfaces. The processor, memory, input / output interface, and communication interface are interconnected within the device via a bus.
[0093] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting various parts of the terminal device via various interfaces and lines.
[0094] The memory can be used to store the computer program. The processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, SmartMediaCard (SMC), Secure Digital (SD) card, FlashCard, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0095] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned coordinated control method for a photovoltaic-hydrogen storage microgrid system.
[0096] It is understood that the content of the above method embodiments is applicable to the present computer storage medium embodiments. The specific functions implemented by the present computer storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0097] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described coordinated control method for a photovoltaic-hydrogen storage microgrid system.
[0098] It is understood that the content of the above method embodiments is applicable to the embodiments of this computer program product. The specific functions implemented by the embodiments of this computer program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0099] Those skilled in the art will understand that all or some of the steps, apparatuses, or functional modules / units in the methods disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0100] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A coordinated control method for a photovoltaic-hydrogen storage microgrid system, characterized in that, The method includes: Acquire the current operating data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operating data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; The current operating data is input into a preset action prediction network, which extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics based on the current operating data. Then, based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics, the network outputs the current control action. The target control action includes target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system, respectively. Based on the target control action, the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system are controlled respectively; The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action; the candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample; wherein, the reward function is used to evaluate the balance between the safety performance, operational performance and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action.
2. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 1, characterized in that, The power characteristics include: photovoltaic power characteristics, energy storage power characteristics, and hydrogen production system power characteristics; the current characteristics include: photovoltaic current characteristics, energy storage current characteristics, and hydrogen production current characteristics; the target operating state control commands include: operating mode control commands, power control commands, and working state control commands. The step of outputting the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics includes: Based on the characteristics of bus voltage, photovoltaic power, and photovoltaic current, photovoltaic operating condition characteristics are generated. Based on the characteristics of bus voltage, energy storage power, energy storage current, and state of charge, the operating condition characteristics of energy storage are generated. Based on the characteristics of bus voltage, power of hydrogen production system, hydrogen production current, and hydrogen energy storage, the operating conditions of hydrogen production are generated. The photovoltaic operating condition characteristics, the energy storage operating condition characteristics, and the hydrogen production operating condition characteristics are fused together to output system operating condition characteristics that characterize the global operating condition of the photovoltaic-hydrogen storage microgrid system. The system operating condition characteristics, photovoltaic operating condition characteristics, energy storage operating condition characteristics, and hydrogen production operating condition characteristics are weighted and fused, and operating mode control commands for photovoltaic power generation system, power control commands for energy storage system, and working state control commands for hydrogen production system are generated under DC voltage constraints, energy storage state of charge constraints, power balance constraints, and equipment action constraints. The operating mode control command, the power control command, and the working state control command are output as target control actions.
3. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 2, characterized in that, Also includes: When the target control action is detected to have been executed, the response operation data of the photovoltaic hydrogen storage microgrid system in the next time period is acquired; Substitute the response operation data into the reward function to calculate the target reward function value corresponding to the target control action; The current operating data, target control actions, operating data of the photovoltaic-hydrogen storage microgrid system for the next time period, and objective function values are used as the target action interaction trajectory and added to the strategy experience pool corresponding to the action prediction network; wherein, the strategy experience pool is used to update the network parameters of the action prediction network.
4. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 3, characterized in that, Also includes: When the number of action interaction trajectories in the strategy experience pool exceeds a preset trajectory number threshold, several action interaction trajectories in the strategy experience pool are used as a trajectory set; wherein, the trajectory set includes: several first action interaction trajectories; the first action interaction trajectory includes: first running data of a first time period, a first control action corresponding to the first running data, a first reward function value corresponding to the first control action, and first response running data corresponding to the next time period corresponding to the first time period; For each first action interaction trajectory, the expected benefit value of the first operating data and the expected benefit value of the first response operating data are determined through the evaluation network corresponding to the action prediction network; wherein, the evaluation network is used to evaluate the role of the operating data in a time period on the operating benefit of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, the first reward function value corresponding to the first control action, the expected reward value of the first running data, and the expected reward value corresponding to the first response running data are substituted into the action advantage function to determine the action advantage function value of the first control action; wherein, the action advantage function is used to evaluate the role of the first control action in the operational benefits of the photovoltaic-hydrogen storage microgrid system. For each first action interaction trajectory, a single-sample pruning strategy objective term is calculated based on the selection probability ratio corresponding to the first control action, the action dominance function value, and a preset probability ratio pruning interval. This objective term characterizes the contribution of the first action interaction trajectory to the optimization of the action prediction network. The single-sample pruning strategy objective term indicates the degree of optimization contribution of the first action interaction trajectory when the network parameters of the action prediction network are updated. The selection probability ratio characterizes the ratio of the probability of selecting the first control action under the current network parameters of the action prediction network, given the same first running data, to the probability of selecting the first control action when the network parameters of the action prediction network were not updated previously. Based on the single-sample pruning strategy objective item corresponding to each first action interaction trajectory and the total number of action interaction trajectories in the trajectory set, the mean value of the strategy optimization objective function corresponding to the trajectory set is calculated. The network parameters of the current action prediction network are updated with the goal of maximizing the mean of the objective function of the strategy.
5. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 4, characterized in that, The target term of the single-sample pruning strategy is calculated according to the following formula: ; in, Indicates in First control action at all times The target item of the single-sample pruning strategy, Indicates in First control action at all times The ratio of the probability of selection, Indicates in First control action at all times The action advantage function value, This represents the clipping function. Indicates will The value of is restricted to The value is retrieved from within. For the preset selection probability range, This represents the preset hyperparameters.
6. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 5, characterized in that, The process of constructing the reward function includes: Based on the real-time voltage data and rated voltage data of the DC bus, a voltage stability bonus data item is constructed; Based on the state of charge data of energy storage batteries and multiple different charge warning level thresholds, a reward data item for safe operation of charged batteries is constructed. Based on the pressure status data of the hydrogen storage tank and multiple different pressure warning level thresholds, construct reward data items for the safe operation of the hydrogen storage system. Based on the switching data of photovoltaic power generation system operation mode, the switching data of energy storage system power operation level, and the switching data of hydrogen production system working status, an action switching penalty data item is constructed. Based on the operating power data of the hydrogen production system, the operating power data of the photovoltaic power generation system, the operating power data of the energy storage system, the upper limit power of the photovoltaic system, the photovoltaic efficiency, the converter efficiency, and the electrolyzer efficiency of the hydrogen production system, a system operating efficiency bonus data item is constructed. Based on the standard deviation data of DC voltage, the standard deviation data of hydrogen production power, and the rated value of hydrogen production power, a system operation stability bonus data item is constructed. Based on the current data of the photovoltaic system, the rated current of the photovoltaic system, the current data of the electrolyzer of the hydrogen production system, the rated current of the electrolyzer, the hydrogen production power, and the rated hydrogen production power, construct equipment life protection incentive data items; Based on the preset hydrogen production reward weighting coefficient, actual hydrogen production data, theoretical hydrogen production data, hydrogen production power, and rated hydrogen production power, a hydrogen production reward data item is constructed. Based on the voltage stability reward data item, the first weight value corresponding to the voltage stability reward data item, the charged safe operation reward data item, the second weight value corresponding to the charged safe operation reward data item, the hydrogen storage system safe operation reward data item, and the third weight value corresponding to the hydrogen storage system safe operation reward data item, a first data item for evaluating the safety performance of the photovoltaic hydrogen storage microgrid system is constructed. Based on the action switching penalty data item, the fourth weight value corresponding to the action switching penalty data item, the system operation efficiency reward data item, the fifth weight value corresponding to the system operation efficiency reward data item, the system operation stability reward data item, and the sixth weight value corresponding to the system operation stability reward data item, a second data item is constructed to evaluate the operation performance of the photovoltaic hydrogen storage microgrid system. Based on the equipment life protection reward data item, the seventh weight value corresponding to the equipment life protection reward data item, the hydrogen production reward data item, and the eighth weight value corresponding to the hydrogen production reward data item, a third data item is constructed to evaluate the energy utilization rate of the photovoltaic-hydrogen storage microgrid system. The reward function is constructed based on the first data item, the second data item, and the third data item.
7. The coordinated control method for a photovoltaic-hydrogen storage microgrid system according to claim 6, characterized in that, The reward function includes: ; ; ; ; ; ; ; ; ; ; in, This represents the reward function value corresponding to the reward function expression. The data item is awarded as a reward for voltage stability. This is a data item for rewarding safe operation of charged batteries. Data items awarded for the safe operation of hydrogen storage systems. To switch the penalty data item for the action, Data items are awarded to improve system operating efficiency. Data items awarded for system operational stability This is a data item for equipment lifespan protection rewards. This is a data item for hydrogen production rewards. As the first weight value, This is the second weight value. As the third weight value, It is the fourth weight value. It is the fifth weight value. It is the sixth weight value. It is the seventh weight value. It is the eighth weight value. This is the real-time voltage of the DC bus. The rated voltage of the DC bus. This refers to the state of charge of the energy storage battery. The pressure status of the hydrogen storage tank; For the first The penalty value for each switching action. For indicator functions, when hour, A value of 1 indicates that a switching action occurred between the current time t and the previous time. The corresponding penalty value is valid; when hour, A value of 0 indicates that there was no switching action between the current time t and the previous time. The corresponding penalty value is invalid; , , ; The overall operating efficiency of the photovoltaic-hydrogen storage microgrid system, For hydrogen production capacity, This represents the upper limit of the photovoltaic system's power. This represents the actual power of the photovoltaic system. For the power of the storage battery, For photovoltaic efficiency, For converter efficiency, The efficiency of the electrolyzer in the hydrogen production system is given by the coefficient 0.95, which represents the system's loss factor. The standard deviation of DC voltage, The standard deviation of hydrogen production power. This is the rated value for hydrogen production capacity. For the current of the photovoltaic system, This refers to the rated current of the photovoltaic system. This refers to the current in the electrolyzer of the hydrogen production system. This is the rated current of the electrolytic cell. This represents the actual production of hydrogen. This represents the theoretical hydrogen production.
8. A coordinated control device based on a photovoltaic-hydrogen storage microgrid system, characterized in that, The device includes: The operation data acquisition module is used to acquire the current operation data of the photovoltaic-hydrogen storage microgrid system in the current time period; wherein, the current operation data includes: current power data, current current data, DC bus voltage, current state of charge of the energy storage system, and current hydrogen energy storage pressure status of the hydrogen production system; The control action generation module is used to input the current operating data into a preset action prediction network, so that the action prediction network extracts power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics based on the current operating data, and then outputs the target control action corresponding to the current operating data based on the power characteristics, current characteristics, bus voltage characteristics, state of charge characteristics, and hydrogen energy storage characteristics; wherein, the target control action includes: target operating state control commands corresponding to the photovoltaic power generation system, energy storage system, and hydrogen production system respectively; The system control module is used to control the operating status of the photovoltaic power generation system, energy storage system and hydrogen production system in the photovoltaic-hydrogen storage microgrid system according to the target control action. The generation process of the action prediction network includes: The action prediction network to be trained is iteratively trained by taking the operation data samples from different historical time periods and the actual control action corresponding to each operation data sample as inputs and the predicted control action corresponding to each operation data sample as output, until the model converges and the preset action prediction network is generated. The process of determining the actual control actions corresponding to the operational data samples includes: For each candidate control action of each operational data sample, the response operational data sample of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action is substituted into a preset reward function to calculate the candidate reward function value corresponding to the candidate control action; the candidate control action with the largest candidate reward function value is taken as the actual control action corresponding to the operational data sample; wherein, the reward function is used to evaluate the balance between the safety performance, operational performance and energy utilization rate of the photovoltaic-hydrogen storage microgrid system after executing the candidate control action.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the coordinated control method of a photovoltaic hydrogen storage microgrid system as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a coordinated control method for a photovoltaic hydrogen storage microgrid system as described in any one of claims 1 to 7.