A method for optimizing the operation of hydrogen-containing building energy systems using a multi-role large model
By employing a safe Markov decision process assisted by a multi-role large model and a proximal policy optimization algorithm, the operation strategy of a hydrogen-containing building energy system is optimized. This solves the problems of relying on expert experience and difficulty in adjusting reward function parameters in existing technologies, and achieves efficient and economical operation and reliable power and heat supply in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing optimization methods for hydrogen-containing building energy systems suffer from problems such as reliance on expert experience making them difficult to adapt to complex environments, high computational complexity, and reliance on manual parameter tuning for reward function design. These issues result in poor system performance in multi-objective optimization and power and heat supply reliability.
A safe Markov decision process assisted by a multi-role large model is adopted, combined with a near-end policy optimization algorithm and a rule action correction mechanism, to optimize the operation strategy of a multi-energy system in a hydrogen-containing building. The system's economy and power supply and heating reliability are synergistically optimized through online decision-making by intelligent agents.
It significantly improves the system's adaptability and economy in complex environments, solves the problem of difficult reward function parameter tuning, achieves multi-objective balanced optimization, and ensures high power and heat supply reliability of the system.
Smart Images

Figure CN121544087B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of building energy system optimization and control technology, specifically a multi-role large model-assisted method for optimizing the operation of hydrogen-containing building energy systems. Background Technology
[0002] Research on the operational optimization of hydrogen-containing building energy systems helps reduce system carbon emissions and energy costs while improving user thermal comfort. Existing research has proposed various methods for optimizing the operation of hydrogen-containing building energy systems, mainly including heuristic rule-based control methods, traditional optimization methods based on mathematical programming, and intelligent decision-making methods based on deep reinforcement learning. While these methods have achieved some success, they each have significant limitations. Heuristic rule-based methods, although simple to implement, heavily rely on expert experience for control performance, making them difficult to adapt to complex and changing operating environments. Mathematical programming methods, while able to obtain theoretically optimal solutions under ideal conditions, require accurate system models and complete predictive information, and have high computational complexity, making online application difficult. Existing deep reinforcement learning-based methods, while avoiding reliance on accurate models, rely on manual parameter tuning for reward function design, making it difficult to adaptively balance multi-objective optimization needs.
[0003] In summary, existing methods for optimizing the operation of multi-energy systems in hydrogen-containing buildings have significant shortcomings. There is an urgent need to research a new operation optimization method to reduce the operating costs of hydrogen-containing building energy systems while ensuring high reliability of power and heat supply. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a multi-role large model-assisted method for optimizing the operation of hydrogen-containing building energy systems. Its purpose is to achieve synergistic optimization of the operational economy and power supply / heating reliability of multi-energy systems in hydrogen-containing buildings. This method overcomes the shortcomings of existing rule-based control methods that rely on expert experience and are difficult to adapt to complex operating environments. It also solves the technical problems of existing deep reinforcement learning-based methods, which require manual tuning of penalty parameters, resulting in slow convergence speed and difficulty in balancing multiple objectives.
[0005] To achieve the above objectives, the present invention employs the following technical solution:
[0006] This invention provides a multi-role large-scale model-assisted method for optimizing the operation of hydrogen-containing building energy systems, including:
[0007] (1) To address the issue of minimizing the operating cost of hydrogen-containing multi-energy systems in off-grid operation mode;
[0008] (2) The problem of minimizing operating costs is remodeled as a safe Markov decision process related to the operation optimization of multi-energy systems in hydrogen-containing buildings;
[0009] (3) Based on the multi-role large language model-assisted proximal policy optimization algorithm, solve the modeled safe Markov decision process to obtain the agent operation strategy related to the hydrogen-containing building multi-energy system;
[0010] (4) The intelligent agent makes online decisions based on the operation strategy of the hydrogen-containing building multi-energy system obtained through training, and applies the decisions to the actual hydrogen-containing building multi-energy system.
[0011] Furthermore, the established problem of minimizing the operating cost of a hydrogen-containing building multi-energy system under off-grid operation mode includes the objective function, decision variables, and constraints, namely:
[0012] (1) Objective function: The objective function is to minimize the expected time average of both system operating cost and energy loss, which is:
[0013] (1),
[0014] in, For system operating costs, For electrical loss, Heat loss is calculated using the following formula:
[0015] (2),
[0016] (3),
[0017] (4),
[0018] (5),
[0019] (6),
[0020] (7),
[0021] In equation (1), For expectation operators; The total operating cost of the system; in equation (2), , , In time The component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost; in formula (3), , , , , , These are the cost coefficients for photovoltaic modules, electrolyzers, fuel cells, electric boilers, ground source heat pumps, and thermal storage tanks, respectively. , , , , , , In time Photovoltaic output power and time Input power and time of the electrolytic cell Fuel cell output power and time Input power of electric boiler, input power of ground source heat pump, time Heating input power and time of the thermal storage tank The heat dissipation output power of the thermal storage tank; in equation (4), In time The purchase price of hydrogen. In time The volume of hydrogen purchased, in formula (5), This is the cost reduction coefficient for photovoltaics. In time The maximum power generation of photovoltaics; in equation (6), Let be the electrical loss penalty function, in equation (7), This is the heat loss penalty function;
[0022] (2) Decision variables: Decision variables include each time slot Photovoltaic power ,time Electrolytic cell input power ,time fuel cell output power ,time Electric boiler input power ,time Ground source heat pump input power ,time Thermal storage tank heating input power ,time Thermal storage tank heat output power and the amount of hydrogen purchased from the hydrogen market ;
[0023] (3) Constraints: Constraints include power balance and equipment physical constraints:
[0024] Photovoltaic constraints:
[0025] (8),
[0026] (9),
[0027] Hydrogen storage tanks and hydrogen balance constraints:
[0028] (10)
[0029] (11),
[0030] (12)
[0031] (13)
[0032] (14)
[0033] (15)
[0034] (16)
[0035] (17)
[0036] Electrolytic cell constraints:
[0037] (18)
[0038] (19)
[0039] (20)
[0040] Fuel cell constraints:
[0041] (twenty one),
[0042] (twenty two),
[0043] (twenty three),
[0044] (twenty four),
[0045] (25)
[0046] Thermal storage tank constraints:
[0047] (26)
[0048] (27)
[0049] (28)
[0050] (29)
[0051] (30)
[0052] Constraints of ground source heat pumps and geothermal wells:
[0053] (31),
[0054] (32),
[0055] (33),
[0056] (34),
[0057] (35),
[0058] (36)
[0059] Electric boiler constraints:
[0060] (37)
[0061] (38),
[0062] System supply and demand balance constraints:
[0063] (39)
[0064] (40)
[0065] In equation (8), The photovoltaic efficiency coefficient. The total area of the photovoltaic panels. for Solar radiation intensity in time slot; in equation (10), For time slots The hydrogen storage capacity of the hydrogen storage tank. This represents the amount of hydrogen produced by the electrolyzer within a given time period. This represents the hydrogen capacity consumed by the fuel cell within a given time period. For time slots The purchased hydrogen capacity, in formula (11) The maximum hydrogen storage capacity of the hydrogen storage tank is given in equation (12). To maximize the amount of hydrogen purchased, in equation (13), It is the gas compressibility factor. , For the fitting parameters, in equation (14), For time slots Hydrogen storage tank pressure, This refers to the volume of the hydrogen storage tank. Let be the ideal gas constant. For time slots Hydrogen storage tank temperature, Let be the molar mass of hydrogen gas, in equation (15), For the safety margin pressure of the hydrogen storage tank, It is a binary variable. To relax the constraints, in equation (16), For the maximum pressure limit of the hydrogen storage tank, in equation (17), Let be the total number of time periods during which the pressure in the hydrogen storage tank exceeds the safety margin, as given in equation (18). The capacity of hydrogen produced by the electrolyzer. The hydrogen production efficiency of the electrolyzer. For time slots Input power of the internal electrolytic cell, For the time interval, in equation (19), For the maximum input power of the electrolytic cell, in equation (20), The ramp rate of the electrolytic cell. They represent the input power of the electrolytic cell in the previous time slot, respectively. In equation (21), The amount of hydrogen consumed by the fuel cell. For time slots The output power of the internal fuel cell, For the hydrogen-to-electricity conversion efficiency of the fuel cell, in equation (22), For the maximum output power of the fuel cell, in equation (23), For the ramp rate of the fuel cell, For time slots The output power of the fuel cell at that time is given by equation (25). For time slots Waste heat generated by fuel cells, The electrothermal conversion ratio of the fuel cell, For the heat recovery coefficient, in equation (26), For time slots The thermal storage capacity of the thermal storage tank For time slots The charging power of the thermal storage tank For time slots The heat dissipation capacity of the thermal storage tank For heat charging efficiency, For heat release efficiency, in equation (27), For the maximum heat storage capacity of the heat storage tank, in equation (28), For the maximum heat charging power, in equation (29), For the maximum heat release power, in equation (31), For time slots The heat output power of a ground source heat pump For time slots The electrical power input to the ground source heat pump, For the efficiency of a ground source heat pump, in equation (32), For the maximum input electrical power of the ground source heat pump, in equation (33), For time slots The heat storage level of geothermal wells For time slots The waste heat from the fuel cell injected into the geothermal well, in equation (34), For the maximum heat storage capacity of the geothermal well, in equation (35), represents the waste heat injected into the geothermal well. No more than the waste heat generated by the fuel cell In equation (37), For time slots The heat generated by the electric boiler For time slots The electrical power input to the electric boiler, For the efficiency of the electric boiler, in equation (38), For the maximum input electrical power of the electric boiler, in equation (39), For time slots electrical load, Time slot The electrical power consumed by the electrolytic cell For time slots The electrical power consumed by a ground source heat pump For time slots The electrical loss, in equation (40), For time slots heat load, For time slots The heat consumed during the charging of the thermal storage tank. For time slots Waste heat injected into geothermal wells For time slots Waste heat generated by fuel cells, For time slots The heat generated by the ground source heat pump For time slots The heat released by the thermal storage tank For time slots Heat loss.
[0066] Furthermore, the problem of minimizing operating costs is remodeled as a Markov decision process with a safety correction mechanism related to the operational optimization of multi-energy systems in hydrogen-containing buildings, defined as a tuple. .in, For the state space related to the operation of multi-energy systems in hydrogen-containing buildings, For the original action space, This is the state transition function. For the final reward function, To solve this Markov decision process, a rule-based action correction mechanism is introduced, using the discount factor. As a pre-module for strategy execution. The specific expression is as follows:
[0067] (41),
[0068] (42),
[0069] (43),
[0070] In equation (41), The current running hour; each action component is normalized to the range of [0,1] or [-1,1], in equation (42) The normalized heat storage tank charging and discharging power; , These are the normalized power allocated by the fuel cell to the electrical load, the ground source heat pump, and the electric boiler, respectively. , The normalized amount of hydrogen purchased; in equation (43), For the final reward function, This is the basic reward function.
[0071] Furthermore, the aforementioned To ensure system supply and demand balance and physical constraints in off-grid mode, a rule-based action correction mechanism is introduced in the secure Markov decision process. Original action Corrected to actual action The correction steps are as follows:
[0072] S41, Hot Can Action Correction, according to Determine the actual heat charging power of the thermal storage tank or heat dissipation power and update the remaining heat load gap. ,when At this time, it is the heat input mode of the hot tank. ,and ,when At this time, it is the heat output mode of the hot tank. ,and ,in, For maximum charge / discharge power, At maximum capacity, the heat load gap is updated to... .
[0073] S42. Photovoltaic Consumption and Distribution Correction: Photovoltaic power generation prioritizes meeting the basic electrical load, with remaining photovoltaic power subsequently used to drive ground-source heat pumps, electric boilers, and electrolytic cells. Specifically, the photovoltaic power required to meet the basic electrical load is first calculated. Secondly, calculate the power of the photovoltaic-driven ground source heat pump. Next, calculate the power of the photovoltaic-driven electric boiler. And update the heat load gap to , , The output power of the photovoltaic system is allocated to the power of the ground source heat pump and the electric boiler, respectively.
[0074] S43. Fuel Cell Operation Correction: When photovoltaic output is insufficient, the fuel cell is used to supplement power and heat supply. The correction process must simultaneously meet the physical constraints of the original operation command, the maximum power of the equipment, and the current hydrogen storage capacity. Specifically, first, the power of the fuel cell to meet the remaining electrical load is calculated, i.e. Secondly, calculate the power of the fuel cell driving the ground source heat pump. Finally, the power of the fuel cell-driven electric boiler is calculated. ;in, The heat generated by the fuel cell ultimately determines the total output power of the fuel cell as follows: .
[0075] S44. Electrolyzer Operation Correction: The electrolyzer utilizes surplus photovoltaic power to produce hydrogen and operates only when the fuel cell is not working, to avoid energy conversion conflicts and is limited by the hydrogen storage tank capacity. Specifically, if the total output power of the fuel cell... Then the power of the electrolytic cell ;like
[0076] Then the power of the electrolytic cell ,in After allocating the photovoltaic output power to the ground source heat pump and electric boiler, the remaining power is used to calculate the hydrogen storage capacity in the intermediate state based on the operation of the electrolyzer and fuel cell. .
[0077] S45. Hydrogen purchase action correction: Determine whether to purchase hydrogen from the market based on the hydrogen storage threshold. Specifically, if the intermediate state hydrogen storage... The actual purchase quantity ,like The actual purchase quantity ,in, This is the safe threshold for hydrogen storage capacity.
[0078] In equation (43), the final reward function Includes the basic loss function Electrical loss penalty function Heat loss penalty function The electrical loss penalty function and the thermal loss penalty function both adopt a three-segment piecewise linear structure. The specific forms of the basic loss function, electrical loss penalty function, and thermal loss penalty function are as follows:
[0079] (44),
[0080] (45)
[0081] (46)
[0082] In equation (44), , , In time The calculation formulas for the component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost are shown in equations (3), (4), and (5), respectively; in equation (45), Here is a three-segment electrical loss penalty function, where, The electrical loss penalty functions are for the first, second, and third segments, respectively, and are as follows: First segment: If ,but Second paragraph: If Third paragraph: If ,but ,in, This represents the severe threshold of electrical load loss; This represents the moderate threshold of electrical load loss. , The preset electrical load loss threshold coefficient, The value is the penalty coefficient, and it satisfies the monotonicity constraint. In equation (46), Here is a three-segment heat loss penalty function, where, These are the heat loss penalty functions for the first, second, and third segments, respectively, specifically: First segment: If... ,but Second paragraph: If ,but Third paragraph: If ,but ,in The threshold indicating the severity of heat load loss; This represents the intermediate threshold of heat load loss. , This is the preset heat load loss threshold coefficient; Let be the heat loss penalty coefficient, and satisfy the monotonicity constraint, i.e.: .
[0083] Furthermore, a secure Markov decision process modeled based on a multi-role large language model-assisted proximal policy optimization algorithm is used to solve the problem. This involves alternating iterative training and parameter tuning phases until the penalty function parameters reach a predefined convergence criterion, specifically including:
[0084] S61. Training Phase: The agent interacts with the environment using the current penalty function parameters, specifically including:
[0085] 1) Initialize the policy network and value network ;
[0086] 2) The agent interacts with the environment to collect experience data and calculates the reward value based on the three-stage penalty function;
[0087] 3) Calculate the advantage function based on collected experience. and rewards ;
[0088] 4) Use the combined loss function Update network parameters:
[0089] (47)
[0090] In equation (47), Let the loss function be the shearing strategy. The loss function of the value network, Let the policy entropy reward function be... and This is the balance coefficient;
[0091] S62. Parameter Tuning Phase: After the trained model is evaluated on the test set, the penalty parameters are optimized using a multi-role large language model, specifically including:
[0092] 1) In The performance of the current strategy is evaluated on the test set, and the total operating cost and electrical and heat loss indicators are recorded.
[0093] 2) The performance of the large language model for evaluating roles is qualitatively assessed based on fuzzy interval criteria;
[0094] 3) Parameter tuning: Based on the evaluation results and guided thought chain suggestions, the large language model generates a new three-stage penalty function slope parameter (…). as well as );
[0095] S63. Repeat the above stages, clearing the agent's interaction experience after each round of training and starting training again.
[0096] Furthermore, the large language model for the evaluation role is primarily responsible for:
[0097] (1) Analysis The cumulative test results over the days are used to break down the total operating cost, hydrogen purchase cost, component operating cost, and photovoltaic abandonment cost.
[0098] (2) Based on the preset fuzzy interval evaluation criteria, the electrical loss and thermal loss performance are mapped to multi-level qualitative evaluation labels;
[0099] (3) Analyze the composition of operating costs, identify the main cost drivers, and output structured results including loss assessment and performance analysis.
[0100] Furthermore, the prompt words of the large language model system for parameter tuning roles include six core components:
[0101] (1) Role description: Defined as an expert in power system operation and reinforcement learning, capable of extracting relevant domain knowledge from embedded knowledge;
[0102] (2) Environmental Description: Provide detailed specifications for the integrated multi-energy system, including:
[0103] ①Core components: photovoltaic power generation, electrolyzer, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, thermal storage tank and geothermal well;
[0104] ②Dynamic characteristics: electrical balance, thermal balance, hydrogen management, energy storage operation;
[0105] ③ Operational constraints: efficiency factor, capacity limit, tiered load loss penalty;
[0106] (3) Task description: Outline the main optimization objectives, including ensuring that electrical and heat losses do not exceed specified thresholds, and minimizing total operating costs while maintaining load satisfaction within acceptable limits;
[0107] (4) Output format: Strictly define a fixed output structure, requiring only fixed objects to be output without any additional explanation or code;
[0108] (5) Parameter constraints: enforce mathematical monotonicity conditions as well as And maintain a coordinated relationship where thermal penalty usually exceeds electrical penalty;
[0109] (6) Reward calculation framework: specify the composite reward structure .
[0110] Furthermore, S62 includes a guided chain of thought processes, which enhances the logical consistency of parameter adjustments through multi-stage analysis:
[0111] S91. When the assessment results of electrical or thermal losses do not reach the preset satisfaction level, and the total operating cost exceeds the preset threshold... At this time, a penalty enhancement strategy or a fine-tuning strategy is adopted to increase the slope of the first segment of the corresponding loss function. Or multiple slope parameters;
[0112] S92. When the assessment results of electrical loss or heat loss reach the preset satisfaction level, but the total operating cost exceeds the preset threshold... At that time, a penalty relaxation strategy or parameter optimization strategy can be adopted to reduce the slope parameter of the corresponding loss function to improve economic efficiency;
[0113] S93. When there is a difference between the assessment results of electrical loss and heat loss, a multi-problem collaborative processing strategy is adopted, prioritizing the adjustment of the penalty parameter of the side with the worse assessment result, or a focus strategy is adopted to keep one side stable while optimizing the other side.
[0114] S94. During the adjustment process, always maintain the monotonicity constraint of the parameters, and determine the magnitude of parameter adjustment based on the training reward value.
[0115] Furthermore, when making online decisions, the agent only needs to invoke the trained policy network, that is:
[0116] (48)
[0117] In equation (48), For the trained policy network, The current state observation value, For the generated optimized actions.
[0118] Compared to existing technologies, the beneficial effects of this invention are as follows: This invention combines the advantages of large models and deep reinforcement learning, and the beneficial effects achieved by this invention are as follows:
[0119] (1) Compared with existing rule-based control methods, the method of the present invention can autonomously explore the optimal strategy in complex operating environments based on the reinforcement learning decision framework, which significantly improves the adaptability and economy of the system under changing operating conditions and overcomes the limitations of rule-based control strategies that are fixed and difficult to cope with uncertainty.
[0120] (2) Compared with existing methods based on traditional reinforcement learning, this invention achieves adaptive optimization of penalty parameters by using a multi-role large language model to assist in the design of the reward function, effectively solving the pain point of difficulty in optimizing reward function parameters in traditional deep reinforcement learning. At the same time, by introducing a chain thinking mechanism, the logic and interpretability of parameter adjustment are enhanced, enabling the system to better balance multiple objectives such as economic operation and power and heat supply reliability. Attached Figure Description
[0121] Figure 1 This is a flowchart of a multi-role large model-assisted method for optimizing the operation of a hydrogen-containing building energy system, as proposed in this invention.
[0122] Figure 2 The electrical loss convergence curve of an optimization method for the operation of a hydrogen-containing building energy system assisted by a multi-role large model;
[0123] Figure 3 The heat loss convergence curve of a multi-role large model-assisted optimization method for the operation of hydrogen-containing building energy systems;
[0124] Figure 4 The reward function convergence curve is presented as an optimization method for the operation of a hydrogen-containing building energy system assisted by a multi-role large model. Detailed Implementation
[0125] The present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for illustrating the technical solutions of the present invention more clearly, and should not be construed as limiting the scope of protection of the present invention.
[0126] Example 1:
[0127] like Figure 1 As shown, the design flowchart of a multi-role large model-assisted method for optimizing the operation of a hydrogen-containing building energy system provided by the present invention includes the following steps:
[0128] The established problem of minimizing the operating cost of a hydrogen-containing building multi-energy system under off-grid operation mode includes the objective function, decision variables, and constraints, namely:
[0129] (1) Objective function: The objective function is to minimize the expected time average of both system operating cost and energy loss, which is:
[0130] (1),
[0131] in, For system operating costs, For electrical loss, Heat loss is calculated using the following formula:
[0132] (2),
[0133] (3),
[0134] (4),
[0135] (5),
[0136] (6),
[0137] (7),
[0138] In equation (1), For expectation operators; The total operating cost of the system; in equation (2), , , In time The component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost; in formula (3), , , , , , These are the cost coefficients for photovoltaic modules, electrolyzers, fuel cells, electric boilers, ground source heat pumps, and thermal storage tanks, respectively. , , , , , , In time The output power of photovoltaic cells, the input power of electrolyzers, the output power of fuel cells, the input power of electric boilers, the input power of ground source heat pumps, the heating input power of thermal storage tanks, and the heat release output power of thermal storage tanks; in equation (4), In time The purchase price of hydrogen. In time The volume of hydrogen purchased, in formula (5), This is the cost reduction coefficient for photovoltaics. In time The maximum power generation of photovoltaics; in equation (6), Let be the electrical loss penalty function, in equation (7), This is the heat loss penalty function;
[0139] (2) Decision variables: Decision variables include each time slot Photovoltaic power Electrolytic cell input power Fuel cell output power Electric boiler input power Ground source heat pump input power Thermal storage tank charging power Heat storage tank heat release power and the amount of hydrogen purchased from the hydrogen market ;
[0140] (3) Constraints: Constraints include power balance and equipment physical constraints:
[0141] Photovoltaic constraints:
[0142] (8),
[0143] (9),
[0144] Hydrogen storage tanks and hydrogen balance constraints:
[0145] (10)
[0146] (11),
[0147] (12)
[0148] (13)
[0149] (14)
[0150] (15)
[0151] (16)
[0152] (17)
[0153] Electrolytic cell constraints:
[0154] (18)
[0155] (19)
[0156] (20)
[0157] Fuel cell constraints:
[0158] (twenty one),
[0159] (twenty two),
[0160] (twenty three),
[0161] (twenty four),
[0162] (25)
[0163] Thermal storage tank constraints:
[0164] (26)
[0165] (27)
[0166] (28)
[0167] (29)
[0168] (30)
[0169] Constraints of ground source heat pumps and geothermal wells:
[0170] (31),
[0171] (32),
[0172] (33),
[0173] (34),
[0174] (35),
[0175] (36)
[0176] Electric boiler constraints:
[0177] (37)
[0178] (38),
[0179] System supply and demand balance constraints:
[0180] (39)
[0181] (40)
[0182] In equation (8), The photovoltaic efficiency coefficient. The total area of the photovoltaic panels. for Solar radiation intensity in time slot; in equation (10), For time slots The hydrogen storage capacity of the hydrogen storage tank. This represents the amount of hydrogen produced by the electrolyzer within a given time period. This represents the hydrogen capacity consumed by the fuel cell within a given time period. For time slots The purchased hydrogen capacity, in formula (11) The maximum hydrogen storage capacity of the hydrogen storage tank is given in equation (12). To maximize the amount of hydrogen purchased, in equation (13), It is the gas compressibility factor. , For the fitting parameters, in equation (14), For time slots Hydrogen storage tank pressure, This refers to the volume of the hydrogen storage tank. Let be the ideal gas constant. For time slots Hydrogen storage tank temperature, Let be the molar mass of hydrogen gas, in equation (15), For the safety margin pressure of the hydrogen storage tank, It is a binary variable. To relax the constraints, in equation (16), For the maximum pressure limit of the hydrogen storage tank, in equation (17), Let be the total number of time periods during which the pressure in the hydrogen storage tank exceeds the safety margin, as given in equation (18). The capacity of hydrogen produced by the electrolyzer. The hydrogen production efficiency of the electrolyzer. For time slots Input power of the internal electrolytic cell, For the time interval, in equation (19), For the maximum input power of the electrolytic cell, in equation (20), The ramp rate of the electrolytic cell. They represent the input power of the electrolytic cell in the previous time slot, respectively. In equation (21), The amount of hydrogen consumed by the fuel cell. For time slots The output power of the internal fuel cell, For the hydrogen-to-electricity conversion efficiency of the fuel cell, in equation (22), For the maximum output power of the fuel cell, in equation (23), For the ramp rate of the fuel cell, For time slots The output power of the internal fuel cell is given by equation (25). For time slots The heat generated by the fuel cell The electrothermal conversion ratio of the fuel cell, For the heat recovery coefficient, in equation (26), For time slots The thermal storage capacity of the thermal storage tank For time slots The charging power of the thermal storage tank For time slots The heat dissipation capacity of the thermal storage tank For heat charging efficiency, For heat release efficiency, in equation (27), For the maximum heat storage capacity of the heat storage tank, in equation (28), For the maximum heat charging power, in equation (29), For the maximum heat release power, in equation (31), For time slots The heat output power of a ground source heat pump For time slots The electrical power input to the ground source heat pump, For the efficiency of a ground source heat pump, in equation (32), For the maximum input electrical power of the ground source heat pump, in equation (33), For time slots The heat storage level of geothermal wells For time slots The waste heat from the fuel cell injected into the geothermal well, in equation (34), For the maximum heat storage capacity of the geothermal well, in equation (35), represents the waste heat injected into the geothermal well. No more than the waste heat generated by the fuel cell In equation (37), For time slots The heat generated by the electric boiler For time slots The electrical power input to the electric boiler, For the efficiency of the electric boiler, in equation (38), For the maximum input electrical power of the electric boiler, in equation (39), For time slots electrical load, Time slot The electrical power consumed by the electrolytic cell For time slots The electrical power consumed by a ground source heat pump For time slots The electrical loss, in equation (40), For time slots heat load, For time slots The heat consumed during the charging of the thermal storage tank. For time slots Waste heat injected into geothermal wells For time slots Waste heat power generated by fuel cells For time slots The heat generated by the ground source heat pump For time slots The heat released by the thermal storage tank For time slots Heat loss.
[0183] Furthermore, in step 2, the problem of minimizing operating costs is remodeled as a Markov decision process with a safety correction mechanism related to the operational optimization of multi-energy systems in hydrogen-containing buildings, defined as a tuple. .in, For the state space related to the operation of multi-energy systems in hydrogen-containing buildings, For the original action space, This is the state transition function. For the final reward function, To solve this Markov decision process, a rule-based action correction mechanism is introduced, using the discount factor. As a pre-module for strategy execution. The specific expression is as follows:
[0184] (41),
[0185] (42),
[0186] (43),
[0187] In equation (41), The current running hour; each action component is normalized to the range of [0,1] or [-1,1], in equation (42) , The normalized heat storage tank charging and discharging power; , These are the normalized power allocated by the fuel cell to the electrical load, the ground source heat pump, and the electric boiler, respectively. , The normalized amount of hydrogen purchased; in equation (43), For the final reward function, This is the basic reward function.
[0188] Furthermore, the aforementioned To ensure system supply and demand balance and physical constraints in off-grid mode, a rule-based action correction mechanism is introduced in the secure Markov decision process. Original action Corrected to actual action The correction steps are as follows:
[0189] S41, Hot Can Action Correction, according to Determine the actual heat charging power of the thermal storage tank or heat dissipation power and update the remaining heat load gap. ,when At this time, it is the heat input mode of the hot tank. ,and ,when At this time, it is the heat output mode of the hot tank. ,and ,in, For maximum charge / discharge power, At maximum capacity, the heat load gap is updated to... .
[0190] S42. Photovoltaic Consumption and Distribution Correction: Photovoltaic power generation prioritizes meeting the basic electrical load, with remaining photovoltaic power subsequently used to drive ground-source heat pumps, electric boilers, and electrolytic cells. Specifically, the photovoltaic power required to meet the basic electrical load is first calculated. Secondly, calculate the power of the photovoltaic-driven ground source heat pump. Next, calculate the power of the photovoltaic-driven electric boiler. And update the heat load gap to , , The output power of the photovoltaic system is allocated to the power of the ground source heat pump and the electric boiler, respectively.
[0191] S43. Fuel Cell Operation Correction: When photovoltaic output is insufficient, the fuel cell is used to supplement power and heat supply. The correction process must simultaneously meet the physical constraints of the original operation command, the maximum power of the equipment, and the current hydrogen storage capacity. Specifically, first, the power of the fuel cell to meet the remaining electrical load is calculated, i.e. Secondly, calculate the power of the fuel cell driving the ground source heat pump:
[0192] ;in, Finally, to calculate the power of the fuel cell-driven electric boiler, considering the heat generated by the fuel cell. The total output power of the fuel cell was ultimately determined to be... .
[0193] S44. Electrolyzer Operation Correction: The electrolyzer utilizes surplus photovoltaic power to produce hydrogen and operates only when the fuel cell is not working, to avoid energy conversion conflicts and is limited by the hydrogen storage tank capacity. Specifically, if the total output power of the fuel cell... Then the power of the electrolytic cell ;like
[0194] Then the power of the electrolytic cell ; Calculate the hydrogen storage capacity in the intermediate state based on the operation of the electrolyzer and fuel cell. .
[0195] S45. Hydrogen purchase action correction: Determine whether to purchase hydrogen from the market based on the hydrogen storage threshold. Specifically, if the intermediate state hydrogen storage... The actual purchase quantity ,like The actual purchase quantity ,in, This is the safe threshold for hydrogen storage capacity.
[0196] Furthermore, in equation (43), the final reward function Includes the basic loss function Electrical loss penalty function Heat loss penalty function The electrical loss penalty function and the thermal loss penalty function both adopt a three-segment piecewise linear structure. The specific forms of the basic loss function, electrical loss penalty function, and thermal loss penalty function are as follows:
[0197] (44),
[0198] (45)
[0199] (46)
[0200] In equation (44), , , In time The calculation formulas for the component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost are shown in equations (3), (4), and (5), respectively; in equation (45), Here is a three-segment electrical loss penalty function, where, These are the penalty functions for electrical losses in the first, second, and third segments, respectively. The current power loss values are as follows: First segment: If ,but Second paragraph: If Third paragraph: If ,but ,in, This represents the severe threshold of electrical load loss; This represents the moderate threshold of electrical load loss. , The preset electrical load loss threshold coefficient, The value is the penalty coefficient, and it satisfies the monotonicity constraint. In equation (46), Here is a three-segment heat loss penalty function, where, These are the heat loss penalty functions for the first, second, and third segments, respectively. The current heat loss values are as follows: First segment: If ,but Second paragraph: If ,but Third paragraph: If ,but ,in The threshold indicating the severity of heat load loss; This represents the intermediate threshold of heat load loss. , This is the preset heat load loss threshold coefficient; Let be the heat loss penalty coefficient, and satisfy the monotonicity constraint, i.e.: ;
[0201] Furthermore, in step 3, the secure Markov decision process modeled using a multi-role large language model-assisted proximal policy optimization algorithm is solved by alternating training and parameter tuning phases until the penalty function parameters reach a predefined convergence criterion. Specifically, this includes:
[0202] S61. Training Phase: The agent interacts with the environment using the current penalty function parameters, specifically including:
[0203] 1) Initialize the policy network and value network ;
[0204] 2) The agent interacts with the environment to collect experience data and calculates the reward value based on the three-stage penalty function;
[0205] 3) Calculate the advantage function based on collected experience. and rewards ;
[0206] 4) Use the combined loss function Update network parameters:
[0207] (47)
[0208] In equation (47), Let the loss function be the shearing strategy. The loss function of the value network, Let the policy entropy reward function be... and This is the balance coefficient;
[0209] S62. Parameter Tuning Phase: After the trained model is evaluated on the test set, the penalty parameters are optimized using a multi-role large language model, specifically including:
[0210] 1) In The performance of the current strategy is evaluated on the test set, and the total operating cost and electrical and heat loss indicators are recorded.
[0211] 2) The performance of the large language model for evaluating roles is qualitatively assessed based on fuzzy interval criteria;
[0212] 3) Parameter tuning: Based on the evaluation results and guided thought chain suggestions, the large language model generates a new three-stage penalty function slope parameter (…). as well as );
[0213] S63. Repeat the above stages, clearing the agent's interaction experience after each round of training and starting training again.
[0214] Furthermore, the large language model for the evaluation role is primarily responsible for:
[0215] (1) Analysis The cumulative test results over the days are used to break down the total operating cost, hydrogen purchase cost, component operating cost, and photovoltaic abandonment cost.
[0216] (2) Based on the preset fuzzy interval evaluation criteria, the electrical loss and thermal loss performance are mapped to multi-level qualitative evaluation labels;
[0217] (3) Analyze the composition of operating costs, identify the main cost drivers, and output structured results including loss assessment and performance analysis.
[0218] Furthermore, the prompt words of the large language model system for parameter tuning roles include six core components:
[0219] (1) Role description: Defined as an expert in power system operation and reinforcement learning, capable of extracting relevant domain knowledge from embedded knowledge;
[0220] (2) Environmental Description: Provide detailed specifications for the integrated multi-energy system, including:
[0221] ①Core components: photovoltaic power generation, electrolyzer, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, thermal storage tank and geothermal well;
[0222] ②Dynamic characteristics: electrical balance, thermal balance, hydrogen management, energy storage operation;
[0223] ③ Operational constraints: efficiency factor, capacity limit, tiered load loss penalty;
[0224] (3) Task description: Outline the main optimization objectives, including ensuring that electrical and heat losses do not exceed specified thresholds, and minimizing total operating costs while maintaining load satisfaction within acceptable limits;
[0225] (4) Output format: Strictly define a fixed output structure, requiring only fixed objects to be output without any additional explanation or code;
[0226] (5) Parameter constraints: enforce mathematical monotonicity conditions as well as And maintain a coordinated relationship where thermal penalty usually exceeds electrical penalty;
[0227] (6) Reward calculation framework: specify the composite reward structure .
[0228] Furthermore, S62 includes a guided chain of thought processes, which enhances the logical consistency of parameter adjustments through multi-stage analysis:
[0229] S91. When the assessment results of electrical or thermal losses do not reach the preset satisfaction level, and the total operating cost exceeds the preset threshold... At this time, a penalty enhancement strategy or a fine-tuning strategy is adopted to increase the slope of the first segment of the corresponding loss function. Or multiple slope parameters;
[0230] S92. When the assessment results of electrical loss or heat loss reach the preset satisfaction level, but the total operating cost exceeds the preset threshold... At that time, a penalty relaxation strategy or parameter optimization strategy can be adopted to reduce the slope parameter of the corresponding loss function to improve economic efficiency;
[0231] S93. When there is a difference between the assessment results of electrical loss and heat loss, a multi-problem collaborative processing strategy is adopted, prioritizing the adjustment of the penalty parameter of the side with the worse assessment result, or a focus strategy is adopted to keep one side stable while optimizing the other side.
[0232] S94. During the adjustment process, always maintain the monotonicity constraint of the parameters, and determine the magnitude of parameter adjustment based on the training reward value.
[0233] Furthermore, in step 4, when the agent makes online decisions, it only needs to call the trained policy network, that is:
[0234] (48)
[0235] In equation (48), For the trained policy network, The current state observation value, For the generated optimized actions.
[0236] To verify the effectiveness and advancement of the multi-role large model-assisted optimization method for hydrogen-containing building energy systems proposed in this invention, three sets of comparative schemes were set up.
[0237] Comparison with Option 1: This option employs a heuristic operation strategy based on engineering experience. The core strategy prioritizes meeting electricity and heat load demands, adhering to fixed equipment start-up and shutdown rules and energy management regulations. Specifically, photovoltaic power generation is first used to meet electricity load demands, with surplus power allocated to the ground source heat pump, electric boiler, and electrolyzer according to a fixed priority order. Hydrogen storage tanks and thermal storage tanks are controlled based on preset charge / discharge thresholds, while fuel cells and electrolyzers operate on a mutually exclusive principle to avoid energy conversion losses.
[0238] Comparison with Scheme 2: This scheme uses the standard near-end strategy optimization algorithm, but its reward function includes penalties for electrical and thermal losses. as well as For fixed values, LLM (Limited Least Mean Time) was not introduced for closed-loop tuning. In specific implementation, the agent is trained based on the same state space and action space, and the reward function adopts the same mathematical form as in this invention, but all penalty parameters are determined before training begins and remain unchanged throughout the training process. By comparing with this baseline, the specific contribution of the multi-role large model-assisted optimization method for hydrogen-containing building energy systems proposed in this paper to the final performance improvement can be clearly separated and quantified.
[0239] Comparison Scheme 3: This scheme uses the commercial software GAMS in conjunction with the CPLEX solver to construct the system scheduling problem as a mixed-integer linear programming model for direct solution. In this scheme, it is assumed that uncertain parameters such as solar radiation intensity, electrical and thermal load demand, and hydrogen price for all future periods can be perfectly predicted. The performance of this scheme provides the performance upper limit for the method of this invention.
[0240] Figure 2 and Figure 3 The convergence process of electrical loss and thermal loss during training is shown separately. Figure 4 The reward convergence curve during training is shown. The graphical results demonstrate the convergence of the multi-role large language model-assisted reinforcement learning process.
[0241] Table 1 - Comparison of Costs and Safety of the Method of this Invention and Other Solutions
[0242]
[0243] Table 1 shows a comparison of the experimental results of the method of the present invention and the three comparative schemes mentioned above. Compared with comparative scheme one and comparative scheme two, the method of the present invention has lower operating costs and heat load loss, and the electrical loss is 0. Specifically, the method of the present invention can reduce operating costs by 89.04% and 33.76%, and reduce heat loss by 98.16% and 96.98%, respectively. Moreover, compared with comparative scheme three under perfect prediction information, the relative error of the operating cost of the method of the present invention is 4.04%, and the absolute value of heat loss is 599.32 kWh (equivalent to a cutoff of 0.8 kWh per hour), indicating that the method of the present invention has excellent near-optimal performance.
Claims
1. A method for optimizing the operation of a hydrogen-containing building energy system using a multi-role large model, characterized in that, Includes the following steps: Step 1: Establish the problem of minimizing the operating cost and energy loss of hydrogen-containing building multi-energy systems under off-grid operation mode; Step 2: Remodel the operating cost minimization problem as a Markov decision process with a safety correction mechanism related to the operation optimization of multi-energy systems in hydrogen-containing buildings, defined as tuples. ,in, For the state space related to the operation of multi-energy systems in hydrogen-containing buildings, For the original action space, This is the state transition function. For the final reward function, A rule-based action correction mechanism is introduced as the discount factor. As a pre-module for strategy execution; Final reward function Includes the basic loss function Electrical loss penalty function Heat loss penalty function Among them, both the electrical loss penalty function and the heat loss penalty function adopt a three-segment piecewise linear structure; Step 3: Solve the modeled safe Markov decision process based on the multi-role large language model-assisted proximal policy optimization algorithm to obtain the agent operation strategy related to the hydrogen-containing building multi-energy system. The training phase and parameter adjustment phase are alternately iterated until the penalty function parameter reaches the predefined convergence criterion. Parameter tuning phase: After the trained model is evaluated on the test set, the penalty parameters are optimized using a multi-role large language model, specifically including: 1) In The performance of the current strategy is evaluated on the test set, and the total operating cost and electrical and heat loss indicators are recorded. 2) The performance of the large language model for evaluating roles is qualitatively assessed based on fuzzy interval criteria; 3) Parameter tuning: Based on the evaluation results and guided thought chain suggestions, the large language model generates a new three-stage penalty function slope parameter. as well as ; The large language model responsible for evaluating roles is: (1) Analysis The cumulative test results over the days are used to break down the total operating cost, hydrogen purchase cost, component operating cost, and photovoltaic abandonment cost. (2) Based on the preset fuzzy interval evaluation criteria, the electrical loss and thermal loss performance are mapped to multi-level qualitative evaluation labels; (3) Analyze the composition of operating costs, identify cost drivers, and output structured results including loss assessment and performance analysis; Step 4: The intelligent agent makes online decisions based on the hydrogen-containing building multi-energy system operation strategy obtained through training, and applies the decisions to the actual hydrogen-containing building multi-energy system.
2. The method according to claim 1, characterized in that, The problem of minimizing the operating cost of a hydrogen-containing building multi-energy system under off-grid operation mode established in Step 1 includes the objective function, decision variables, and constraints, as follows: (1) Objective function: The objective function is to minimize the expected time average of both system operating cost and energy loss, which is: (1), in, The total operating cost of the system, For electrical loss, Heat loss is calculated using the following formula: (2), (3), (4), (5), (6), (7), In equation (1), For expectation operators; The total operating cost of the system; in equation (2), , , In time The component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost; in formula (3), , , , , , These are the cost coefficients for photovoltaic modules, electrolyzers, fuel cells, electric boilers, ground source heat pumps, and thermal storage tanks, respectively. , , , , , , In time Photovoltaic output power and time Input power and time of the electrolytic cell Fuel cell output power and time Input power and time of electric boiler Input power and time of ground source heat pump Heating input power and time of the thermal storage tank The heat dissipation output power of the thermal storage tank; in equation (4), In time The purchase price of hydrogen. In time The volume of hydrogen purchased, in formula (5), This is the cost reduction coefficient for photovoltaics. In time The maximum power generation of photovoltaics; in equation (6), Let be the electrical loss penalty function, in equation (7), This is the heat loss penalty function; (2) Decision variables: Decision variables include time Photovoltaic output power ,time Electrolytic cell input power ,time fuel cell output power ,time Electric boiler input power ,time Ground source heat pump input power ,time Thermal storage tank heating input power ,time Thermal storage tank heat output power and in time Purchased hydrogen volume ; (3) Constraints: Constraints include power balance and equipment physical constraints: Photovoltaic constraints: (8), (9), Hydrogen storage tanks and hydrogen balance constraints: (10), (11), (12), (13), (14), (15), (16), (17), Electrolytic cell constraints: (18), (19), (20), Fuel cell constraints: (21), (22), (23), (24), (25), Thermal storage tank constraints: (26), (27), (28), (29), (30), Constraints of ground source heat pumps and geothermal wells: (31), (32), (33), (34), (35), (36), Electric boiler constraints: (37), (38), System supply and demand balance constraints: (39), (40), In equation (8), The photovoltaic efficiency coefficient. The total area of the photovoltaic panels. for Solar radiation intensity in time slot; in equation (10), For time slots The hydrogen storage capacity of the hydrogen storage tank. This represents the amount of hydrogen produced by the electrolyzer within a given time period. This represents the hydrogen capacity consumed by the fuel cell within a given time period. For time slots The volume of hydrogen purchased, in formula (11) The maximum hydrogen storage capacity of the hydrogen storage tank is given in equation (12). To maximize the amount of hydrogen purchased, in equation (13), The gas compressibility factor, , For the fitting parameters, in equation (14), For time slots Hydrogen storage tank pressure, This refers to the volume of the hydrogen storage tank. Let be the ideal gas constant. For time slots Hydrogen storage tank temperature, Let be the molar mass of hydrogen gas, in equation (15), For the safety margin pressure of the hydrogen storage tank, It is a binary variable. To relax the constraints, in equation (16), For the maximum pressure limit of the hydrogen storage tank, in equation (17), Let be the total number of time periods during which the pressure in the hydrogen storage tank exceeds the safety margin, as given in equation (18). The capacity of hydrogen produced by the electrolyzer. The hydrogen production efficiency of the electrolyzer. For time slots Input power of the internal electrolytic cell, For the time interval, in equation (19), For the maximum input power of the electrolytic cell, in equation (20), The ramp rate of the electrolytic cell. They represent the input power of the electrolytic cell in the previous time slot, respectively. In equation (21), The amount of hydrogen consumed by the fuel cell. For time slots The output power of the internal fuel cell, For the hydrogen-to-electricity conversion efficiency of the fuel cell, in equation (22), For the maximum output power of the fuel cell, in equation (23), For the ramp rate of the fuel cell, Indicates time slot The output electrical power of the internal fuel cell is given by equation (25). For time slots Waste heat generated by fuel cells The electrothermal conversion ratio of the fuel cell, For the heat recovery coefficient, in equation (26), For time slots The thermal storage capacity of the thermal storage tank For time slots The charging power of the thermal storage tank For time slots The heat dissipation capacity of the thermal storage tank For heat charging efficiency, For heat release efficiency, in equation (27), For the maximum heat storage capacity of the heat storage tank, in equation (28), For the maximum heat charging power, in equation (29), For the maximum heat release power, in equation (31), For time slots The heat output power of a ground source heat pump For time slots The electrical power input to the ground source heat pump, For the efficiency of a ground source heat pump, in equation (32), For the maximum input electrical power of the ground source heat pump, in equation (33), For time slots The thermal storage level of geothermal wells For time slots The waste heat from the fuel cell injected into the geothermal well, in equation (34), For the maximum heat storage capacity of the geothermal well, in equation (35), represents the waste heat injected into the geothermal well. No more than the waste heat generated by the fuel cell In equation (37), For time slots The heat generated by the electric boiler For time slots The electrical power input to the electric boiler, For the efficiency of the electric boiler, in equation (38), For the maximum input electrical power of the electric boiler, in equation (39), For time slots electrical load, The electrical power consumed by the electrolytic cell in time slot t. For time slots The electrical power input to the ground source heat pump, For time slots Fuel cell output power, For time slots The electrical loss, in equation (40), For time slots heat load, For time slots The heat consumed during the charging of the thermal storage tank. For time slots Waste heat injected into geothermal wells For time slots Waste heat generated by fuel cells For time slots The heat output power of a ground source heat pump For time slots The heat released by the thermal storage tank For time slots Heat loss.
3. The method according to claim 2, characterized in that, In step 2, the problem of minimizing operating costs is remodeled as a Markov decision process with a safety correction mechanism related to the operation optimization of multi-energy systems in hydrogen-containing buildings, defined as a tuple. ,in, For the state space related to the operation of multi-energy systems in hydrogen-containing buildings, For the original action space, This is the state transition function. For the final reward function, A rule-based action correction mechanism is introduced as the discount factor. As a prerequisite module for strategy execution, the specific expression is as follows: (41), (42), (43), In equation (41), The current running hour; each action component is normalized to the range of [0,1] or [-1,1], in equation (42) , The normalized heat storage tank charging and discharging power; , These are the normalized power allocated by the fuel cell to the electrical load, the ground source heat pump, and the electric boiler, respectively. , The normalized amount of hydrogen purchased; in equation (43), For the final reward function, This is the basic reward function.
4. The method according to claim 3, characterized in that, Introducing a rule-based action correction mechanism Original action Corrected to actual action The correction steps are as follows: S41, Hot Can Action Correction, according to Determine the actual heat charging power of the thermal storage tank or heat dissipation power and update the remaining heat load gap. ,when At this time, it is the heat input mode of the hot tank. ,and ,when At this time, it is the heat output mode of the hot tank. ,and ,in, For maximum charge / discharge power, At maximum capacity, the heat load gap is updated to... ; S42. Photovoltaic Consumption and Distribution Correction: Photovoltaic power generation prioritizes meeting the basic electrical load. Remaining photovoltaic power is then used to drive ground-source heat pumps, electric boilers, and electrolytic cells. Specifically, the power generated by photovoltaics to meet the basic electrical load is first calculated. Secondly, calculate the power of the photovoltaic-driven ground source heat pump. ; Next, calculate the power of the photovoltaic-driven electric boiler. And update the heat load gap to , , The output power of the photovoltaic system is allocated to the power of the ground source heat pump and the electric boiler, respectively. S43. Fuel Cell Operation Correction: When photovoltaic output is insufficient, the fuel cell is used to supplement power and heat supply. The correction process must simultaneously meet the physical constraints of the original operation command, the maximum power of the equipment, and the current hydrogen storage capacity. Specifically, first, the power of the fuel cell to meet the remaining electrical load is calculated, i.e. Secondly, calculate the power of the fuel cell driving the ground source heat pump. Finally, the power of the fuel cell-driven electric boiler was calculated. ,in, The heat generated by the fuel cell; the total output power of the fuel cell was ultimately determined to be... ; S44. Electrolyzer operation correction: The electrolyzer utilizes surplus photovoltaic power to produce hydrogen and operates only when the fuel cell is not working. Specifically, if the total output power of the fuel cell... Then the power of the electrolytic cell ;like Then the power of the electrolytic cell ;in After allocating the photovoltaic output power to the ground source heat pump and electric boiler, the remaining power is used to calculate the hydrogen storage capacity in the intermediate state based on the operation of the electrolyzer and fuel cell. ; S45. Hydrogen purchase action correction: Determine whether to purchase hydrogen from the market based on the hydrogen storage threshold. Specifically, if the intermediate hydrogen storage level... The actual purchase quantity ,like The actual purchase quantity ,in, This is the safe threshold for hydrogen storage capacity.
5. The method according to claim 4, characterized in that, In equation (43), the final reward function Includes the basic loss function Electrical loss penalty function Heat loss penalty function The electrical loss penalty function and the thermal loss penalty function both adopt a three-segment piecewise linear structure. The specific forms of the basic loss function, electrical loss penalty function, and thermal loss penalty function are as follows: (44), (45), (46), In equation (44), , , In time The calculation formulas for the component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost are shown in equations (3), (4), and (5), respectively; in equation (45), Here is a three-segment electrical loss penalty function, where, The electrical loss penalty functions are for the first, second, and third segments, respectively, and are as follows: First segment: If ,but Second paragraph: If Third paragraph: If ,but ,in, This represents the severe threshold of electrical load loss; This represents the moderate threshold of electrical load loss. , The preset electrical load loss threshold coefficient, The value is the penalty coefficient, and it satisfies the monotonicity constraint. In equation (46), Here is a three-segment heat loss penalty function, where, These are the heat loss penalty functions for the first, second, and third segments, respectively, specifically: First segment: If... ,but Second paragraph: If ,but Third paragraph: If ,but ,in The threshold indicating the severity of heat load loss; This represents the intermediate threshold of heat load loss. , This is the preset heat load loss threshold coefficient; Let be the heat loss penalty coefficient, and satisfy the monotonicity constraint, i.e.: .
6. The method according to claim 5, characterized in that, Step 3, which uses a multi-role large language model-assisted proximal policy optimization algorithm to solve the modeled secure Markov decision process, employs alternating iterative training and parameter tuning phases until the penalty function parameters reach a predefined convergence criterion. Specifically, this includes: S61. Training Phase: The agent interacts with the environment using the current penalty function parameters, specifically including: 1) Initialize the policy network and value network ; 2) The agent interacts with the environment to collect experience data and calculates the reward value based on the three-stage penalty function; 3) Calculate the advantage function based on collected experience. and rewards ; 4) Use the combined loss function Update network parameters: (47), In equation (47), Let the loss function be the shearing strategy. The loss function of the value network, Let the policy entropy reward function be... and This is the balance coefficient; S62. Repeat the above stages, clearing the agent's interaction experience after each round of training and starting training again.
7. The method according to claim 6, characterized in that, The large language model system prompt words for parameter tuning roles include six core components: (1) Role description: Defined as an expert in power system operation and reinforcement learning; (2) Environmental Description: Provide detailed specifications for the integrated multi-energy system, including: Core components: photovoltaic power generation, electrolyzer, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, thermal storage tank and geothermal well; Dynamic characteristics: electrical balance, thermal balance, hydrogen management, and energy storage operation; Operational constraints: efficiency factor, capacity limit, and tiered load loss penalty; (3) Task description: Outline the optimization objectives, including ensuring that electrical and heat losses do not exceed specified thresholds, and minimizing total operating costs while maintaining load satisfaction within acceptable limits; (4) Output format: Strictly define a fixed output structure, requiring only fixed objects to be output without any additional explanation or code; (5) Parameter constraints: enforce mathematical monotonicity conditions as well as And maintain a coordinated relationship where thermal penalty exceeds electrical penalty; (6) Reward calculation framework: specifying the composite reward structure .
8. The method according to claim 1, characterized in that, The parameter adjustment phase includes a guided chain of thought, enhancing the logical consistency of the parameter adjustment steps through multi-stage analysis: S91. When the assessment results of electrical or thermal losses do not reach the preset satisfaction level, and the total operating cost exceeds the preset threshold... At this time, a penalty enhancement strategy or a fine-tuning strategy is adopted to increase the slope of the first segment of the corresponding loss function. Or multiple slope parameters; S92. When the assessment results of electrical loss or heat loss reach the preset satisfaction level, but the total operating cost exceeds the preset threshold... At that time, a penalty relaxation strategy or parameter optimization strategy is adopted to reduce the slope parameter of the corresponding loss function; S93. When there is a difference between the assessment results of electrical loss and heat loss, a multi-problem collaborative processing strategy is adopted, prioritizing the adjustment of the penalty parameter of the side with the worse assessment result, or a focus strategy is adopted to keep one side stable while optimizing the other side. S94. During the adjustment process, always maintain the monotonicity constraint of the parameters, and determine the magnitude of parameter adjustment based on the training reward value.
9. The method according to claim 1, characterized in that, In step 4, when the agent makes online decisions, it only needs to call the trained policy network, that is: (48), In equation (48), For the trained policy network, The current state observation value, For the generated optimized actions.
Citation Information
Patent Citations
Narrow space mechanical arm path planning method based on large language model
CN120755879A
Distributed energy storage charging multifunctional cooperative control method and system
CN120879718A