Method for optimizing operation of hydrogen-containing building energy system assisted by multi-role large model

By employing a multi-role large model-assisted safe Markov decision process and a proximal policy optimization algorithm, this study addresses the problems of relying on expert experience and difficulty in adjusting reward function parameters in existing optimization methods for hydrogen-containing building energy systems. It achieves coordinated optimization of the system's economy and power supply/heating reliability in complex environments.

CN121544087AActive Publication Date: 2026-02-17NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202610086039.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-02-17
Estimated Expiration
2046-01-22

AI Technical Summary

Technical Problem

Existing optimization methods for hydrogen-containing building energy systems suffer from several drawbacks: reliance on expert experience makes them difficult to adapt to complex environments, high computational complexity, and reliance on manual parameter tuning for reward function design. These issues make it difficult to achieve synergistic optimization of system economy and power and heating reliability.

Method used

A safe Markov decision process assisted by a multi-role large model is adopted, combined with a proximal policy optimization algorithm. By establishing an objective function and constraints, introducing a rule-based action correction mechanism, and optimizing the reward function design, the system can realize online decision-making and operation strategies.

Benefits of technology

It significantly improves the system's adaptability and economy under varying operating conditions, solves the problem of difficult parameter tuning of reward function in traditional methods, and achieves multi-objective balance and power supply and heating reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544087A_ABST
    Figure CN121544087A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-role large model assisted hydrogen-containing building energy system operation optimization method, and belongs to the technical field of building energy system optimization control, and the method comprises the steps: firstly, building a hydrogen-containing building multi-energy system operation cost minimization problem in an off-grid operation mode; secondly, re-modeling the problem into a security Markov decision process, and defining a system state space, an action space and a composite reward function; then, solving a safety Markov decision process of modeling based on a multi-role large language model assisted near-end strategy optimization algorithm, and obtaining an intelligent agent operation strategy related to the hydrogen-containing building multi-energy system; finally, the intelligent agent makes an online decision based on the obtained optimization strategy, the decision acts on the actual hydrogen-containing building multi-energy system, the system operation cost can be effectively reduced, and the energy supply reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of building energy system optimization and control technology, specifically a multi-role large model-assisted method for optimizing the operation of hydrogen-containing building energy systems. Background Technology

[0002] Research on the operational optimization of hydrogen-containing building energy systems helps reduce system carbon emissions and energy costs while improving user thermal comfort. Existing research has proposed various methods for optimizing the operation of hydrogen-containing building energy systems, mainly including heuristic rule-based control methods, traditional optimization methods based on mathematical programming, and intelligent decision-making methods based on deep reinforcement learning. While these methods have achieved some success, they each have significant limitations. Heuristic rule-based methods, although simple to implement, heavily rely on expert experience for control performance, making them difficult to adapt to complex and changing operating environments. Mathematical programming methods, while able to obtain theoretically optimal solutions under ideal conditions, require accurate system models and complete predictive information, and have high computational complexity, making online application difficult. Existing deep reinforcement learning-based methods, while avoiding reliance on accurate models, rely on manual parameter tuning for reward function design, making it difficult to adaptively balance multi-objective optimization needs.

[0003] In summary, existing methods for optimizing the operation of multi-energy systems in hydrogen-containing buildings have significant shortcomings. There is an urgent need to research a new operation optimization method to reduce the operating costs of hydrogen-containing building energy systems while ensuring high reliability of power and heat supply. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a multi-role large model-assisted method for optimizing the operation of hydrogen-containing building energy systems. Its purpose is to achieve synergistic optimization of the operational economy and power supply / heating reliability of multi-energy systems in hydrogen-containing buildings. This method overcomes the shortcomings of existing rule-based control methods that rely on expert experience and are difficult to adapt to complex operating environments. It also solves the technical problems of existing deep reinforcement learning-based methods, which require manual tuning of penalty parameters, resulting in slow convergence speed and difficulty in balancing multiple objectives.

[0005] To achieve the above objectives, the present invention employs the following technical solution:

[0006] This invention provides a multi-role large-scale model-assisted method for optimizing the operation of hydrogen-containing building energy systems, including:

[0007] (1) To address the issue of minimizing the operating cost of hydrogen-containing multi-energy systems in off-grid operation mode;

[0008] (2) The problem of minimizing operating costs is remodeled as a safe Markov decision process related to the operation optimization of multi-energy systems in hydrogen-containing buildings;

[0009] (3) Based on the multi-role large language model, the safety Markov decision process of modeling is solved by the near-end strategy optimization algorithm, and the agent operation strategy related to the hydrogen-containing building multi-energy system is obtained;

[0010] (4) The agent makes online decisions based on the operation strategy of the hydrogen-containing building multi-energy system obtained by training, and applies the decision to the actual hydrogen-containing building multi-energy system.

[0011] Further, the established off-grid operation mode hydrogen-containing building multi-energy system operation cost minimization problem includes objective function, decision variable and constraint condition, that is:

[0012] (1) Objective function: simultaneously minimize the time average expected value of system operation cost and energy supply loss, which is:

[0013] (1),

[0014] Among them, is the system operation cost, is the electric loss, is the heat loss, which is calculated by the following formula:

[0015] (2),

[0016] (3),

[0017] (4),

[0018] (5),

[0019] (6),

[0020] (7),

[0021] Among them, in formula (1), is the expectation operator; is the total operation cost of the system; in formula (2), , , are the component operation cost, hydrogen purchase cost and photovoltaic disposal cost at time ; in formula (3), , , , , , are the cost coefficients of photovoltaic components, electrolytic cells, fuel cells, electric boilers, ground source heat pumps and heat storage tanks, respectively, , , , , , , is the output power of the photovoltaic at time is the input power of the electrolyzer at time is the output power of the fuel cell at time is the input power of the electric boiler at time is the input power of the ground source heat pump at time is the heating input power of the thermal storage tank at time is the discharging output power of the thermal storage tank; in equation (4), is the hydrogen purchase price at time , is the volume of hydrogen purchased at time , in equation (5), is the photovoltaic curtailment cost coefficient, is the maximum power generation of the photovoltaic at time ; in equation (6), is the electrical loss penalty function, is the thermal loss penalty function; in equation (7),

[0022] (2) Decision variables: the decision variables include the photovoltaic usage power at each time slot , the electrolyzer input power at time , the fuel cell output electrical power at time , the electric boiler input electrical power at time , the ground source heat pump input electrical power at time , the thermal storage tank heating input power at time , the thermal storage tank discharging output power at time , and the amount of hydrogen purchased from the hydrogen market ;

[0023] (3) Constraint conditions: the constraint conditions include power balance and equipment physical constraints:

[0024] Photovoltaic constraint:

[0025] (8),

[0026] (9),

[0027] Hydrogen storage tank and hydrogen balance constraint:

[0028] (10),

[0029] (11),

[0030] (12),

[0031] (13),

[0032] (14),

[0033] (15),

[0034] (16),

[0035] (17),

[0036] Electrolyzer constraints:

[0037] (18),

[0038] (19),

[0039] (20),

[0040] Fuel cell constraints:

[0041] (21),

[0042] (22),

[0043] (23),

[0044] (24),

[0045] (25),

[0046] Thermal storage tank constraints:

[0047] (26),

[0048] (27),

[0049] (28),

[0050] (29),

[0051] (30)

[0052] Constraints of ground source heat pumps and geothermal wells:

[0053] (31),

[0054] (32),

[0055] (33),

[0056] (34),

[0057] (35),

[0058] (36)

[0059] Electric boiler constraints:

[0060] (37)

[0061] (38)

[0062] System supply and demand balance constraints:

[0063] (39)

[0064] (40),

[0065] In equation (8), The photovoltaic efficiency coefficient. The total area of ​​the photovoltaic panels. for Solar radiation intensity in time slot; in equation (10), For time slots The hydrogen storage capacity of the hydrogen storage tank. This represents the amount of hydrogen produced by the electrolyzer within a given time period. This represents the hydrogen capacity consumed by the fuel cell within a given time period. For time slots The purchased hydrogen capacity, in formula (11) The maximum hydrogen storage capacity of the hydrogen storage tank is given in equation (12). To maximize the amount of hydrogen purchased, in equation (13), The gas compressibility factor, , For the fitting parameters, in equation (14), For time slots Hydrogen storage tank pressure, This refers to the volume of the hydrogen storage tank. R is the ideal gas constant, t is the time slot T is the hydrogen storage tank temperature, M is the molar mass of hydrogen, in equation (15), P is the hydrogen storage tank safety margin pressure, B is a binary variable, R is the relaxation constraint, in equation (16), P is the hydrogen storage tank ultimate maximum pressure, in equation (17), N is the total number of time periods in which the hydrogen storage tank pressure exceeds the safety margin, in equation (18), H is the hydrogen production capacity of the electrolyzer, E is the hydrogen production efficiency of the electrolyzer, t is the time slot P is the input power to the internal electrolyzer, T is the time interval, in equation (19), P is the maximum input power to the electrolyzer, in equation (20), R is the ramp rate of the electrolyzer, P is the input power to the electrolyzer in the previous time slot, in equation (21), H is the amount of hydrogen consumed by the fuel cell, t is the time slot P is the output electrical power of the fuel cell, E is the hydrogen-to-electricity efficiency of the fuel cell, in equation (22), P is the maximum output power of the fuel cell, in equation (23), R is the ramp rate of the fuel cell, t is the time slot P is the output electrical power of the fuel cell at time t, in equation (25), t is the time slot Q is the waste heat generated by the fuel cell, E is the electricity-to-heat conversion ratio of the fuel cell, R is the heat recovery coefficient, in equation (26), t is the time slot H is the heat storage level of the thermal storage tank, t is the time slot P is the charging power of the thermal storage tank, t is the time slot P is the discharging power of the thermal storage tank, E is the charging efficiency, E is the discharging efficiency, in equation (27), H is the maximum heat storage capacity of the thermal storage tank, in equation (28), P is the maximum charging power, in equation (29), P is the maximum discharging power, in equation (31), t is the time slot Heat power output by the ground source heat pump, is the time slot Electric power input by the ground source heat pump, is the ground source heat pump efficiency, in equation (32), is the maximum electric power input by the ground source heat pump, in equation (33), is the time slot Heat storage level of the geothermal well, is the time slot Fuel cell waste heat injected into the geothermal well, in equation (34), is the maximum heat storage capacity of the geothermal well, in equation (35), indicating the waste heat injected into the geothermal well No more than the waste heat generated by the fuel cell , in equation (37), is the time slot Heat generated by the electric boiler, is the time slot Electric power input by the electric boiler, is the electric boiler efficiency, in equation (38), is the maximum electric power input by the electric boiler, in equation (39), is the time slot Electric load, is the time slot Electric power consumed by the electrolyzer, is the time slot Electric power consumed by the ground source heat pump, is the time slot Electric loss, in equation (40), is the time slot Thermal load, is the time slot Heat consumed by the heat storage tank to charge, is the time slot Waste heat injected into the geothermal well, is the time slot Waste heat generated by the fuel cell, is the time slot Heat generated by the ground source heat pump, is the time slot Heat provided by the heat storage tank to discharge, is the time slot Thermal loss.

[0066] Further, the operation cost minimization problem is re-modeled as a Markov decision process with safety correction mechanism related to the operation optimization of the hydrogen-containing building multi-energy system, defined as a tuple . Wherein, is the state space related to the operation of the hydrogen-containing building multi-energy system, is the original action space, is the state transition function, is the final reward function, is the discount factor, in order to solve the Markov decision process, a rule-based action modification mechanism is introduced as a pre-module of policy execution. The specific expression is as follows:

[0067] (41),

[0068] (42),

[0069] (43),

[0070] wherein, in formula (41), is the current operating hours; each action component is normalized to the range of [0, 1] or [-1, 1], and in formula (42) is the normalized heat charging and discharging power of the heat storage tank; , are the normalized powers of the fuel cell allocated to the electrical load, the ground source heat pump and the electric boiler respectively; , is the normalized hydrogen purchase amount; in formula (43), is the final reward function, is the basic reward function.

[0071] Further, the is a rule-based action modification mechanism in the safe Markov decision process, in order to ensure the balance between supply and demand of the system in the off-grid mode and the physical constraints, a rule-based action modification mechanism is introduced the original action is modified to the actual execution action , and the modification steps are as follows:

[0072] S41, heat tank action modification, according to determine the actual heat charging power or heat discharging power of the heat storage tank, and update the remaining heat load gap , when , it is the heat input mode of the heat tank, , and , when , it is the heat output mode of the heat tank , and , wherein, is the maximum heat charging and discharging power, is the maximum capacity, at this time, the heat load gap is updated to .

[0073] S42, photovoltaic consumption and distribution correction, photovoltaic power generation priority to meet the electrical load, the remaining photovoltaic power is used to drive the ground source heat pump, electric boiler and electrolytic cell in turn. Specifically, first calculate the power of photovoltaic to meet the basic electrical load Secondly, the power of photovoltaic driving ground source heat pump is calculated Thirdly, the power of photovoltaic driving electric boiler is calculated And the heat load gap is updated to , , ,

[0074] S43, fuel cell action correction, when photovoltaic output is insufficient, use fuel cell to supplement power supply and heat supply, the correction process needs to meet the original action instruction, the maximum power of the device and the physical constraint of the current hydrogen storage. Specifically, first calculate the power of fuel cell to meet the remaining electrical load, that is Secondly, the power of fuel cell driving ground source heat pump is calculated Finally, the power of fuel cell driving electric boiler is calculated ; Wherein The heat generated by the fuel cell is finally determined as the total output power of the fuel cell .

[0075] S44, electrolytic cell action correction, electrolytic cell uses photovoltaic residual power to produce hydrogen, and only runs when the fuel cell does not work, in order to avoid energy conversion conflict and be limited by hydrogen storage tank capacity. Specifically, if the total output power of the fuel cell is Then the electrolytic cell power is If

[0076] Then the electrolytic cell power is Wherein The intermediate state of hydrogen storage is calculated according to the action of electrolytic cell and fuel cell .

[0077] S45, hydrogen purchase action correction, according to the hydrogen storage threshold to determine whether to buy hydrogen from the market. Specifically, if the intermediate state of hydrogen storage is The actual purchase amount is If The actual purchase amount is Wherein, The hydrogen storage safety threshold.

[0078] In equation (43), the final reward function Includes the basic loss function , the power loss penalty function , thermal loss penalty function , wherein the electric loss penalty function and the thermal loss penalty function are both in a three-segment piecewise linear structure, and the basic loss function, the electric loss penalty function, and the thermal loss penalty function have specific forms as follows:

[0079] (44),

[0080] (45),

[0081] (46),

[0082] wherein in formula (44), , , are respectively the component operation cost, the hydrogen purchase cost, and the photovoltaic disposal cost at time t, and the calculation formulas are shown in formula (3), formula (4), and formula (5) respectively; in formula (45), is a three-segment electric loss penalty function, wherein are respectively the first segment, the second segment, and the third segment electric loss penalty function, and are specifically as follows: the first segment: if , then , the second segment: if , the third segment: if , wherein represents a severe threshold value of electric load loss; represents a moderate threshold value of electric load loss, , are preset electric load loss threshold coefficients, is a penalty coefficient, and satisfies a monotonicity constraint ; in formula (46), is a three-segment thermal loss penalty function, wherein are respectively the first segment, the second segment, and the third segment thermal loss penalty function, and are specifically as follows: the first segment: if , then , the second segment: if , then wherein represents a severe threshold value of thermal load loss; represents a moderate threshold value of thermal load loss, , are preset thermal load loss threshold coefficients; is a thermal loss penalty coefficient, and satisfies a monotonicity constraint, i.e.: .

[0083] Further, based on the multi-role large language model, the safety Markov decision process is solved by the near-end strategy optimization algorithm, and an alternating iteration training stage and parameter adjustment stage are adopted until the penalty function parameters reach the predefined convergence standard, specifically including:

[0084] S61, training stage: the agent interacts with the environment using the current penalty function parameters, specifically including:

[0085] 1) Initialize the policy network and the value network ;

[0086] 2) The agent interacts with the environment to collect experience data, and calculates the reward value according to the three-section penalty function;

[0087] 3) Calculate the advantage function and reward based on the collected experience;

[0088] 4) Update the network parameters using the combined loss function :

[0089] (47),

[0090] In equation (47), is the clipping policy loss function, is the loss function of the value network, is the policy entropy reward function, and are balance coefficients;

[0091] S62, parameter adjustment stage: after the trained model is evaluated on the test set, the multi-role large language model is used to optimize the penalty parameters, specifically including:

[0092] 1) Evaluate the current policy performance on the day test set, record the total running cost and electric and heat loss indicators;

[0093] 2) The role of the large language model is evaluated based on the fuzzy interval standard to qualitatively evaluate the performance;

[0094] 3) The parameter optimization role of the large language model generates new three-section penalty function slope parameters and according to the evaluation results and guided thinking chain suggestions;

[0095] S63, the above stages are executed in a loop, and the interaction experience of the agent is cleared after each training to start training again.

[0096] Further, the large language model for evaluating the role is mainly responsible for:

[0097] (1) Parsing The total test results are decomposed into total operating cost, hydrogen purchase cost, component operating cost, and photovoltaic disposal cost;

[0098] (2) Based on the preset fuzzy interval evaluation standard, the electrical loss and thermal loss performance are mapped to multi-level qualitative evaluation labels;

[0099] (3) Analyze the composition of operating costs, identify the main cost drivers, and output structured results containing loss evaluation and performance analysis.

[0100] Further, the large language model system of the parameter tuning role includes six core components:

[0101] (1) Role description: defined as an expert in power system operation and reinforcement learning, capable of extracting relevant domain knowledge from embedded knowledge;

[0102] (2) Environment description: provides detailed specifications for integrated multi-energy systems, including:

[0103] ① Core components: photovoltaic power generation, electrolytic cell, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, heat storage tank, and geothermal well;

[0104] ② Dynamic characteristics: electrical balance, thermal balance, hydrogen management, and energy storage operation;

[0105] ③ Operating constraints: efficiency factors, capacity limits, and hierarchical load loss penalties;

[0106] (3) Task description: summarizes the main optimization objectives, including ensuring that electrical loss and thermal loss do not exceed specified thresholds, and minimizing total operating costs while maintaining load satisfaction within acceptable limits;

[0107] (4) Output format: strictly defines a fixed output structure, requiring only fixed objects to be output without additional explanations or code;

[0108] (5) Parameter constraints: enforce mathematical monotonicity conditions and , and maintain the coordinated relationship that thermal penalties usually exceed electrical penalties;

[0109] (6) Reward calculation framework: specifies a composite reward structure .

[0110] Further, the S62 includes guiding a chain of thought chain to enhance the logical consistency of parameter adjustment through multi-stage analysis:

[0111] S91, when the evaluation result of the electric loss or the heat loss does not reach the preset satisfaction level, and the total operation cost exceeds the preset threshold , a punishment enhancement strategy or a fine adjustment strategy is adopted to increase the slope parameter of the first segment or the multi-segment slope of the corresponding loss function .

[0112] S92, when the evaluation result of the electric loss or the heat loss reaches the preset satisfaction level, but the total operation cost exceeds the preset threshold , a punishment relaxation strategy or a parameter optimization strategy is adopted to reduce the slope parameter of the corresponding loss function to improve the economy;

[0113] S93, when the evaluation results of the electric loss and the heat loss are different, a multi-problem collaborative processing strategy is adopted to preferentially adjust the punishment parameter of the side with the poorer evaluation result, or a focused strategy is adopted to keep one side stable and optimize the other side;

[0114] S94, during the adjustment process, the monotonicity constraint of the parameter is always maintained, and the magnitude of the parameter adjustment is determined according to the size of the training reward value.

[0115] Further, when the intelligent agent makes online decisions, only the trained policy network needs to be called, that is:

[0116] (48),

[0117] In formula (48), is the trained policy network, is the current state observation value, is the generated optimized action.

[0118] Compared with the prior art, the beneficial effects of the present application are as follows:

[0119] (1) Compared with the existing rule-based control method, the method of the present application can autonomously explore the optimal strategy in a complex operating environment based on a reinforcement learning decision framework, significantly improving the adaptability and economy of the system under variable operating conditions, and overcoming the limitations of rule-based control strategies, such as rigidity and difficulty in dealing with uncertainty.

[0120] (2) Compared with the existing method based on traditional reinforcement learning, the present application realizes the adaptive optimization of the punishment parameter by using a multi-role large language model to assist in the design of the reward function, effectively solving the pain point of difficulty in parameter tuning of the reward function in traditional deep reinforcement learning. At the same time, by introducing a chain thinking mechanism, the logicality and interpretability of parameter adjustment are enhanced, so that the system can better balance the multiple target demands such as economic operation and power supply and heating reliability. BRIEF DESCRIPTION OF DRAWINGS

[0121] Figure 1 A flow chart of a multi-role large model assisted hydrogen-containing building energy system operation optimization method according to the present application;

[0122] Figure 2 An electrical loss convergence curve of a multi-role large model assisted hydrogen-containing building energy system operation optimization method;

[0123] Figure 3 A thermal loss convergence curve of a multi-role large model assisted hydrogen-containing building energy system operation optimization method;

[0124] Figure 4 A reward function convergence curve of a multi-role large model assisted hydrogen-containing building energy system operation optimization method. DETAILED DESCRIPTION

[0125] The present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely intended to more clearly illustrate the technical solutions of the present application, and should not be used to limit the protection scope of the present application.

[0126] Example 1:

[0127] As shown in the design flow chart of a multi-role large model assisted hydrogen-containing building energy system operation optimization method according to the present application, the method comprises the following steps: Figure 1 The established hydrogen-containing building multi-energy system operation cost minimization problem under off-grid operation mode includes objective function, decision variable and constraint condition, i.e.:

[0128] (1) Objective function: simultaneously minimize the time average expected value of system operation cost and energy supply loss, which is:

[0129]

[0130] (1),

[0131] wherein, is the system operation cost, is the electrical loss, is the thermal loss, which is calculated by the following formula:

[0132] (2),

[0133] (3),

[0134] (4),

[0135] (5),

[0136] ​(6),

[0137] (7),

[0138] wherein in formula (1), is a desired operator; is the total operating cost of the system; in formula (2), , , are the component operating cost, hydrogen purchase cost, and photovoltaic curtailment cost at time , respectively; in formula (3), , , , , , are the cost coefficients of the photovoltaic component, the electrolyzer, the fuel cell, the electric boiler, the ground source heat pump, and the thermal storage tank, respectively, , , , , , , are the output power of the photovoltaic, the input power of the electrolyzer, the output power of the fuel cell, the input power of the electric boiler, the input power of the ground source heat pump, the heating input power of the thermal storage tank, and the heat output power of the thermal storage tank at time , respectively; in formula (4), is the hydrogen purchase price at time , is the hydrogen volume purchased at time ; in formula (5), is the photovoltaic curtailment cost coefficient, is the maximum power generation of the photovoltaic at time ; in formula (6), is the electrical loss penalty function, and in formula (7), is the thermal loss penalty function;

[0139] (2) Decision variables: the decision variables include the photovoltaic use power , the electrolyzer input power , the fuel cell output power , the electric boiler input power , the ground source heat pump input power , the thermal storage tank charging power , the thermal storage tank discharging power , and the hydrogen amount purchased from the hydrogen market at each time slot ;

[0140] (3) Constraints: Constraints include power balance and equipment physical constraints:

[0141] Photovoltaic constraints:

[0142] (8),

[0143] (9),

[0144] Hydrogen storage tank and hydrogen balance constraints:

[0145] (10),

[0146] (11),

[0147] (12),

[0148] (13),

[0149] (14),

[0150] (15),

[0151] (16),

[0152] (17),

[0153] Electrolyzer constraints:

[0154] (18),

[0155] (19),

[0156] (20),

[0157] Fuel cell constraints:

[0158] (21),

[0159] (22),

[0160] (23),

[0161] (24),

[0162] (25),

[0163] Thermal storage tank constraints:

[0164] (26),

[0165] (27),

[0166] (28),

[0167] (29),

[0168] (30),

[0169] Ground source heat pump and geothermal well constraints:

[0170] (31),

[0171] (32),

[0172] (33),

[0173] (34),

[0174] (35),

[0175] (36),

[0176] Electric boiler constraints:

[0177] (37),

[0178] (38),

[0179] System supply and demand balance constraints:

[0180] (39),

[0181] (40),

[0182] In formula (8), is the photovoltaic efficiency coefficient, is the total area of the photovoltaic panel, is the solar radiation intensity of the time slot; in formula (10), is the time slot hydrogen storage capacity of the hydrogen storage tank, is the hydrogen production capacity of the electrolytic tank in a given time period, is the hydrogen consumption capacity of the fuel cell in a given time period, is the time slot purchased hydrogen capacity, in equation (11) is the maximum hydrogen storage amount of the hydrogen storage tank, in equation (12) is the maximum purchased hydrogen amount, in equation (13), is the gas compressibility factor, , is a fitting parameter, in equation (14), is the time slot is the hydrogen storage tank pressure, is the hydrogen storage tank volume, is the ideal gas constant, is the time slot is the hydrogen storage tank temperature, is the molar mass of hydrogen, in equation (15), is the hydrogen storage tank safety margin pressure, is a binary variable, is the relaxation constraint, in equation (16), is the hydrogen storage tank ultimate maximum pressure, in equation (17), is the total number of time periods in which the hydrogen storage tank pressure exceeds the safety margin, in equation (18), is the hydrogen generation capacity of the electrolyzer, is the hydrogen production efficiency of the electrolyzer, is the time slot is the input power of the electrolyzer, is the time interval, in equation (19), is the electrolyzer maximum input power, in equation (20), is the ramp rate of the electrolyzer, is the input power of the electrolyzer in the previous time slot, in equation (21), is the hydrogen consumption of the fuel cell, is the time slot is the output electric power of the fuel cell, is the hydrogen-to-electricity efficiency of the fuel cell, in equation (22), is the fuel cell maximum output power, in equation (23), is the ramp rate of the fuel cell, is the time slot is the fuel cell output electric power, in equation (25), is the time slot is the heat generated by the fuel cell, is the electric-to-thermal conversion ratio of the fuel cell, is the heat recovery coefficient, in equation (26), is the time slot is the thermal storage level of the thermal storage tank, is the time slot is the thermal charging power of the thermal storage tank, for the time slot the heat release power of the thermal storage tank, for the charging efficiency, for the heat release efficiency, in equation (27), for the maximum thermal storage capacity of the thermal storage tank, in equation (28), for the maximum charging power, in equation (29), for the maximum heat release power, in equation (31), for the time slot the heat power output by the ground source heat pump, for the time slot the electric power input to the ground source heat pump, for the ground source heat pump efficiency, in equation (32), for the maximum electric power input to the ground source heat pump, in equation (33), for the time slot the thermal storage level of the geothermal well, for the time slot the fuel cell waste heat injected into the geothermal well, in equation (34), for the maximum thermal storage capacity of the geothermal well, in equation (35), indicating the waste heat injected into the geothermal well not exceeding the waste heat generated by the fuel cell , in equation (37), for the time slot the heat generated by the electric boiler, for the time slot the electric power input to the electric boiler, for the electric boiler efficiency, in equation (38), for the maximum electric power input to the electric boiler, in equation (39), for the time slot the electric load, for the time slot the electric power consumed by the electrolyzer, for the time slot the electric power consumed by the ground source heat pump, for the time slot the electric loss, in equation (40), for the time slot the thermal load, for the time slot the heat consumed by the thermal storage tank charging, for the time slot the waste heat injected into the geothermal well, for the time slot the waste heat power generated by the fuel cell, for the time slot the heat generated by the ground source heat pump, for the time slot heat provided by the heat storage tank discharging, for time slot thermal loss.

[0183] Further, the step 2 minimizes the operation cost problem is re-modeled as a Markov decision process with safety correction mechanism related to hydrogen-containing building multi-energy system operation optimization, defined as a tuple . Wherein, is the state space related to the operation of the hydrogen-containing building multi-energy system, is the original action space, is the state transition function, is the final reward function, is the discount factor, in order to solve the Markov decision process, the rule-based action correction mechanism is introduced as a pre-module of policy execution. The specific expression is as follows:

[0184] (41),

[0185] (42),

[0186] (43),

[0187] Wherein, in formula (41), is the current operating hours; each action component is normalized to the range of [0, 1] or [-1, 1], and in formula (42) , is the normalized heat charging and discharging power of the heat storage tank; , are the normalized power of the fuel cell allocated to the electrical load, the ground source heat pump and the electric boiler respectively; , is the normalized hydrogen purchase amount; in formula (43), is the final reward function, is the basic reward function.

[0188] Further, the step 2 minimizes the operation cost problem is re-modeled as a Markov decision process with safety correction mechanism related to hydrogen-containing building multi-energy system operation optimization, defined as a tuple is the rule-based action correction mechanism in the safety Markov decision process, in order to ensure the balance between supply and demand of the system in off-grid mode and physical constraints, the rule-based action correction mechanism is introduced to correct the original action to the actual execution action , and the correction step is as follows:

[0189] S41, heat tank action correction, according to determine the actual heat charging power or heat discharging power and update the remaining heat load gap when is the heat input mode of the hot tank, and when is the heat output mode of the hot tank and wherein, is the maximum charge and discharge heat power, is the maximum capacity, at this time, the heat load gap is updated to .

[0190] S42, photovoltaic accommodation and distribution correction, photovoltaic power generation priority meets the electrical load, the remaining photovoltaic power is used to drive the ground source heat pump, electric boiler and electrolytic cell in turn. Specifically, first calculate the power of photovoltaic to meet the basic electrical load , secondly, calculate the power of photovoltaic to drive the ground source heat pump ; Thirdly, calculate the power of photovoltaic to drive the electric boiler , and update the heat load gap to , , are the power of photovoltaic output to the ground source heat pump and electric boiler respectively.

[0191] S43, fuel cell action correction, when the photovoltaic output is insufficient, the fuel cell is used to supplement power supply and heat supply, and the correction process needs to meet the original action instruction, the maximum power of the equipment and the physical constraint of the current hydrogen storage amount. Specifically, first calculate the power of fuel cell to meet the remaining electrical load, that is ; Secondly, calculate the power of fuel cell to drive the ground source heat pump:

[0192] ; wherein, is the heat generated by the fuel cell, and finally, calculate the power of fuel cell to drive the electric boiler ; Finally, the total output power of the fuel cell is determined as .

[0193] S44, electrolytic cell action correction, electrolytic cell uses the remaining power of photovoltaic to produce hydrogen, and only runs when the fuel cell does not work, so as to avoid energy conversion conflict and be limited by the capacity of hydrogen storage tank. Specifically, if the total output power of the fuel cell is , the power of the electrolytic cell is ; If

[0194] , the power of the electrolytic cell is ; According to the action of electrolytic cell and fuel cell, the hydrogen storage amount of the intermediate state is calculated .

[0195] S45, hydrogen purchase action correction, according to the hydrogen storage amount threshold value to determine whether to purchase hydrogen from the market. Specifically, if the intermediate state hydrogen storage amount , the actual purchase amount , if , the actual purchase amount , wherein is the hydrogen storage amount safety threshold.

[0196] Further, in the formula (43), the final reward function includes the basic loss function , the electric loss penalty function , and the thermal loss penalty function , wherein the electric loss penalty function and the thermal loss penalty function both adopt a three-segment piecewise linear structure, and the basic loss function, the electric loss penalty function, and the thermal loss penalty function have the following specific forms:

[0197] (44),

[0198] (45),

[0199] (46),

[0200] wherein in the formula (44), , , are the component operating cost, the hydrogen purchase cost, and the photovoltaic disposal cost at time , and the calculation formulas are shown in the formulas (3), (4), and (5) respectively; in the formula (45), is a three-segment electric loss penalty function, wherein are the first segment, the second segment, and the third segment electric loss penalty functions, is the current electric loss value, and specifically: the first segment: if , the second segment: if the third segment: if , wherein indicates the severe threshold value of the electric load loss; indicates the moderate threshold value of the electric load loss, , are preset electric load loss threshold coefficients, is a penalty coefficient, and satisfies the monotonicity constraint ; in the formula (46), is a three-segment thermal loss penalty function, wherein respectively, are a first segment, a second segment, and a third segment heat loss penalty function, is a current heat loss value, specifically, respectively, is a first segment: if , is a second segment: if , is a third segment: if , wherein represents a severe threshold of heat load loss; represents a moderate threshold of heat load loss, , is a preset heat load loss threshold coefficient; is a heat loss penalty coefficient, and satisfies a monotonicity constraint, that is: ;

[0201] Further, the step 3 of solving the modeling safe Markov decision process based on the multi-role large language model assisted near-end strategy optimization algorithm adopts an alternating iteration of a training phase and a parameter adjustment phase until the penalty function parameter reaches a predefined convergence standard, and specifically includes:

[0202] S61, training phase: the agent interacts with the environment using the current penalty function parameter, specifically including:

[0203] 1) initialize the policy network and the value network ;

[0204] 2) the agent interacts with the environment to collect experience data, and calculates the reward value according to the three-segment penalty function;

[0205] 3) calculate the advantage function and the reward based on the collected experience;

[0206] 4) update the network parameters using the combined loss function :

[0207] (47),

[0208] In formula (47), is a clipping policy loss function, is a loss function of the value network, is a policy entropy reward function, and are balance coefficients;

[0209] S62, parameter adjustment phase: after the trained model is evaluated on the test set, the multi-role large language model is used to optimize the penalty parameter, specifically including:

[0210] 1) On the test set of 1 day, evaluate the performance of the current strategy, record the total running cost and the indicators of electricity and heat loss;

[0211] 2) Evaluate the role of large language model based on fuzzy interval standard to qualitatively evaluate the performance;

[0212] 3) Parameter optimization role large language model generates new three-section penalty function slope parameters (a, b, c) according to the evaluation results and guided thinking chain suggestions; And );

[0213] S63, the above stages are executed in a loop, and the interaction experience of the agent is cleared after each round of training, and the training is restarted.

[0214] Further, the evaluation role of large language model is mainly responsible for:

[0215] (1) Analyze the cumulative test results of 1 day, and decompose the total running cost, hydrogen purchase cost, component running cost and photovoltaic disposal cost;

[0216] (2) Based on the preset fuzzy interval evaluation standard, map the performance of electricity loss and heat loss into multi-level qualitative evaluation labels;

[0217] (3) Analyze the composition of running cost, identify the main cost driving factors, and output the structured results containing loss evaluation and performance analysis.

[0218] Further, the parameter optimization role of large language model system prompt words includes six core components:

[0219] (1) Role description: defined as an expert in power system operation and reinforcement learning, capable of extracting relevant domain knowledge from embedded knowledge;

[0220] (2) Environment description: provides detailed specifications of integrated multi-energy systems, including:

[0221] ① Core components: photovoltaic power generation, electrolytic cell, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, heat storage tank and geothermal well;

[0222] ② Dynamic characteristics: power balance, heat balance, hydrogen management, energy storage operation;

[0223] ③ Operation constraints: efficiency factors, capacity limits, hierarchical load loss penalties;

[0224] (3) Task description: summarizes the main optimization objectives, including ensuring that the electricity loss and heat loss do not exceed the specified threshold, and minimizing the total running cost while maintaining the load satisfaction within the acceptable limits;

[0225] ​​(4) Output format: strictly define fixed output structure, require only output fixed objects without additional explanation or code;

[0226] (5) Parameter constraints: enforce mathematical monotonicity conditions and maintain the coordinated relationship that the thermal penalty is usually higher than the electrical penalty;

[0227] (6) Reward calculation framework: specify the composite reward structure .

[0228] Further, the S62 comprises guiding the chain of thought chain, enhancing the logical consistency of parameter adjustment through multi-stage analysis:

[0229] S91, when the evaluation results of electrical loss or thermal loss do not reach the preset satisfaction level, and the total operating cost exceeds the preset threshold , a penalty enhancement strategy or fine adjustment strategy is adopted to increase the first segment slope or multi-segment slope parameter of the corresponding loss function;

[0230] S92, when the evaluation results of electrical loss or thermal loss reach the preset satisfaction level, but the total operating cost exceeds the preset threshold , a penalty relaxation strategy or parameter optimization strategy is adopted to reduce the slope parameter of the corresponding loss function to improve economy;

[0231] S93, when the evaluation results of electrical loss and thermal loss differ, a multi-problem collaborative processing strategy is adopted to preferentially adjust the penalty parameter of the side with worse evaluation results, or a focused strategy is adopted to maintain one side stable while optimizing the other side;

[0232] S94, during the adjustment process, the monotonicity constraint of the parameter is always maintained, and the magnitude of parameter adjustment is determined according to the size of the training reward value.

[0233] Further, when the agent makes online decisions in step 4, only the trained policy network needs to be called, that is:

[0234] (48),

[0235] In formula (48), is the trained policy network, is the current state observation value, is the generated optimized action.

[0236] In order to verify the effectiveness and advancement of the multi-role large model assisted hydrogen-containing building energy system operation optimization method proposed in the present application, three groups of comparison schemes are set.

[0237] Comparative scheme one: this scheme adopts a heuristic operation strategy system based on engineering experience. The core strategy adopted aims to prioritize meeting the electricity and heat load demand, and follow fixed device start-stop and energy management rules. Specifically, the photovoltaic power is first used to meet the electricity load demand, and the remaining power is allocated to the ground source heat pump, electric boiler and electrolytic cell in a fixed priority order; the hydrogen storage tank and the heat storage tank are controlled based on the preset energy charging and discharging threshold, and the fuel cell and the electrolytic cell follow the mutual exclusion operation principle to avoid energy conversion loss.

[0238] Comparative scheme two: this scheme adopts a standard near-end strategy optimization algorithm, but the electricity and heat loss penalty coefficients in the reward function and are fixed values, and LLM is not introduced for closed-loop optimization. In specific implementation, the agent is trained based on the same state space and action space, and the reward function adopts the same mathematical form as the present invention, but all penalty parameters are determined before the start of training and remain unchanged throughout the training process. By comparing with this baseline, the specific contribution of the multi-role large model assisted hydrogen-containing building energy system operation optimization method proposed in this paper to the final performance improvement can be clearly separated and quantified.

[0239] Comparative scheme three: this scheme adopts commercial software GAMS combined with CPLEX solver, and builds the system scheduling problem as a mixed integer linear programming model for direct solution. In this scheme, it is assumed that all future time periods of solar radiation intensity, electricity and heat load demand, hydrogen price and other uncertain parameters can be perfectly predicted. The performance under this scheme provides an upper limit for the performance of the present invention.

[0240] Figure 2 and Figure 3 respectively show the convergence process of electricity loss and heat loss during training. Figure 4 shows the reward convergence curve during training. The results of the figure show the convergence of the multi-role large language model assisted reinforcement learning process.

[0241] Table 1- Comparison table of various costs and safety of the present invention method and other schemes

[0242]

[0243] Table 1 shows the comparison of the experimental results of the method of the present application and the above three comparative schemes. Compared with the first and second comparative schemes, the method of the present application has lower operation cost and heat loss, and the electric loss is 0. Specifically, the method of the present application can reduce the operation cost by 89.04% and 33.76% respectively, and reduce the heat loss by 98.16% and 96.98% respectively. Moreover, compared with the third comparative scheme under perfect prediction information, the relative error of the operation cost of the method of the present application is 4.04%, and the absolute value of the heat loss is 599.32 kWh (equivalent to 0.8 kWh of cut-off per hour), indicating that the method of the present application has excellent near-optimal performance.

Claims

1. A method for optimizing the operation of a hydrogen-containing building energy system using a multi-role large model, characterized in that, Includes the following steps: Step 1: Establish the problem of minimizing the operating cost of hydrogen-containing multi-energy systems in off-grid operation mode; Step 2: Remodel the operating cost minimization problem as a safe Markov decision process related to the operation optimization of multi-energy systems in hydrogen-containing buildings; Step 3: Solve the modeled safe Markov decision process based on the multi-role large language model-assisted proximal policy optimization algorithm to obtain the agent operation strategy related to the hydrogen-containing building multi-energy system; Step 4: The intelligent agent makes online decisions based on the hydrogen-containing building multi-energy system operation strategy obtained through training, and applies the decisions to the actual hydrogen-containing building multi-energy system.

2. The method according to claim 1, characterized in that, The problem of minimizing the operating cost of a hydrogen-containing building multi-energy system under off-grid operation mode established in Step 1 includes the objective function, decision variables, and constraints, as follows: (1) Objective function: The objective function is to minimize the expected time average of both system operating cost and energy loss, which is: (1), in, The total operating cost of the system, For electrical loss, Heat loss is calculated using the following formula: (2), (3), (4), (5), (6), (7), In equation (1), For expectation operators; The total operating cost of the system; in equation (2), , , In time The component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost; in formula (3), , , , , , These are the cost coefficients for photovoltaic modules, electrolyzers, fuel cells, electric boilers, ground source heat pumps, and thermal storage tanks, respectively. , , , , , , In time Photovoltaic output power and time Input power and time of the electrolytic cell Fuel cell output power and time Input power and time of electric boiler Input power and time of ground source heat pump Heating input power and time of the thermal storage tank The heat dissipation output power of the thermal storage tank; in equation (4), In time The purchase price of hydrogen. In time The volume of hydrogen purchased, in formula (5), This is the cost reduction coefficient for photovoltaics. In time The maximum power generation of photovoltaics; in equation (6), Let be the electrical loss penalty function, in equation (7), This is the heat loss penalty function; (2) Decision variables: Decision variables include time Photovoltaic output power ,time Electrolytic cell input power ,time fuel cell output power ,time Electric boiler input power ,time Ground source heat pump input power ,time Thermal storage tank heating input power ,time Thermal storage tank heat output power and in time Purchased hydrogen volume ; (3) Constraints: Constraints include power balance and equipment physical constraints: Photovoltaic constraints: (8), (9), Hydrogen storage tanks and hydrogen balance constraints: (10), (11), (12), (13), (14), (15), (16), (17), Electrolytic cell constraints: (18), (19), (20), Fuel cell constraints: (21), (22), (23), (24), (25), Thermal storage tank constraints: (26), (27), (28), (29), (30), Constraints of ground source heat pumps and geothermal wells: (31), (32), (33), (34), (35), (36), Electric boiler constraints: (37), (38), System supply and demand balance constraints: (39), (40), In equation (8), The photovoltaic efficiency coefficient. The total area of ​​the photovoltaic panels. for Solar radiation intensity in time slot; in equation (10), For time slots The hydrogen storage capacity of the hydrogen storage tank. This represents the amount of hydrogen produced by the electrolyzer within a given time period. This represents the hydrogen capacity consumed by the fuel cell within a given time period. For time slots The volume of hydrogen purchased, in formula (11) The maximum hydrogen storage capacity of the hydrogen storage tank is given in equation (12). To maximize the amount of hydrogen purchased, in equation (13), The gas compressibility factor, , For the fitting parameters, in equation (14), For time slots Hydrogen storage tank pressure, This refers to the volume of the hydrogen storage tank. Let be the ideal gas constant. For time slots Hydrogen storage tank temperature, Let be the molar mass of hydrogen gas, in equation (15), For the safety margin pressure of the hydrogen storage tank, It is a binary variable. To relax the constraints, in equation (16), For the maximum pressure limit of the hydrogen storage tank, in equation (17), Let be the total number of time periods during which the pressure in the hydrogen storage tank exceeds the safety margin, as given in equation (18). The capacity of hydrogen produced by the electrolyzer. The hydrogen production efficiency of the electrolyzer. For time slots Input power of the internal electrolytic cell, For the time interval, in equation (19), For the maximum input power of the electrolytic cell, in equation (20), The ramp rate of the electrolytic cell. They represent the input power of the electrolytic cell in the previous time slot, respectively. In equation (21), The amount of hydrogen consumed by the fuel cell. For time slots The output power of the internal fuel cell, For the hydrogen-to-electricity conversion efficiency of the fuel cell, in equation (22), For the maximum output power of the fuel cell, in equation (23), For the ramp rate of the fuel cell, Indicates time slot The output electrical power of the internal fuel cell is given by equation (25). For time slots Waste heat generated by fuel cells The electrothermal conversion ratio of the fuel cell, For the heat recovery coefficient, in equation (26), For time slots The thermal storage capacity of the thermal storage tank For time slots The charging power of the thermal storage tank For time slots The heat dissipation capacity of the thermal storage tank For heat charging efficiency, For heat release efficiency, in equation (27), For the maximum heat storage capacity of the heat storage tank, in equation (28), For the maximum heat charging power, in equation (29), For the maximum heat release power, in equation (31), For time slots The heat output power of a ground source heat pump For time slots The electrical power input to the ground source heat pump, For the efficiency of a ground source heat pump, in equation (32), For the maximum input electrical power of the ground source heat pump, in equation (33), For time slots The heat storage level of geothermal wells For time slots The waste heat from the fuel cell injected into the geothermal well, in equation (34), For the maximum heat storage capacity of the geothermal well, in equation (35), represents the waste heat injected into the geothermal well. No more than the waste heat generated by the fuel cell In equation (37), For time slots The heat generated by the electric boiler For time slots The electrical power input to the electric boiler, For the efficiency of the electric boiler, in equation (38), For the maximum input electrical power of the electric boiler, in equation (39), For time slots electrical load, The electrical power consumed by the electrolytic cell in time slot t. For time slots The electrical power input to the ground source heat pump, For time slots Fuel cell output power, For time slots The electrical loss, in equation (40), For time slots heat load, For time slots The heat consumed during the charging of the thermal storage tank. For time slots Waste heat injected into geothermal wells For time slots Waste heat generated by fuel cells For time slots The heat output power of a ground source heat pump For time slots The heat released by the thermal storage tank For time slots Heat loss.

3. The method according to claim 2, characterized in that, In step 2, the problem of minimizing operating costs is remodeled as a Markov decision process with a safety correction mechanism related to the operation optimization of multi-energy systems in hydrogen-containing buildings, defined as a tuple. ,in, For the state space related to the operation of multi-energy systems in hydrogen-containing buildings, For the original action space, This is the state transition function. For the final reward function, A rule-based action correction mechanism is introduced as the discount factor. As a prerequisite module for strategy execution, the specific expression is as follows: (41), (42), (43), In equation (41), The current running hour; each action component is normalized to the range of [0,1] or [-1,1], in equation (42) , The normalized heat storage tank charging and discharging power; , These are the normalized power allocated by the fuel cell to the electrical load, the ground source heat pump, and the electric boiler, respectively. , The normalized amount of hydrogen purchased; in equation (43), For the final reward function, This is the basic reward function.

4. The method according to claim 3, characterized in that, Introducing a rule-based action correction mechanism Original action Corrected to actual action The correction steps are as follows: S41, Hot Can Action Correction, according to Determine the actual heat charging power of the thermal storage tank or heat dissipation power and update the remaining heat load gap. ,when At this time, it is the heat input mode of the hot tank. ,and ,when At this time, it is the heat output mode of the hot tank. ,and ,in, For maximum charge / discharge power, At maximum capacity, the heat load gap is updated to... ; S42. Photovoltaic Consumption and Distribution Correction: Photovoltaic power generation prioritizes meeting the basic electrical load. Remaining photovoltaic power is then used to drive ground-source heat pumps, electric boilers, and electrolytic cells. Specifically, the power generated by photovoltaics to meet the basic electrical load is first calculated. Secondly, calculate the power of the photovoltaic-driven ground source heat pump. ; Next, calculate the power of the photovoltaic-driven electric boiler. And update the heat load gap to , , The output power of the photovoltaic system is allocated to the power of the ground source heat pump and the electric boiler, respectively. S43. Fuel Cell Operation Correction: When photovoltaic output is insufficient, the fuel cell is used to supplement power and heat supply. The correction process must simultaneously meet the physical constraints of the original operation command, the maximum power of the equipment, and the current hydrogen storage capacity. Specifically, the power of the fuel cell to meet the remaining electrical load is first calculated, i.e. Secondly, calculate the power of the fuel cell driving the ground source heat pump. Finally, the power of the fuel cell-driven electric boiler was calculated. ,in, The heat generated by the fuel cell; the total output power of the fuel cell was ultimately determined to be... ; S44. Electrolyzer operation correction: The electrolyzer utilizes surplus photovoltaic power to produce hydrogen and operates only when the fuel cell is not working. Specifically, if the total output power of the fuel cell... Then the power of the electrolytic cell ;like Then the power of the electrolytic cell ;in After allocating the photovoltaic output power to the ground source heat pump and electric boiler, the remaining power is used to calculate the hydrogen storage capacity in the intermediate state based on the operation of the electrolyzer and fuel cell. ; S45. Hydrogen purchase action correction: Determine whether to purchase hydrogen from the market based on the hydrogen storage threshold. Specifically, if the intermediate hydrogen storage level... The actual purchase quantity ,like The actual purchase quantity ,in, This is the safe threshold for hydrogen storage capacity.

5. The method according to claim 4, characterized in that, In equation (43), the final reward function Includes the basic loss function Electrical loss penalty function Heat loss penalty function The electrical loss penalty function and the thermal loss penalty function both adopt a three-segment piecewise linear structure. The specific forms of the basic loss function, electrical loss penalty function, and thermal loss penalty function are as follows: (44), (45), (46), In equation (44), , , In time The calculation formulas for the component operating cost, hydrogen purchase cost, and photovoltaic abandonment cost are shown in equations (3), (4), and (5), respectively; in equation (45), Here is a three-segment electrical loss penalty function, where, The electrical loss penalty functions are for the first, second, and third segments, respectively, and are as follows: First segment: If ,but Second paragraph: If Third paragraph: If ,but ,in, This represents the severe threshold of electrical load loss; This represents the moderate threshold of electrical load loss. , The preset electrical load loss threshold coefficient, The value is the penalty coefficient, and it satisfies the monotonicity constraint. In equation (46), Here is a three-segment heat loss penalty function, where, These are the heat loss penalty functions for the first, second, and third segments, respectively, specifically: First segment: If... ,but Second paragraph: If ,but Third paragraph: If ,but ,in The threshold indicating the severity of heat load loss; This represents the intermediate threshold of heat load loss. , This is the preset heat load loss threshold coefficient; Let be the heat loss penalty coefficient, and satisfy the monotonicity constraint, i.e.: .

6. The method according to claim 5, characterized in that, Step 3, which uses a multi-role large language model-assisted proximal policy optimization algorithm to solve the modeled secure Markov decision process, employs alternating iterative training and parameter tuning phases until the penalty function parameters reach a predefined convergence criterion. Specifically, this includes: S61. Training Phase: The agent interacts with the environment using the current penalty function parameters, specifically including: 1) Initialize the policy network and value network ; 2) The agent interacts with the environment to collect experience data and calculates the reward value based on the three-stage penalty function; 3) Calculate the advantage function based on collected experience. and rewards ; 4) Use the combined loss function Update network parameters: (47), In equation (47), Let the loss function be the shearing strategy. The loss function of the value network, Let the policy entropy reward function be... and This is the balance coefficient; S62. Parameter Tuning Phase: After the trained model is evaluated on the test set, the penalty parameters are optimized using a multi-role large language model, specifically including: 1) In The performance of the current strategy is evaluated on the test set, and the total operating cost and electrical and heat loss indicators are recorded. 2) The performance of the large language model for evaluating roles is qualitatively assessed based on fuzzy interval criteria; 3) Parameter tuning: Based on the evaluation results and guided thought chain suggestions, the large language model generates a new three-stage penalty function slope parameter. as well as ; S63. Repeat the above stages, clearing the agent's interaction experience after each round of training and starting training again.

7. The method according to claim 6, characterized in that, The large language model responsible for evaluating roles is: (1) Analysis The cumulative test results over the days are used to break down the total operating cost, hydrogen purchase cost, component operating cost, and photovoltaic abandonment cost. (2) Based on the preset fuzzy interval evaluation criteria, the electrical loss and thermal loss performance are mapped to multi-level qualitative evaluation labels; (3) Analyze the composition of operating costs, identify cost drivers, and output structured results including loss assessment and performance analysis.

8. The method according to claim 7, characterized in that, The large language model system prompt words for parameter tuning roles include six core components: (1) Role description: Defined as an expert in power system operation and reinforcement learning; (2) Environmental Description: Provide detailed specifications for the integrated multi-energy system, including: Core components: photovoltaic power generation, electrolyzer, fuel cell, ground source heat pump, electric boiler, hydrogen storage tank, thermal storage tank and geothermal well; Dynamic characteristics: electrical balance, thermal balance, hydrogen management, and energy storage operation; Operational constraints: efficiency factor, capacity limit, and tiered load loss penalty; (3) Task description: Outline the optimization objectives, including ensuring that electrical and heat losses do not exceed specified thresholds, and minimizing total operating costs while maintaining load satisfaction within acceptable limits; (4) Output format: Strictly define a fixed output structure, requiring only fixed objects to be output without any additional explanation or code; (5) Parameter constraints: enforce mathematical monotonicity conditions as well as And maintain a coordinated relationship where thermal penalty exceeds electrical penalty; (6) Reward calculation framework: specify the composite reward structure .

9. The method according to claim 7, characterized in that, Step S62 includes a guided chain of thought, enhancing the logical consistency of parameter adjustments through multi-stage analysis: S91. When the assessment results of electrical or thermal losses do not reach the preset satisfaction level, and the total operating cost exceeds the preset threshold... At this time, a penalty enhancement strategy or a fine-tuning strategy is adopted to increase the slope of the first segment of the corresponding loss function. Or multiple slope parameters; S92. When the assessment results of electrical loss or heat loss reach the preset satisfaction level, but the total operating cost exceeds the preset threshold... At that time, a penalty relaxation strategy or parameter optimization strategy is adopted to reduce the slope parameter of the corresponding loss function; S93. When there is a difference between the assessment results of electrical loss and heat loss, a multi-problem collaborative processing strategy is adopted, prioritizing the adjustment of the penalty parameter of the side with the worse assessment result, or a focus strategy is adopted to keep one side stable while optimizing the other side. S94. During the adjustment process, always maintain the monotonicity constraint of the parameters, and determine the magnitude of parameter adjustment based on the training reward value.

10. The method according to claim 1, characterized in that, In step 4, when the agent makes online decisions, it only needs to call the trained policy network, that is: (48), In equation (48), For the trained policy network, The current state observation value, For the generated optimized actions.

Citation Information

Patent Citations

  • Narrow space mechanical arm path planning method based on large language model

    CN120755879A

  • Distributed energy storage charging multifunctional cooperative control method and system

    CN120879718A

  • Method and device for cooperative control and intelligent optimization of urban rail transit regional passenger flow

    CN121328986A

Cited By

  • Prediction and decision-making integrated resource scheduling method and system for electricity-hydrogen-heat integrated energy system

    CN122026525A