Risk perception scheduling method for multi-energy microgrid based on distributed Pareto reinforcement learning of large language model
Through the combination of large language model and distributed Pareto reinforcement learning, the uncertainty problem of complex environments in multi-energy microgrid scheduling is solved, effective management of extreme risks and personalized strategy selection is achieved, and the robustness and robustness of the system are improved.
Patent Information
- Application Number
- CN202510257469.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively deal with complex and dynamic uncertain environments in multi-energy microgrid scheduling. Traditional methods cannot accurately capture extreme risks, and lack diversified strategic support for agents with different risk preferences, resulting in insufficient adaptability of optimization results when facing extreme risks.
The method of large language model (LLM) combined with distributed Pareto reinforcement learning (DRL) is used to simulate real climate scenarios, build a state transfer function matrix, design the Markov decision-making process, and use the LLM-DPSAC algorithm to determine the optimization target, and realize risk-aware scheduling of multi-energy microgrids.
It significantly improves the modeling ability and flexible response ability of complex environments, can effectively capture low-probability and high-consequence events, provide personalized optimization strategies, and improves the robustness and robustness of the system under extreme conditions.
Smart Images

Figure CN120258507A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of microgrid risk perception scheduling, and specifically to a risk perception scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model. Background Technique
[0002] The operation of a multi-energy microgrid faces many uncertainties. Conditional Value at Risk (CVaR) can solve extreme weather and tail risk situations in the scheduling optimization of a multi-energy microgrid. CVaR is an effective risk measurement tool, especially suitable for scenarios considering tail risk and extreme losses. Different from traditional expected value optimization, CVaR focuses on the impact of extreme losses or rare events and quantifies the potential losses that the system may face in the worst case, effectively capturing high-volatility, low-probability risks that may lead to significant consequences and ensuring the stability of the system even under extreme conditions. When solving optimization problems related to CVaR, traditional methods such as mixed-integer programming and heuristic algorithms cannot fully explore the solution space. Especially when dealing with complex non-linear problems, they often lead to low efficiency and inaccuracy. In contrast, Deep Reinforcement Learning (DRL) can learn an approximate optimal policy through interaction with the environment, making it particularly suitable for managing the uncertainties and dynamics inherent in the decision-making process.
[0003] The existing technology has at least the following disadvantages: First, traditional uncertainty modeling methods often rely on predefined rules or statistical models, and these methods usually cannot effectively generate an uncertain environment in complex and dynamically changing real scenarios. Especially when facing changing external conditions and highly volatile systems, traditional methods often cannot accurately capture and reflect the uncertainties in the system, resulting in large errors and unreliability in practical applications. Second, current methods that combine CVaR with economic benefit optimization usually only focus on the average expected return and ignore low-probability, high-consequence extreme events. When the system faces large fluctuations or adverse situations, simply relying on the expected return as the optimization goal may lead to ineffective risk control and thus unforeseen major losses. In addition, existing DRL methods mainly focus on linear preferences based on expected values and lack support for diverse strategies of agents when facing different risk preferences. Reinforcement learning fails to fully consider the diversity and uncertainty distribution of risks, resulting in insufficient adaptability of the optimization results when facing extreme risks. Therefore, based on the current technology, how to effectively combine uncertainty modeling, risk management, and diverse strategy selection has become a key problem to be solved urgently. Summary of the Invention
[0004] The object of the present invention is to provide a risk perception scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model, including the following steps:
[0005] 1) Use the LLM to simulate and generate real climate scenarios for the multi - energy microgrid;
[0006] 2) Update the state - transition function P(s'|s,a) corresponding to each real climate scenario, thereby constructing a state - transition function matrix;
[0007] 3) Design a risk - aware scheduling framework for the multi - energy microgrid based on the Markov decision process;
[0008] 4) Obtain the current real climate scenario of the multi - energy microgrid and determine the state - transition function corresponding to the current real climate scenario according to the state - transition function matrix;
[0009] 5) According to the state - transition function corresponding to the current real climate scenario, use the LLM - DPSAC algorithm to determine the optimization objective in the risk - aware scheduling framework of the multi - energy microgrid;
[0010] 6) Based on the optimization objective, determine and execute the action A' corresponding to the new state S', realizing the risk - aware scheduling of the multi - energy microgrid.
[0011] Furthermore, set the Markov decision process MDP=(S, A, P, P0, R, γ, T); where S is the state space of the multi - energy microgrid, A is the action space of the multi - energy microgrid, P is the state - transition function, P0(S) is the initial state distribution, R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0012] Furthermore, the LLM is trained through historical meteorological data, energy production records, and load demand data.
[0013] Furthermore, the input of the LLM includes the state S and the action A, and the output is the new state S'.
[0014] Furthermore, in step 2), the state - transition function matrix is used to determine the probability of transitioning from state S to state S' under a specific action A.
[0015] Furthermore, in step 3), the risk - aware scheduling framework of the multi - energy microgrid is defined as (S, A, P, P0, R, γ, T), where S is the state space of the multi - energy microgrid, A is the action space of the multi - energy microgrid, P is the state - transition function, P0(S) is the initial state distribution, R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0016] Furthermore, in step 5), the steps of using the LLM - DPSAC algorithm to determine the optimization objective in the risk - aware scheduling framework of the multi - energy microgrid include:
[0017] 5.1) Define the set of candidate utility functions f θ = {f θ1 , f θ2 ,..., f θM}; f θM is the conditional risk value C M under the automatic adjustment factor β obj ;
[0018] 5.2) Define the learning objective J(θ θi ) for the given utility function f i ), that is:
[0019] J(θ i ) = αJ value (θ i ) + (1 - α)J grad (θ i )(1)
[0020] In the formula, J value (θ i ) is the value network, J grad (θ i ) is the policy network, and α is the correlation factor;
[0021] 5.3) On the randomized policy π*, train the optimal policy using the SAC algorithm, that is:
[0022]
[0023] In the formula, is the energy consumption; R(s t , a t ) is the reward; ρ π is the parameter of the policy π*; β is the adjustment factor;
[0024] Among them, the information entropy H(π·|s t ) is as follows:
[0025] H(π(·|s i )) = -∑π(·|s i ) log(π(·|s i ))(3)
[0026] 5.4) Construct the Bellman equation containing the maximum entropy, that is:
[0027]
[0028] In the formula, V(s t ) is the state value function, and γ is the reward discount factor;
[0029] 5.5) Based on the Bellman equation, the optimal policy π is constrained within a specific set π by minimizing the KL divergence, i.e.: new That is:
[0030]
[0031] In the formula represents the constant factor used to normalize the Q distribution; represents the value function under the previous policy; D KL is the KL divergence;
[0032] 5.6) Select the action A from the policy π according to the state S, and update the parameters of the Actor network and the Critic network. new Specifically, the parameter update methods of the Actor network and the Critic network are as follows:
[0033] Furthermore, the parameter update methods of the Actor network and the Critic network are as follows:
[0034]
[0035] Among them, and θ are network parameters; E is the energy consumption; R is the reward; is the state value function; Z is the constant factor; Q is the value function.
[0036] The technical effect of the present invention is beyond doubt. By combining the large language model distributed Pareto Soft Actor-Critic (LLM-DPSAC) algorithm considering CVaR, the limitations of traditional methods in economic benefit optimization are solved. Specifically, the LLM can simulate and generate realistic ocean climate scenarios, and dynamically generate the state transition function P(s'|s,a) based on real-time data changes, which effectively overcomes the limitations of traditional physical model-based and statistical methods. Through this method, the system can more accurately capture the dynamic changes in the complex environment, significantly improving the modeling ability of the real environment and the ability to flexibly respond to changes.
[0037] In addition, the present invention further introduces CVaR to address low-probability high-consequence events, especially the risks of extreme weather and sudden surges in demand. This improvement enables the optimization model in this paper to consider the balance between economic benefits and risk management, and establishes the corresponding objective function to ensure robustness and effectiveness in extreme situations.
[0038] Most importantly, the LLM-DPSAC algorithm models the distribution of policy returns directly, taking into account both the average expected value and effectively capturing the diverse policy choices of agents with different risk preferences. Through this innovative algorithm, the present invention not only improves the modeling ability of risk sensitivity but also provides personalized optimization strategies for different types of agents, enabling flexible adaptation to changing environments and task requirements.
[0039] In summary, the present invention combines LLM with distributed Pareto optimization to address numerous challenges in the prior art when facing complex and uncertain environments. Experimental results demonstrate that the proposed model and algorithm outperform traditional methods in terms of performance, showing higher efficiency and stronger robustness in the scheduling of multi-energy microgrids under complex climate conditions and risk-sensitive environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flow diagram of the process for generating typical weather characteristics using the large language model of the present invention;
[0041] Figure 2 is the algorithm flowchart of LLM-DPSAC of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject matter scope of the present invention is limited to the following embodiments. Without departing from the above technical ideas of the present invention, various substitutions and modifications made according to ordinary technical knowledge and customary means in the art should be included within the protection scope of the present invention.
[0043] Embodiment 1:
[0044] Refer to Figures 1 to 2 , a risk-aware scheduling method for a multi-energy microgrid based on large language model-driven distributed Pareto reinforcement learning, comprising the following steps:
[0045] 1) Use the LLM to simulate and generate real climate scenarios of the multi-energy microgrid;
[0046] 2) Update the state transition function P(s'|s,a) corresponding to each real climate scenario to construct a state transition function matrix;
[0047] 3) Design a risk-aware scheduling framework for the multi-energy microgrid based on the Markov decision process;
[0048] 4) Obtain the current real climate scenario of the multi-energy microgrid and determine the state transition function corresponding to the current real climate scenario according to the state transition function matrix;
[0049] 5) Determine the optimization objective in the risk-aware scheduling framework of the multi-energy microgrid using the LLM-DPSAC algorithm according to the state transition function corresponding to the current real climate scenario;
[0050] 6) Based on the optimization objective, determine and execute the action A' corresponding to the new state S' to achieve the risk-aware scheduling of the multi-energy microgrid.
[0051] Set the Markov decision process MDP = (S, A, P, P0, R, γ, T); where S is the state space of the multi-energy microgrid, A is the action space of the multi-energy microgrid, P is the state transition function, P0(S) is the initial state distribution, and R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0052] The LLM is trained using historical meteorological data, energy production records, and load demand data.
[0053] The input of the LLM includes the state S and the action A, and the output is the new state S'.
[0054] In step 2), the state transition function matrix is used to determine the probability of transitioning from state S to state S' under a specific action A.
[0055] In step 3), the risk-aware scheduling framework of the multi-energy microgrid is defined as (S, A, P, P0, R, γ, T), where S is the state space of the multi-energy microgrid, A is the action space of the multi-energy microgrid, P is the state transition function, P0(S) is the initial state distribution, and R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0056] In step 5), the steps of determining the optimization objective in the risk-aware scheduling framework of the multi-energy microgrid using the LLM-DPSAC algorithm include:
[0057] 5.1) Define the set of candidate utility functions f θ = {f θ1 , f θ2 ,..., f θM}; f θM is the conditional risk value C M under the automatic adjustment factor β obj ;
[0058] 5.2) Define the learning objective J(θ θi ) for a given utility function f i , that is:
[0059] J(θ i ) = αJ value (θ i ) + (1 - α)Jgrad (θ i )(1)
[0060] In the formula, J value (θ i ) is the value network, J grad (θ i ) is the policy network, and α is the correlation factor;
[0061] 5.3) On the randomized policy π*, use the SAC algorithm to train the optimal policy, that is:
[0062]
[0063] In the formula, is the energy consumption; R(s t , a t ) is the reward; ρ π is the parameter of the policy π*; β is the adjustment factor;
[0064] Among them, the entropy H(π·|s t ) is as follows:
[0065] H(π(·|s i )) = -∑π(·|s i ) log(π(·|s i ))(3)
[0066] 5.4) Construct the Bellman equation including the maximum entropy, that is:
[0067]
[0068] In the formula, V(s t ) is the state value function, and γ is the reward discount factor;
[0069] 5.5) Based on the Bellman equation, constrain the optimal policy π within a specific set π new by minimizing the KL divergence, that is:
[0070]
[0071] In the formula represents the constant factor used to normalize the Q distribution; represents the value function under the previous policy; D KL is the KL divergence;
[0072] 5.6) Select the action A from the policy π new according to the state S, and update the parameters of the Actor network and the Critic network.
[0073] The parameter update methods of the Actor network and the Critic network are as follows:
[0074]
[0075] Among them, and θ are network parameters; E is energy consumption; R is the reward; is the state value function; Z is a constant factor; Q is the value function.
[0076] Example 2:
[0077] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, including the following steps:
[0078] 1) Use the LLM to simulate and generate real climate scenarios of the multi-energy microgrid;
[0079] 2) Update the state transition function P(s'|s,a) corresponding to each real climate scenario, thereby constructing a state transition function matrix;
[0080] 3) Design a risk-aware scheduling framework for the multi-energy microgrid based on the Markov decision process;
[0081] 4) Obtain the current real climate scenario of the multi-energy microgrid, and determine the state transition function corresponding to the current real climate scenario according to the state transition function matrix;
[0082] 5) According to the state transition function corresponding to the current real climate scenario, use the LLM-DPSAC algorithm to determine the optimization objective in the risk-aware scheduling framework of the multi-energy microgrid;
[0083] 6) Based on the optimization objective, determine and execute the action A' corresponding to the new state S', and realize the risk-aware scheduling of the multi-energy microgrid.
[0084] Example 3:
[0085] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as that of Example 2. Further, set the Markov decision process MDP=(S, A, P, P0, R, γ, T); where S is the state space of the multi-energy microgrid, A is the action space of the multi-energy microgrid, P is the state transition function, P0(S) is the initial state distribution, R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0086] Example 4:
[0087] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-3. Further, the LLM is trained with historical meteorological data, energy production records, and load demand data.
[0088] Embodiment 5:
[0089] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-4. Further, the input of the LLM includes state S and action A, and the output is the new state S'.
[0090] Embodiment 6:
[0091] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-5. Further, in step 2), the state transition function matrix is used to determine the possibility of transitioning from state S to state S' under a specific action A.
[0092] Embodiment 7:
[0093] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-6. Further, in step 3), the multi-energy microgrid risk-aware scheduling framework is defined as (S, A, P, P0, R, γ, T), where S is the state space of the multi-energy microgrid, A is the action space of the multi-energy microgrid, P is the state transition function, P0(S) is the initial state distribution, R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0094] Embodiment 8:
[0095] A risk-aware scheduling method for a multi-energy microgrid based on large language model-based distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-7. Further, in step 5), the steps of using the LLM-DPSAC algorithm to determine the optimization objective in the multi-energy microgrid risk-aware scheduling framework include:
[0096] 5.1) Define the set of candidate utility functions f θ = {f θ1 , f θ2 ,..., f θM}; f θM is the conditional risk value C M under the automatic adjustment factor β obj ;
[0097] 5.2) Define the given utility function f θiThe learning objective J(θ i ), that is:
[0098] J(θ i ) = αJ value (θ i )+(1 - α)J grad (θ i )(1)
[0099] In the formula, J value (θ i ) is the value network, J grad (θ i ) is the policy network, and α is the correlation factor;
[0100] 5.3) On the randomized policy π*, use the SAC algorithm to train the optimal policy, that is:
[0101]
[0102] In the formula, E (s,at)~ρπ is the energy consumption; R(s t , a t ) is the reward; ρ π is the parameter of the policy π*;
[0103] Among them, the information entropy H(π·|s t ) is as follows:
[0104] H(π(·|s i )) = -∑π(·|s i )log(π(·|s i ))(3)
[0105] 5.4) Construct the Bellman equation containing the maximum entropy, that is:
[0106]
[0107] In the formula, V(s t ) is the state value function, and γ is the reward discount factor;
[0108] 5.5) Based on the Bellman equation, constrain the optimal policy π within a specific set by minimizing the KL divergence, that is:
[0109]
[0110] In the formula represents the constant factor used to normalize the Q distribution; represents the value function under the previous policy;
[0111] 5.6) According to the state S from the policy π newSelect action A and update the parameters of the Actor network and the Critic network.
[0112] Embodiment 9:
[0113] A risk-aware scheduling method for a multi-energy microgrid based on large language model distributed Pareto reinforcement learning, the technical content is the same as any one of Embodiments 2-8. Further, the parameter update methods of the Actor network and the Critic network are as follows:
[0114]
[0115] Wherein, and θ are network parameters.
[0116] Embodiment 10:
[0117] A large language model - distributed Pareto Soft Actor-Critic algorithm (LLM-DPSAC) is used for optimal scheduling of a multi-energy microgrid, as Figure 2 shown. Specifically, the LLM is used to simulate and generate real climate scenarios to construct the dynamic changes in the real environment, as Figure 1 shown. In addition, a set of non-linear utility functions are learned to determine the optimization objective and change the Markov decision process (MDP). The actor network generates an action based on the currently observed operating state s t of the multi-energy microgrid, and then the agent transmits the state after this action to obtain the state s t+1 at the next instant. The agent samples from the new energy output and load probability distributions, and then selects the action executed by s * according to the policy π t+1 and stores the returned response. During the entire training phase, the error is backpropagated to update the network parameters. This loop continues to enable the agent to adapt to the learning process, where the reward value gradually increases and stabilizes.
[0118] First, set the MDP of the LLM-DPSAC algorithm. The MDP is defined as (S, A, P, P0, R, γ, T), where S is the state space of the multi-energy microgrid, A is the action space of the multi-energy microgrid, P is the state transition function, P0(S) is the initial state distribution, R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
[0119] The specific form of the state transition function P(s'|s,a) depends on the physical characteristics of the environment, such as weather changes and fluctuations in energy demand. In traditional models, these transition probabilities are usually defined based on physical laws or statistical models, but these methods may not fully capture the dynamic changes in real scenarios, especially when the environment is highly uncertain and variable.
[0120] LLMs have powerful generation capabilities, providing new solutions for the scheduling optimization of multi - energy microgrids. The advantage of LLMs lies in their ability to construct a more detailed and realistic environment, reflecting the uncertainty and variability of energy resources and demands in the real world. By generating scenarios consistent with the current environmental conditions based on historical data input, LLMs help decision - makers make more flexible decisions in a highly complex environment. By generating environmental descriptions that truly reflect dynamic changes, the system can not only predict common environmental changes but also adapt to sudden extreme conditions. This adaptive environment - building ability provides a more flexible and robust response strategy for the microgrid scheduling system. With the help of the LLM, the structural framework of the state transition function is generated as Figure 1 shown.
[0121] 1. Data collection and pre - processing: Collect a dataset including historical meteorological data, energy production records, and load demand data. Use this data to train the LLM so that the LLM can learn and understand the relationships between various inputs (state S and action A) and outputs (new state S'). This enables the LLM to master how various factors (such as weather conditions, changes in energy supply, and load fluctuations) jointly affect the change of the system state.
[0122] 2. Scenario generation: Use the trained LLM to generate various possible environmental scenarios. These scenarios illustrate the potential new states that the system may transition to under specific state S and action A conditions. Compared with traditional methods, LLMs can construct a more detailed and realistic environment, reflecting the uncertainty and temporal variability of energy resources and demands in the real world.
[0123] 3. Probability assessment and update: For each generated scenario, evaluate the probability of transitioning to the new state. Update the state transition function P(s'|s,a) according to the output of the LLM. This step involves integrating all possible scenarios generated by the LLM and their respective probabilities to form a comprehensive and updated state transition function matrix. This matrix is used to determine the likelihood of transitioning from state S to state S' under a specific action A.
[0124] 4. Decision - making simulation and system feedback: Use the updated transition probabilities to conduct decision - making simulations within the microgrid management system, observe the results, and make adjustments according to the actual operation of the system. This includes further updating and optimizing the uncertainty model based on new data to more accurately simulate and predict the uncertainty changes occurring in the real environment.
[0125] The main idea of LLM-DPSAC is to learn a set of non-linear utility functions to help the agent discover the optimal policy and thus find the optimal solution for multiple objectives. By learning a series of different and plausible utility functions, the agent is guided towards the optimal policy. The utility function is used to quantify the degree to which the return distribution of the policy meets the user's preferences. In DRL and distributed optimization, the role of the utility function is to convert complex multi-dimensional returns into scalars for easy comparison and optimization. In this algorithm, the utility function plays two key roles:
[0126] 1. Measuring the quality of the distribution: converting the complex return distribution into a comparable scalar;
[0127] 2. Guiding policy optimization: determining the optimization objective through the utility function.
[0128] Specifically, define f θ1 , f θ2 , …, f θM as the set of candidate utility functions, parameterized as θ1, …, θ M , where f θM is defined as C obj under different β. For any given utility function f θi , the learning objective is defined as:
[0129] J(θ i ) = αJ value (θ i ) + (1 - α)J grad (θ i )
[0130] Optimizing this objective function can prompt the generated utility functions to exhibit greater diversity in values and derivatives within the normalized range [0, 1] K , thus covering potential user preferences more comprehensively. In non-decreasing neural networks, gradient descent is used to minimize the objective function to generate N sets of utility functions to optimize the policy to maximize the expected utility. It effectively balances the trade-off between multiple objectives and distributions and finally generates a policy more likely to be consistent with user preferences. In short, the state space is expanded using cumulative rewards, and multi-dimensional rewards are converted into scalar rewards through the differences in utility functions.
[0131] Soft Actor-Critic (SAC) is a continuous action space reinforcement learning algorithm based on the maximum entropy theory. By maximizing the expected reward of the environment and the maximum entropy of the policy, the SAC algorithm achieves intelligent decision-making and learning in the continuous action space while maintaining exploration. SAC enhances the robustness of the algorithm and is able to explore better communication scheduling strategies. The optimal policy of the SAC algorithm is trained on the randomized policy policyπ*, and its expression is:
[0132]
[0133] The information entropy H(π·|s t ) combines the cumulative reward with the entropy regularization component of the objective function, enabling the agent to thoroughly explore the action space during training to maximize the reward, as shown in the following formula:
[0134] H(π(t|s i ))=-∑π(·|s i )log(π(·|s i ))
[0135] The Bellman equation containing the maximum entropy is shown in the following formula:
[0136]
[0137] where V(s t ) is the state value function and γ is the reward discount factor. To simplify management and improve returns, when optimizing and updating parameters, these policies are constrained within a specific set by minimizing the KL divergence, as shown in the following formula:
[0138]
[0139] where represents the constant factor used to normalize the Q distribution. represents the value function under the previous policy. A deep neural network is used to represent the q value function and the policy distribution. The actor network serves as the policy network:
[0140]
[0141] where and θ are the network parameters. The factor β is automatically adjusted. Finally, the Adam algorithm is used to optimize the parameters through gradient descent to obtain the update of the network parameters. When the number of iterations reaches the predefined threshold or the optimal decisions obtained from consecutive iterations are the same and the update increment of the value function is less than the predefined threshold, the algorithm terminates. The optimized policy and the value function used to evaluate and optimize the policy are output.
[0142] The LLM-DPSAC algorithm first learns a set of non-linear utility functions, determines the optimization objective, and modifies the MDP to provide a basis for optimization. Then, the SAC algorithm is trained and combined with an entropy-driven evolutionary algorithm to obtain the weight distribution of the optimization problem. Based on the LLM-DPSAC algorithm, the distribution Pareto frontier between economic benefits and CvaR can be effectively approximated.
Claims
1. A risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of large language models, characterized in that, Including the following steps: 1) Use the LLM to simulate and generate real climate scenarios for the multi - energy microgrid. 2) Update the state - transition function P(s'|s,a) corresponding to each real climate scenario, thereby constructing a state - transition function matrix; 3) Design a risk - aware scheduling framework for the multi - energy microgrid based on the Markov decision process; 4) Obtain the current real climate scenario of the multi - energy microgrid, and determine the state - transition function corresponding to the current real climate scenario according to the state - transition function matrix; 5) According to the state - transition function corresponding to the current real climate scenario, use the LLM - DPSAC algorithm to determine the optimization objective in the risk - aware scheduling framework of the multi - energy microgrid; 6) Based on the optimization objective, determine and execute the action A' corresponding to the new state S', realizing the risk - aware scheduling of the multi - energy microgrid.
2. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, characterized in that, Set the Markov decision process MDP = (S, A, P, P0, R, γ, T); where S is the state space of the multi - energy micro - grid, A is the action space of the multi - energy micro - grid, P is the state transition function, P0(S) is the initial state distribution, and R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
3. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, characterized in that, The LLM is trained with historical meteorological data, energy production records, and load demand data.
4. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, characterized in that, The input of the LLM includes the state S and the action A, and the output is the new state S'.
5. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, wherein In step 2), the state - transition function matrix is used to determine the probability of transitioning from state S to state S' under a specific action A.
6. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, characterized in that, In step 3), the risk perception scheduling framework of the multi - energy micro - grid is defined as (S, A, P, P0, R, γ, T), where S is the state space of the multi - energy micro - grid, A is the action space of the multi - energy micro - grid, P is the state transition function, P0(S) is the initial state distribution, and R is the objective function C obj , γ is the discount factor, and T is the total number of time steps.
7. The risk-aware scheduling method for a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 1, characterized in that, In step 5), the steps of using the LLM - DPSAC algorithm to determine the optimization objective in the risk - aware scheduling framework of the multi - energy microgrid include: 5.1) Define the set of candidate utility functions f θ = {f θ1 , f θ2 ,..., f θM}; f θM is the objective function C M under the automatic adjustment factor β obj ; 5.2) Define the learning objective J(θ θi ) for a given utility function f i ), i.e.: J(θ i ) = αJ value (θ i )+(1 - α)J grad (θ i )(1) where J value (θ i ) is the value network, J grad (θ i ) is the policy network, and α is the correlation factor; 5.3) On the randomized policy π*, use the SAC algorithm to train the optimal policy, that is: In the formula, is the energy consumption; R(s t ,a t ) = Cobj is the reward function; ρ π is the parameter of the policy π*; β is the adjustment factor; Among them, the information entropy H(π·|s t ) is as follows: H(π(·|s i )) = -∑π(·|s i ) log(π(·|s i )) (3) 5.4) Construct a Bellman equation containing the maximum entropy, that is: where V(s t ) is the state value function and γ is the reward discount factor; 5.5) Based on the Bellman equation, the optimal policy π is constrained within a specific set π by minimizing the KL divergence, i.e.: new where: wherein, represents a constant factor for normalizing the Q distribution; represents the value function under the previous policy; D KL is the KL divergence; 5.6) Select action A according to state S from policy π new and update the parameters of the Actor network and the Critic network.
8. The risk-aware scheduling method of a multi-energy microgrid based on distributed Pareto reinforcement learning of a large language model according to claim 7, characterized in that, The parameter update methods of the Actor network and the Critic network are as follows: Among them, and θ are network parameters; E is the energy consumption; R is the reward; V θ (s t+1 ) is the state value function; Z is a constant factor; Q is the value function.
Citation Information
Cited By
Micro-grid energy management and control system based on large language model and deep reinforcement learning
CN122203443A
Large model guided multi-energy microgrid risk scheduling method and system
CN122371352A
A large model guided multi-energy microgrid risk scheduling method and system
CN122371352B