Multi-element load hierarchical collaborative scheduling method based on multi-agent game learning

By building a multi-load energy interaction framework and Stackelberg-Nash game model, combined with the Nash game-MADDPG algorithm, the problem of lack of a reasonable framework and privacy security in multi-load coordinated scheduling is solved, and the interests balance of multiple subjects and user response are achieved, and the efficiency and security of coordinated scheduling are improved.

CN120601441APending Publication Date: 2025-09-05NORTH CHINA ELECTRIC POWER UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444143.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing multi-load collaborative scheduling methods lack a reasonable collaborative scheduling framework and efficient privacy and security scheduling algorithms, making it difficult to achieve global optimization of multi-subjects and accuracy of user-side response. In addition, deep reinforcement learning is inefficient in learning and unstable training under multi-subject conditions.

Method used

A multi-load stratified and hierarchical collaborative scheduling method based on multi-subject game learning is constructed. By constructing a multi-load energy interaction framework, considering economic incentives and user psychology, the Stackelberg-Nash game model and Nash game-MADDPG algorithm are adopted, and combined with KKT conditions and an experience replay mechanism that enhances memory, the multi-subject interests balance and optimal cooperation between UA is achieved.

Benefits of technology

It has achieved balanced interests of multiple subjects and mutual benefit and win-win results, improved the accuracy and privacy and security of coordinated scheduling of multiple loads, reduced the complexity of model calculations, suppressed the impact of uncertainty, and improved the efficiency of user-side demand response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120601441A_ABST
    Figure CN120601441A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-element load hierarchical collaborative scheduling method based on multi-agent game learning, and belongs to the technical field of telecommunication. UA is introduced, and a basic framework of multi-element load energy interaction is established. The economic incentive level and the psychology of the user are considered, a user response comfort model is constructed, and the willingness that the user side multi-element load participates in demand response is accurately reflected. Decision-making models of micro-grid operators and UA are respectively constructed, and based on the game theory, a multi-element load hierarchical collaborative scheduling problem is converted into a game problem, so that multi-subject benefit balance and mutual benefit and win-win are realized. Based on the KKT condition, the upper layer game problem is solved, and the model calculation complexity is reduced. An upper layer game plan is transmitted to a lower layer game model, based on an actual source load uncertainty state, an algorithm is proposed to solve a Nash game problem and stabilize uncertainty influence, a transaction result with optimal cooperation between UAs is obtained, and Nash equilibrium of UA alliances is efficiently solved by introducing cooperation cost and an experience playback mechanism based on enhanced memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning, and belongs to the technical field of electrical communications. Background Art

[0002] With the widespread integration of new loads such as electric vehicles and distributed power sources, as well as the development of technologies such as cogeneration and user-side management, multi-load collaborative scheduling technologies have emerged. These technologies often rely on regional energy systems such as industrial, commercial, and residential areas. Through source-load-storage collaborative scheduling and user-side demand response, they improve economic efficiency and promote the widespread adoption of new energy. However, the complex inter-operability relationships among multi-loads place higher demands on the fairness, real-time nature, and privacy of multi-agent collaborative scheduling. The volatility of distributed resource output and the randomness of user-side demand response also pose challenges to the safe operation of distribution networks. Therefore, there is an urgent need to research real-time, accurate, and secure multi-load collaborative scheduling methods.

[0003] Existing multi-layered coordinated scheduling methods suffer from two major problems: First, there is a lack of a reasonable coordinated scheduling framework. Existing research often adopts a single master-slave game or cooperative game framework. The master-slave game framework prioritizes the interests of the upper layer, ignoring the cooperative ability of lower-level entities, making it difficult to achieve a global optimal solution for multiple entities. The cooperative game framework causes the lower layer to lose bargaining power and be unable to influence upper-level decisions. Furthermore, both frameworks ignore user-side response comfort, affecting the accuracy of demand response. Second, there is a lack of efficient scheduling algorithms that balance performance advantages and privacy security. Common model-driven methods require that the parameters of the source, load, and storage models are known, accurate, and fixed. This cannot adapt to the random changes in source and load, and is difficult to adapt to multi-load layered and graded coordinated scheduling scenarios with complex user response behaviors and a high proportion of renewable energy access. Deep reinforcement learning (DRL) can learn optimal strategies through interaction with the environment without global information, but it still struggles to meet the privacy requirements of multiple entities and suffers from low learning efficiency and unstable training.

[0004] In view of the above-mentioned defects, the present invention aims to create a multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning, so as to make it more valuable for industrial utilization. Summary of the Invention

[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning.

[0006] The present invention provides a multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning, and the specific scheduling steps are as follows:

[0007] First, a multi-load energy interaction framework including UA is constructed;

[0008] Based on the above multi-load energy interaction framework, and taking into account the economic incentive level and user psychology, a user load response comfort model with multi-dimensional characteristics is constructed to accurately reflect the willingness of multiple loads on the user side to participate in demand response.

[0009] At the same time, the MO decision model and UA decision model are constructed respectively, and based on game theory, a multi-load Stackelberg-Nash game model is constructed to achieve multi-subject interest balance and mutual benefit and win-win results;

[0010] Based on the KKT condition, the upper-level Stackelberg game problem is solved to reduce the computational complexity of the model;

[0011] The upper-level game plan is transferred to the lower-level Nash game model. Based on the actual source-load uncertainty state, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smooth out the impact of uncertainty, and obtain the optimal transaction result for cooperation among UAs. By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved, ultimately achieving multi-layered and hierarchical collaborative scheduling of multiple loads.

[0012] Furthermore, the multi-load energy interaction framework including UA includes: the upper-level power grid and gas company are the sources of electricity and gas, the microgrid operator (MO) realizes electricity and heat coordination through combined heat and power (CHP) units and gas boilers (GB) to meet the electricity and heat needs of massive users; the microgrid operator manages electricity and heat storage, as well as power-to-gas (P2G) and carbon capture (CCS) equipment to reduce carbon emissions;

[0013] The controllable resources on the user side include household photovoltaics and residential electric loads and thermal loads that can participate in demand response. It is defined that there are I UAs in the multi-load microgrid, and the set is U={1,…,UA i ,…UA I}, each UA i Manage N in a specific area i users; the set of control time periods T = {1,…,t,…T}.

[0014] Furthermore, the user load response comfort model taking into account multi-dimensional characteristics is expressed as:

[0015]

[0016] Where, are the comfort cost coefficients of electrical load and thermal load respectively; UA iThe economic incentive level EIL, incentive timing characteristics ITC and user herd mentality UCP of electric load, and the same applies to thermal load; Can be expressed as:

[0017]

[0018] Where, The energy price of the current period is used to express it. MO is UA for period t i The electricity sales price set; Consider the current period's energy price relative to the overall period's price; Affected by the number of users participating in demand response in previous periods, UA for period t-1 i The number of users participating in electricity load demand response within the control area;

[0019] Similarly, Expressed as

[0020]

[0021] Where, is the selling price of MO during period t; UA for period t-1 i The number of users participating in heat load demand response within the control area;

[0022] The number of users participating in demand response can be expressed as:

[0023]

[0024] Where, Indicates UA i The load response variable of the kth user in the tth period is: indicates participation in demand response, whereas non-participation is indicated. p(·) represents probability, where p0 is the probability of not participating in demand response in the current period due to the influence of herd mentality given that the participant participated in demand response in the previous period. p1 is the probability of not participating in demand response in the previous period due to the influence of herd mentality given that the participant participated in demand response in the current period.

[0025] Furthermore, the MO decision model is expressed as:

[0026]

[0027] Where R ECO (t) is the economic benefit of MO, C LC (t) is the low carbon cost, C COM (t) is the comfort cost; R ECO (t) can be expressed as

[0028]

[0029] Where C b (t) is the time period from the upper power grid and UA to MO. i The cost of purchasing electricity, R s (t) is the MO's report to the upper grid and UA i Revenue from electricity sales, UA i Energy storage loss costs, new energy grid access expenses and CHP operating costs are expressed as follows:

[0030]

[0031] C CHP (t) = λ CHP P CHP (t)+λ P2G P P2G (t)+λ CCS P CCS (t) (11)

[0032] Where, Δt is the length of the control period; are the grid electricity price and on-grid electricity price in period t, MO to UA i The electricity price for users, Hot selling price; It is the purchase and sale of electricity between MO and the upper power grid; UA i Electricity purchase amount, UA i Purchased heat; are the charging and discharging power of electrical and thermal energy storage respectively; E ,λ H is the cost coefficient of electricity and heat storage; PV (t) is the household photovoltaic grid-connected electricity price paid by MO to users, using the full power generation subsidy model, P i PV (t) is the household photovoltaic output power; gas (t) is the gas purchase cost per unit of gas, V gas (t) = V CHP (t)+V GB (t)-V P2G (t) is the gas usage rate, where V CHP (t), V GB (t) is the gas consumption rate of CHP and GB, V P2G (t) is the rate at which the P2G device generates gas; P CHP (t), P P2G (t), PCCS (t) is the discharge power of CHP, P2G and CCS, λ CHP ,λ P2G ,λ CCS is the operation and maintenance cost coefficient of CHP, P2G and CCS;

[0033] For low carbon cost C LC (t), since the current domestic carbon quota allocation method mainly adopts free allocation, UA i The carbon emission quota E*(t) is expressed as

[0034]

[0035] Where κ is the regional carbon emission per unit electricity, which is 590t / (kW·h) here; is the electricity conversion coefficient, H CHP (t) is the thermal power of the CHP unit;

[0036] Considering that actual carbon emissions mostly come from CHP and GB units, and the power purchased by the upper power grid mainly comes from coal-fired power units, the actual carbon emissions of the system during period t are

[0037]

[0038] Where, E(t) is the actual total carbon emissions of MO during period t, E b (t) is the carbon emissions generated by purchasing electricity from the upper power grid, E CHP (t), E GB (t), E CCS (t) is the actual carbon emissions of CHP, GB, and CCS units; η is the carbon emission coefficient of coal-fired power units, κ CHP , κ GB is the carbon emission coefficient of CHP and gas turbine, which is 591t / (kW·h); η CCS is the energy consumption of CCS device to process CO2, η CCS =0.269MW·h / t; thus, the low-carbon cost can be expressed as

[0039] C LC (t) = λ LC (E(t)-E * (t))(14)

[0040] Where λ LC is the carbon trading price coefficient;

[0041] C COM (t) is the comfort cost caused by the change of energy consumption habits when the user side load participates in demand response, which can be expressed as

[0042]

[0043] Where, P i CL (t), UA for period t i Response amount of electrical load and thermal load; is the controllable load response type weight, where For electrical and thermal loads, when UA i When the controllable load is in a reduction-type response, the weight is 1; otherwise, when the controllable load is in an absorption-type response, the weight is [0, 1]. is the set of response comfort cost coefficients of controllable loads; UA i Minimum incentive prices for electrical and thermal loads; is the response type indicator parameter; ⊙ is the Hadamard product, which means the multiplication of corresponding matrix elements.

[0044] Furthermore, the UA decision model takes minimizing the cost of the jurisdiction as the optimization goal, which includes the energy purchase cost, comfort cost, photovoltaic grid connection income, and inter-UA transaction income. The optimization goal is expressed as:

[0045]

[0046] Where, The benefits of energy interaction between users and aggregators, UA i The energy purchase cost is expressed as:

[0047]

[0048] Where λ i→j (t) is UA i With UA j The energy transaction price, P i i→j (t) is UA i To UA j Output power.

[0049] Furthermore, the first stage of the multi-load Stackelberg-Nash game model is the Stackelberg game between the MO and each UA. The MO sets electricity and heat prices based on the differentiated characteristics of each UA to achieve the goal of maximizing MO profits. The second stage is the Nash game of the UA alliance. Each UA minimizes the alliance cost through cooperation among UAs based on the energy price issued by the upper-level MO and the current source and load status.

[0050] Furthermore, the method for solving the upper-level Stackelberg game problem based on the KKT condition is:

[0051] Since the decision variables of the upper-level model can be regarded as constants in the lower-level model, and the lower-level model satisfies the Slater condition, the Lagrangian function of the lower-level model is constructed. For the existence of bilinear terms and complementary slack conditions, the strong duality theory is adopted. Based on the Big M method and McCormick envelope process, the original problem is transformed into a single-level mixed integer linear programming problem that is easy to solve, thereby converting the lower-level problem into the optimality condition of the upper-level problem.

[0052] Furthermore, the steps of proposing the Nash game-MADDPG algorithm to solve the Nash game problem are:

[0053] (2) Construction of Markov decision process model

[0054] The Markov decision process model of the multi-load Nash game mainly includes the state space, action space and real-time reward function considering the coordination cost;

[0055] 1) State space: Define S(t) = {O1(t), ..., O i (t),…,O I (t)}, where O i (t) is UA i The state space of is expressed as:

[0056]

[0057] Where, P i PV,act (t) is the UA of period t i The actual output of household photovoltaics;

[0058] 2) Action Space

[0059] To improve computational efficiency, P i→j (t) is converted to

[0060]

[0061] Where, X i→j (t) is UA i Power interaction between UAs generated by intelligent agents; The upper limit of power interaction between UAs;

[0062] Assume that UA i The action space is Expressed as

[0063]

[0064] 3) Real-time reward function

[0065] In order to improve the training efficiency of local agents, this paper introduces the agent collaboration cost. The real-time reward function is expressed as

[0066]

[0067] Where, UA i The collaboration cost is related to the strategies of other agents. is the collaboration cost coefficient between UAs. Collaboration cost allows agents to have a global awareness and consider the benefits of other agents while optimizing their own benefits;

[0068] (2) Nash Game-MADDPG Algorithm Solution Process

[0069] The Nash game-MADDPG algorithm adopts the actor-critic network structure. The MO side sets up a centralized main Q network and target Q network. Based on the global state information S(t), it guides the UA to learn a strategy that approaches the global Nash equilibrium. Each UA is configured with two policy networks, which rely on the local observation state O i (t) Generate its own strategy and execute it in a distributed manner. Each UA does not exchange its own observation state. At the same time, an experience replay pool that aggregates global information is set up to assist in the training of the agent. An experience replay mechanism based on enhanced memory is proposed. The network settings of the NASH game-MADDPG algorithm are as follows: UA i The main policy network is represented as The centralized main Q network on the MO side is expressed as The corresponding target networks are φ i ′, Q i ′, the network parameters are The process of solving Nash equilibrium based on NASH game-MADDPG is as follows

[0070] 1) UA i The agent generates a strategy based on the behaviorpolicyβ strategy

[0071] 2) Each UA calculates its own real-time reward function r i (t), the state is transferred to O i (t+1), the experience chain Upload to MO and store in the experience replay pool Φ;

[0072] 3) MO calculates the collaboration cost of each UA and asynchronously updates the Q network corresponding to each UA in the centralized main Q network based on the global environment state. The coefficient update formula is:

[0073]

[0074] Where Q i ′(S(t+1),π(t+1)) is the corresponding UA generated by the target Q network in state S(t+1) i The Q value is the cumulative return of all UAs adopting the Nash equilibrium strategy starting from S(t+1), in the form of Pareto optimality;

[0075] Then, samples ||Φ(t)|| are extracted from Φ to update the main Q network parameters. An experience replay mechanism based on enhanced memory is proposed and measured by TD-error, which can be expressed as:

[0076]

[0077] The selection probability of different experience chains is expressed as

[0078]

[0079] Where ε is a small random number to avoid the situation where the TD-error is close to 0 and cannot be selected. The update formula of the centralized main Q network is:

[0080]

[0081] 4) Each UA extracts ||Φ from the MO's experience replay pool i (t)|| experience chains are used for training, and the parameter update process of the main policy network is expressed as:

[0082]

[0083] 5) Update the target Q network and target policy network, the formula is

[0084]

[0085] Where, is the weight of the soft update parameter.

[0086] A multi-load hierarchical and graded collaborative scheduling device based on multi-agent game learning, comprising:

[0087] Framework construction module: used to build a multi-load energy interaction framework including UA;

[0088] Modeling module: Based on the above multi-load energy interaction framework, the module considers economic incentives and user psychology, builds a user load response comfort model that takes into account multi-dimensional characteristics, and accurately reflects the willingness of multiple loads on the user side to participate in demand response.

[0089] At the same time, the MO decision model and UA decision model are constructed respectively, and based on game theory, a multi-load Stackelberg-Nash game model is constructed to achieve multi-subject interest balance and mutual benefit and win-win results;

[0090] Solution module: used to solve the upper-level Stackelberg game problem based on KKT conditions and reduce the computational complexity of the model;

[0091] The upper-level game plan is transferred to the lower-level Nash game model. Based on the actual source-load uncertainty state, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smooth out the impact of uncertainty, and obtain the optimal transaction result for cooperation among UAs. By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved, ultimately achieving multi-layered and hierarchical collaborative scheduling of multiple loads.

[0092] Furthermore, in the framework construction module, the multi-load energy interaction framework including UA includes: the upper-level power grid and gas company are the sources of electricity and gas, the microgrid operator MO achieves electricity and heat coordination through combined heat and power (CHP) units and gas boilers (GB) to meet the electricity and heat needs of massive users; the microgrid operator manages electricity and heat storage, as well as power-to-gas (P2G) and carbon capture (CCS) equipment to reduce carbon emissions;

[0093] The controllable resources on the user side include household photovoltaics and residential electric loads and thermal loads that can participate in demand response. It is defined that there are I UAs in the multi-load microgrid, and the set is U={1,…,UA i ,…UA I}, each UA i Manage N in a specific area i users; the set of control time periods T = {1,…,t,…T}.

[0094] Furthermore, in the modeling module, the user load response comfort model taking into account multi-dimensional characteristics is expressed as:

[0095]

[0096] Where, are the comfort cost coefficients of electrical load and thermal load respectively; UA i The economic incentive level EIL, incentive timing characteristics ITC and user herd mentality UCP of electric load, and the same applies to thermal load; Can be expressed as:

[0097]

[0098] Where, The energy price of the current period is used to express it. MO is UA for period t i The electricity sales price set; Consider the current period's energy price relative to the overall period's price; Affected by the number of users participating in demand response in previous periods, UA for period t-1 i The number of users participating in electricity load demand response within the control area;

[0099] Similarly, Expressed as

[0100]

[0101] Where, is the selling price of MO during period t; UA for period t-1 i The number of users participating in heat load demand response within the control area;

[0102] The number of users participating in demand response can be expressed as:

[0103]

[0104] Where, Indicates UA i The load response variable of the kth user in the tth period is: indicates participation in demand response, whereas non-participation is indicated. p(·) represents probability, where p0 is the probability of not participating in demand response in the current period due to the influence of herd mentality given that the participant participated in demand response in the previous period. p1 is the probability of not participating in demand response in the previous period due to the influence of herd mentality given that the participant participated in demand response in the current period.

[0105] By means of the above solution, the present invention has at least the following advantages:

[0106] 1. This paper proposes a multi-load hierarchical scheduling framework that takes user response comfort into account. First, UAs are introduced to establish a basic framework for multi-load energy interaction. Second, a user response comfort model is constructed, taking into account economic incentives and user psychology, accurately reflecting the willingness of multi-load users to participate in demand response. Finally, decision-making models are constructed for microgrid operators and UAs, respectively. Based on game theory, the multi-load hierarchical coordinated scheduling problem is transformed into a game problem, achieving a balanced and mutually beneficial outcome for all parties.

[0107] 2. This paper proposes a multi-agent reinforcement learning-based multi-load hierarchical collaborative scheduling method. First, based on the Karush-Kuhn-Tucker (KKT) condition, the upper-level Stackelberg game problem is solved to reduce the model's computational complexity. Second, the upper-level game plan is transferred to the lower-level game model. Based on the actual source and load uncertainty, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smoothing out the effects of uncertainty and achieving the optimal transaction results for cooperation between user agents (UAs). By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved.

[0108] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate a certain embodiment of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0110] Figure 1 It is a flow chart of the multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning of the present invention. DETAILED DESCRIPTION

[0111] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0112] This invention proposes a multi-load hierarchical coordinated scheduling method based on multi-agent game learning, which includes two parts: a multi-load hierarchical scheduling framework that takes into account user response comfort and a multi-load hierarchical coordinated scheduling method based on multi-agent reinforcement learning. Figure 1 The specific technical solutions are as follows:

[0113] S1: This paper proposes a multi-load hierarchical scheduling framework that takes user response comfort into account. First, UA is introduced to establish a basic framework for multi-load energy interaction. Second, considering economic incentives and user psychology, a user response comfort model is constructed to accurately reflect the willingness of multi-load users to participate in demand response. Finally, decision-making models are constructed for microgrid operators and UAs respectively. Based on game theory, the multi-load hierarchical coordinated scheduling problem is transformed into a game problem, achieving a balanced and mutually beneficial outcome for multiple stakeholders.

[0114] In a further embodiment, step S1 comprises:

[0115] S1.1 Multi-load energy interaction framework

[0116] In a multi-load energy interaction framework involving UA, the upstream power grid and gas company are the primary sources of electricity and gas. Microgrid operators (MOs) use combined heat and power (CHP) units and gas boilers (GB) to achieve power and heat synergy, meeting the electricity and heat needs of a large number of users. Microgrid operators also manage electricity and heat storage, as well as power-to-gas (P2G) and carbon capture systems (CCS) to reduce carbon emissions.

[0117] The controllable resources on the user side mainly include household photovoltaics and residential electric loads and thermal loads that can participate in demand response. Considering the small scheduling capacity of individual users, the present invention introduces user aggregators to aggregate small and medium-sized user resources with similar geographical distribution and energy consumption characteristics, providing demand response services that are better than those of a single user, so as to fully tap the demand response capabilities of the user side and expand the user's profit space. It is defined that there are a total of I UAs in the multi-load microgrid, and the set is U = {1,...,UA i ,...UA I}, each UA i Manage N in a specific area i Users. The set of control periods T = {1,...,t,...T}. In this energy interaction framework, the MO is responsible for operating multiple loads and interacting with the upstream power grid and gas company. The MO formulates energy control plans, including differentiated energy pricing for each user-assigned user (UA), with the goal of minimizing the operating costs of these multiple loads. UAs form a cooperative alliance, adjust their energy consumption plans based on the energy prices issued by the MO and their current status, and interact with each other within the alliance to achieve optimal and reasonable profit distribution.

[0118] S1.2 User load response comfort model considering multi-dimensional characteristics

[0119] Users participate in demand response by adjusting their own energy consumption, and while gaining benefits, they also incur comfort costs due to changes in energy consumption habits. To this end, the present invention constructs a response comfort model for user electricity and heat loads, which is expressed as

[0120]

[0121] Where, are the comfort cost coefficients of electrical load and thermal load respectively. The lower the value, the smaller the comfort loss caused by load regulation and the stronger the willingness to participate in demand response. UA i The economic incentive level (EIL), incentive timing characteristics (ITC), and user conformity psychology (UCP) of electrical loads, and the same applies to thermal loads. Can be expressed as

[0122]

[0123] Where, The energy price of the current period is used to express it. MO is UA for period t i The electricity sales price is set. The main consideration is the positioning of energy prices in the current period in the overall period prices. Affected by the number of users participating in demand response in previous periods, UA for period t-1 i The number of users participating in the load demand response in the control area (considering the privacy of users, the decision-making situation of users in the current period will not be announced, so select characterization).

[0124] Similarly, Expressed as

[0125]

[0126] Where, is the selling price of MO during period t. UA for period t-1 i The number of users participating in heat load demand response within the control area.

[0127] Considering the influence of herd mentality, the number of users participating in demand response can be expressed as

[0128]

[0129] Where, Indicates UA i The load response variable of the kth user in the tth period is: = represents participation in demand response, whereas ≤ represents non-participation. p(·) represents the probability, where p0 is the probability of a participant participating in demand response in the previous period but being influenced by the herd mentality and not participating in demand response in the current period; and p1 is the probability of a participant not participating in demand response in the previous period but being influenced by the herd mentality and participating in demand response in the current period.

[0130] S1.3 Multi-Load Stackelberg-Nash Game Model

[0131] (1)MO decision model

[0132] First, the MO decision model is constructed. As the leader of the master-slave game, MO takes maximizing its own operating profit as the optimization goal in terms of economy, low carbon and comfort, which can be expressed as

[0133]

[0134] Where R ECO (t) is the economic benefit of MO, C LC (t) is the low carbon cost, C COM (t) is the comfort cost. R ECO (t) can be expressed as

[0135]

[0136] Where C b (t) is the time period from the upper power grid and UA to MO. i The cost of purchasing electricity, R s (t) is the MO's report to the upper grid and UA i Revenue from electricity sales, UA i The energy storage loss cost, new energy grid access expenditure and CHP operation cost are respectively expressed as follows

[0137]

[0138] C CHP (t) = λ CHP P CHP (t)+λ P2G P P2G (t)+λ CCS P CCS (t)(11)

[0139] Where Δt is the length of the control period. are the grid electricity price and on-grid electricity price in period t, MO to UA i The electricity price for users, Hot selling price. It is the amount of electricity purchased and sold between MO and the upper power grid. i b (t) is UA i Electricity purchase amount, UA i Purchase heat. are the charging and discharging power of electrical and thermal energy storage respectively. E ,λ H is the cost coefficient of electricity and heat storage. PV (t) is the household photovoltaic grid-connected electricity price paid by MO to users, using the full power generation subsidy model, P i PV (t) is the household photovoltaic output power. gas (t) is the gas purchase cost per unit of gas, V gas (t) = V CHP (t)+V GB (t)-V P2G (t) is the gas usage rate, where V CHP (t), V GB (t) is the gas consumption rate of CHP and GB, V P2G (t) is the rate at which the P2G device generates gas. CHP (t), P P2G (t), P CCS (t) is the discharge power of CHP, P2G and CCS, λ CHP ,λ P2G ,λ CCS is the operation and maintenance cost coefficient of CHP, P2G and CCS.

[0140] For low carbon cost C LC (t), since the current domestic carbon quota allocation method mainly adopts free allocation, UA i Carbon emission quota E * (t) is expressed as

[0141]

[0142] Where κ is the regional carbon emission per unit electricity, which is 590t / (kW·h) here; is the electricity conversion coefficient, H CHP (t) is the thermal power of the CHP unit.

[0143] Considering that actual carbon emissions mostly come from CHP and GB units, and the power purchased by the upper power grid mainly comes from coal-fired power units, the actual carbon emissions of the system during period t are

[0144]

[0145] Where, E(t) is the actual total carbon emissions of MO during period t, E b (t) is the carbon emissions generated by purchasing electricity from the upper power grid, E CHP (t), E GB (t), E CCS (t) is the actual carbon emissions of CHP, GB, and CCS units. η is the carbon emission coefficient of coal-fired power units, κ CHP , κ GB η is the carbon emission coefficient of CHP and gas turbine, which is 591t / (kW·h). CCS is the energy consumption of CCS device to process CO2, η CCS =0.269MW·h / t. Therefore, the low carbon cost can be expressed as

[0146] C LC (t) = λ LC (E(t)-E * (t)) (14)

[0147] Where λ LC is the carbon trading price coefficient.

[0148] C COM (t) is the comfort cost caused by the change of energy consumption habits when the user side load participates in demand response, which can be expressed as

[0149]

[0150] Where, P i CL (t), UA for period t i Response amount of electrical load and thermal load; is the controllable load response type weight, where For electrical and thermal loads, when UA i When the controllable load is in a reduction-type response, the weight is 1; otherwise, when the controllable load is in an absorption-type response, the weight is [0, 1]. is the set of response comfort cost coefficients of controllable loads. UA i Minimum incentive prices for electrical and thermal loads; is the response type indicator parameter. ⊙ is the Hadamard product, which means the multiplication of corresponding matrix elements.

[0151] The constraints of the MO optimization model include system constraints and equipment operation constraints. The system constraints are as follows:

[0152] 1) Electricity and heat sales price constraints

[0153] To prevent MO from deviating from market reality in order to maximize its own profits and to ensure fairness in transactions, the electricity and heat sales prices must meet the following constraints:

[0154]

[0155] Where x1={1,2} refers to energy price, 1 and 2 represent and are the upper and lower limits of the MO energy interaction price; is the average value requirement of MO energy interaction price.

[0156] 2) Constraints on electricity purchase and sales

[0157]

[0158] Where, is the indicator variable for MO power purchase and sale. If MO purchases power from the distribution network, Similarly, if MO sells electricity to the distribution network,

[0159] 3) Electricity conservation constraint

[0160] The charge conservation constraint can be expressed as

[0161]

[0162] Where, P i CL,o (t) is the UA in period t i Electrical load before regulation.

[0163] 4) Heat conservation constraint

[0164]

[0165] Where, is the charge and discharge amount of thermal energy storage. UA for the tth period i Heat load before regulation.

[0166] The equipment managed by MO mainly includes CHP units, GB units, electric and thermal energy storage, P2G and CCS equipment. Their operation constraints are as follows:

[0167] 1) CHP unit

[0168]

[0169] Where q NG is the calorific value of natural gas. is the heat-to-electricity ratio and power generation efficiency of the CHP unit. It is the upper limit of the power generation capacity of the CHP unit.

[0170] 2) GB unit

[0171]

[0172] Where η GB is the energy conversion efficiency of GB, The upper limit of GB output.

[0173] 3) Electric and thermal energy storage

[0174]

[0175] Where x3 = {1, 2} refers to the energy storage device of MEMG, and 1 and 2 represent electrical energy storage and thermal energy storage, respectively. is the charging and discharging power of the energy storage during period t, It is the upper limit of energy storage charging and discharging power. is the energy storage charging and discharging indicator variable. If the energy storage device is charged, on the contrary

[0176] 4) P2G and CCS equipment

[0177]

[0178] Where η P2G is the natural gas conversion efficiency of the P2G device. CCS is the power conversion coefficient between CCS and P2G.

[0179] (2) UA Decision Model

[0180] UA takes minimizing the cost of the jurisdiction as the optimization goal, which includes the cost of energy purchase, comfort cost, photovoltaic grid access income and inter-UA transaction income. The optimization goal can be expressed as

[0181]

[0182] Where, The benefits of energy interaction between users and aggregators, UA i The energy purchase cost is expressed as

[0183]

[0184] Where λ i→j (t) is UA i With UA j The energy transaction price, P i i→j(t) is UA i To UA j Output power.

[0185] The constraints for UA decision are:

[0186] 1) User-side controllable load constraints

[0187]

[0188] Where, P i CL,min (t), P i CL,max (t) is the UA of period t i Upper and lower limits of electric load response, P i CL When (t) takes a positive value, it is a reduction response, otherwise it is an absorption response; The upper and lower limits of the heat load response.

[0189] 2) Power interaction constraints between UAs

[0190] Considering that the technology of user-side thermal energy interaction is not yet mature, this invention mainly considers the power interaction between UAs, with the following constraints:

[0191]

[0192] P i→j (t)=-P j→i (t) (30)

[0193] Where, It is the upper limit of power exchange between UAs.

[0194] 3) Electricity conservation constraint

[0195]

[0196] 4) Heat conservation constraint

[0197]

[0198] (3) Stackelberg-Nash dual game model

[0199] During the multi-load energy interaction process, the MO serves as a leader at the top level, possessing bargaining and decision-making power. It can formulate interaction strategies with the upper-level power grid and gas company, as well as energy management strategies for conventional units and individual UAs within the microgrid. UAs must respond to the MO's energy pricing and serve as lower-level followers. At the same time, each UA is on an equal footing and must optimize its own profits through cooperation. To this end, the present invention constructs a game model between the MO and UAs. The first stage of the game is the Stackelbelg game between the MO and each UA. The MO sets electricity and heat prices based on the differentiated characteristics of each UA to maximize MO profits. The second stage is the Nash game within the UA alliance. Based on the energy prices issued by the upper-level MO and the current source-load status, each UA minimizes alliance costs through inter-UA cooperation. The Stackelbelg-Nash game enables fair and reasonable distribution of benefits among various entities at multiple levels.

[0200] S2: This paper proposes a multi-layered and hierarchical collaborative scheduling method for multiple loads based on multi-agent reinforcement learning. First, based on the Karush-Kuhn-Tucker (KKT) condition, the upper-level Stackelberg game problem is solved to reduce the model's computational complexity. Second, the upper-level game plan is transferred to the lower-level game model. Based on the actual source and load uncertainty state, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smoothing out the effects of uncertainty and obtaining the optimal transaction results for cooperation between user agents (UAs). By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved.

[0201] In a further embodiment, step S2 comprises:

[0202] S2.1 Solution of Stackelberg Game Based on KKT Conditions

[0203] In the Stackelberg game, the strategies of the upper-level MO and the lower-level UA are coupled, requiring repeated iterations, making the upper-level model difficult to solve directly. Since the decision variables (and the electricity and heat sales prices) of the upper-level model can be regarded as constants in the lower-level model, the lower-level model satisfies the Slater condition. Therefore, the present invention constructs the Lagrangian function of the lower-level model. For the presence of bilinear terms and complementary slack conditions, the strong duality theory is adopted. Based on the Big M method and McCormick envelope process, the original problem is transformed into a single-level mixed integer linear programming problem that is easy to solve. In this way, the lower-level problem is transformed into the optimality condition of the upper problem.

[0204] S2.2 Nash Game Solution Based on Nash Game-MADDPG

[0205] This paper proposes a Nash game-MADDPG algorithm, the detailed steps are as follows

[0206] (3) Construction of Markov decision process model

[0207] The Markov decision process model of the multi-load Nash game mainly includes the state space, action space and real-time reward function considering the collaboration cost.

[0208] 1) State space: Define S(t) = {O1(t), ..., O i (t),...,O I (t)}, where O i (t) is UA i The state space of

[0209]

[0210] Where, P i PV,act (t) is the UA of period t i The actual output of household photovoltaics.

[0211] 2) Action Space

[0212] To improve computational efficiency and satisfy the constraints of equations (29)-(30), P i→j (t) is converted to

[0213]

[0214] Where, X i→j (t) is UA i The power interaction between UAs generated by the intelligent agent. Assume that UA i The action space is Expressed as

[0215]

[0216] 3) Real-time reward function

[0217] In order to improve the training efficiency of local agents, this paper introduces the agent collaboration cost. The real-time reward function is expressed as

[0218]

[0219] Where, UA i The collaboration cost is related to the strategies of other agents. is the collaboration cost coefficient between UAs. The collaboration cost allows the agent to have a global awareness and consider the benefits of other agents while optimizing its own benefits.

[0220] (2) Nash Game-MADDPG Algorithm Solution Process

[0221] The Nash game-MADDPG algorithm adopts the actor-critic network structure. The MO side sets up a centralized main Q network and target Q network. Based on the global state information S(t), it guides UA to learn a strategy that approaches the global Nash equilibrium. Each UA is configured with two policy networks, which rely on the local observation state O i (t) Generate its own strategy and execute it in a distributed manner. Each UA does not exchange its own observation status. In addition, the present invention sets up an experience replay pool that aggregates global information to assist in the training of intelligent agents, and proposes an experience replay mechanism based on enhanced memory. MO has the authority to call all UA experience data, while each UA can only extract its own experience data to effectively protect the privacy and security of users in different regions. The network settings of the NASH game-MADDPG algorithm are as follows: UA i The main policy network is represented as The centralized main Q network on the MO side is expressed as The corresponding target networks are φ i ′, Q i ′, the network parameters are The process of solving Nash equilibrium based on NASH game-MADDPG is as follows

[0222] 1) UA i The agent generates a strategy based on the behaviorpolicyβ strategy The same applies to other intelligent agents.

[0223] 2) Each UA calculates its own real-time reward function r i (t), the state is transferred to O i (t+1), the experience chain Upload to MO and store in the experience replay pool Φ.

[0224] 3) MO calculates the collaboration cost of each UA and asynchronously updates the Q network corresponding to each UA in the centralized main Q network based on the global environment state. The coefficient update formula is:

[0225]

[0226] Where Q i ′(S(t+1),π(t+1)) is the corresponding UA generated by the target Q network in state S(t+1) i The Q value is the cumulative return of all UAs adopting the Nash equilibrium strategy starting from S(t+1). This paper adopts the Pareto optimal form.

[0227] Afterwards, the present invention extracts samples ||Φ(t)|| from Φ and updates the main Q network parameters. Since the MADDPG algorithm involves independent training of multiple agents and global guidance and judgment in solving the Nash equilibrium, convergence is difficult. The present invention proposes an experience replay mechanism based on enhanced memory, which is measured by TD-error and reasonably increases the replay probability of action chains that differ greatly from the predicted value to improve training efficiency. TD-error can be expressed as

[0228]

[0229] The selection probability of different experience chains is expressed as

[0230]

[0231] Where ε is a small random number to avoid the situation where the TD-error is close to 0 and cannot be selected. The update formula of the centralized main Q network is

[0232]

[0233] 4) Each UA extracts ||Φ from the MO's experience replay pool i (t)|| experience chains are trained, and the parameter update process of the main strategy network is expressed as

[0234]

[0235] 5) Update the target Q network and target policy network, the formula is

[0236]

[0237] Where, is the weight of the soft update parameter.

[0238] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning, characterized by The specific scheduling steps are: First, a multi-load energy interaction framework including UA is constructed; Based on the above multi-load energy interaction framework, and taking into account the economic incentive level and user psychology, a user load response comfort model with multi-dimensional characteristics is constructed to accurately reflect the willingness of multiple loads on the user side to participate in demand response. At the same time, the MO decision model and UA decision model are constructed respectively, and based on game theory, a multi-load Stackelberg-Nash game model is constructed to achieve multi-subject interest balance and mutual benefit and win-win results; Based on the KKT condition, the upper-level Stackelberg game problem is solved to reduce the computational complexity of the model; The upper-level game plan is transferred to the lower-level Nash game model. Based on the actual source-load uncertainty state, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smooth out the impact of uncertainty, and obtain the optimal transaction result for cooperation among UAs. By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved, ultimately achieving multi-layered and hierarchical collaborative scheduling of multiple loads.

2. The multi-layered and hierarchical coordinated scheduling method for multiple loads based on multi-agent game learning according to claim 1 is characterized by: The multi-load energy interaction framework with UA includes: the upper-level power grid and gas company are the sources of electricity and gas, the microgrid operator (MO) uses combined heat and power (CHP) units and gas boilers (GB) to achieve electricity and heat coordination to meet the electricity and heat needs of a large number of users; the microgrid operator manages electricity and heat storage, as well as power-to-gas (P2G) and carbon capture (CCS) equipment to reduce carbon emissions; The controllable resources on the user side include household photovoltaics and residential electric loads and thermal loads that can participate in demand response. It is defined that there are I UAs in the multi-load microgrid, and the set is U={1,…,UA i ,…UA I }, each UA i Manage N in a specific area i users; the set of control time periods T = {1,...,t,...T}.

3. The multi-layered and hierarchical coordinated scheduling method for multiple loads based on multi-agent game learning according to claim 1 is characterized by: The user load response comfort model considering multi-dimensional characteristics is expressed as: Where, are the comfort cost coefficients of electrical load and thermal load respectively; UCP i e (t) are UA i The economic incentive level EIL, incentive timing characteristics ITC and user herd mentality UCP of electric load, and the same applies to thermal load; Can be expressed as: Where, The energy price of the current period is used to express it. MO is UA for period t i The electricity sales price set; Consider the current period's energy price relative to the overall period's price; Affected by the number of users participating in demand response in previous periods, UA for period t-1 i The number of users participating in electricity load demand response within the control area; Similarly, Expressed as Where λ i h,s (t) is the selling price of MO during period t; UA for period t-1 i The number of users participating in heat load demand response within the control area; The number of users participating in demand response can be expressed as: Where, Indicates UA i The load response variable of the kth user in the tth period is: indicates participation in demand response, whereas non-participation is indicated. p(·) represents probability, where p0 is the probability of not participating in demand response in the current period due to the influence of herd mentality given that the participant participated in demand response in the previous period. p1 is the probability of not participating in demand response in the previous period due to the influence of herd mentality given that the participant participated in demand response in the current period.

4. The multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning according to claim 1 is characterized by: The MO decision model is expressed as: Where R ECO (t) is the economic benefit of MO, C LC (t) is the low carbon cost, C COM (t) is the comfort cost; R ECO (t) can be expressed as Where C b (t) is the time period from the upper power grid and UA to MO. i The cost of purchasing electricity, R s (t) is the MO's report to the upper grid and UA i Revenue from electricity sales, UA i Energy storage loss costs, new energy grid access expenses and CHP operating costs are expressed as follows: C CHP (t)=λ CHP P CHP (t)+λ P2G P P2G (t)+λ CCS P CCS (t) (11) Where Δt is the length of the control period; is the grid electricity price and on-grid electricity price in period t, MO to UA i The electricity price for users, Hot selling price; P is the amount of electricity purchased and sold between MO and the upper power grid; i b (t) is UA i Electricity purchase amount, UA i Purchased heat; are the charging and discharging power of electrical and thermal energy storage respectively; E ,λ H is the cost coefficient of electricity and heat storage; PV (t) is the household photovoltaic grid-connected electricity price paid by MO to users, using the full power generation subsidy model, P i PV (t) is the household photovoltaic output power; gas (t) is the gas purchase cost per unit of gas, V gas (t) = V CHP (t)+V GB (t)-V P2G (t) is the gas usage rate, where V CHP (t), V GB (t) is the gas consumption rate of CHP and GB, V P2G (t) is the rate at which the P2G device generates gas; P CHP (t), P P2G (t), P CCS (t) is the discharge power of CHP, P2G and CCS, λ CHP ,λ P2G ,λ CCS is the operation and maintenance cost coefficient of CHP, P2G and CCS; For low carbon cost C LC (t), since the current domestic carbon quota allocation method mainly adopts free allocation, UA i Carbon emission quota E * (t) is expressed as Where κ is the regional carbon emission per unit electricity, which is 590t / (kW·h) here; is the electricity conversion coefficient, H CHP (t) is the thermal power of the CHP unit; Considering that actual carbon emissions mostly come from CHP and GB units, and the power purchased by the upper power grid mainly comes from coal-fired power units, the actual carbon emissions of the system during period t are Where, E(t) is the actual total carbon emissions of MO during period t, E b (t) is the carbon emissions generated by purchasing electricity from the upper power grid, E CHP (t), E GB (t), E CCS (t) is the actual carbon emissions of CHP, GB, and CCS units; η is the carbon emission coefficient of coal-fired power units, κ CHP , κ GB is the carbon emission coefficient of CHP and gas turbine, which is 591t / (kW·h); η CCS is the energy consumption of CCS device to process CO2, η CCS =0.269MW·h / t; thus, the low-carbon cost can be expressed as C LC (t)=λ LC (E(t)-E * (t))(14) Where λ LC is the carbon trading price coefficient; C COM (t) is the comfort cost caused by the change of energy consumption habits when the user side load participates in demand response, which can be expressed as Where, P i CL (t), UA for period t i Response amount of electrical load and thermal load; is the controllable load response type weight, where For electrical and thermal loads, when UA i When the controllable load is in a reduction response, the weight is 1; otherwise, when the controllable load is in an absorption response, the weight is [0, 1]. is the set of response comfort cost coefficients of controllable loads; UA i Minimum incentive prices for electrical and thermal loads; is the response type indicator parameter; ⊙ is the Hadamard product, which means the multiplication of corresponding matrix elements.

5. The multi-layered and hierarchical coordinated scheduling method for multiple loads based on multi-agent game learning according to claim 1 is characterized by: The UA decision model aims to minimize the cost of the jurisdiction, which includes the cost of energy purchase, comfort cost, photovoltaic grid connection income, and inter-UA transaction income. The optimization objective is expressed as: Where R i U2U (t) is the revenue from energy interaction between user aggregators, UA i The energy purchase cost is expressed as: Where λ i→j (t) is UA i With UA j The energy transaction price, P i i→j (t) is UA i To UA j Output power.

6. The multi-layered and hierarchical coordinated scheduling method for multiple loads based on multi-agent game learning according to claim 1 is characterized by: The first stage of the multi-load Stackelberg-Nash game model is the Stackelberg game between the MO and each UA. The MO sets electricity and heat prices based on the differentiated characteristics of each UA to maximize MO profits. The second stage is the Nash game of the UA alliance. Based on the energy price issued by the upper-level MO and the current source and load status, each UA minimizes the alliance cost through inter-UA cooperation.

7. The multi-layered and hierarchical coordinated scheduling method for multiple loads based on multi-agent game learning according to claim 1 is characterized by: The method for solving the upper-level Stackelberg game problem based on the KKT condition is: Since the decision variables of the upper-level model can be regarded as constants in the lower-level model, and the lower-level model satisfies the Slater condition, the Lagrangian function of the lower-level model is constructed. For the existence of bilinear terms and complementary slack conditions, the strong duality theory is adopted. Based on the Big M method and McCormick envelope process, the original problem is transformed into a single-level mixed integer linear programming problem that is easy to solve, thereby converting the lower-level problem into the optimality condition of the upper-level problem.

8. The multi-load hierarchical and graded collaborative scheduling method based on multi-agent game learning according to claim 1 is characterized by: The steps of proposing the Nash game-MADDPG algorithm to solve the Nash game problem are: (1) Construction of Markov decision process model The Markov decision process model of the multi-load Nash game mainly includes the state space, action space and real-time reward function considering the coordination cost; 1) State space: Define S(t) = {O1(t), ..., O i (t),...,O I (t)}, where O i (t) is UA i The state space of is expressed as: Where, P i PV,act (t) is the UA of period t i The actual output of household photovoltaics; 2) Action Space To improve computational efficiency, P i→j (t) is converted to Where, X i→j (t) is UA i Power interaction between UAs generated by intelligent agents; The upper limit of power interaction between UAs; Assume that UA i The action space is Expressed as 3) Real-time reward function In order to improve the training efficiency of local agents, this paper introduces the agent collaboration cost. The real-time reward function is expressed as Where, UA i The collaboration cost is related to the strategies of other agents. is the collaboration cost coefficient between UAs. Collaboration cost allows agents to have a global awareness and consider the benefits of other agents while optimizing their own benefits; (2) Nash Game-MADDPG Algorithm Solution Process The Nash game-MADDPG algorithm adopts the actor-critic network structure. The MO side sets up a centralized main Q network and target Q network. Based on the global state information S(t), it guides the UA to learn a strategy that approaches the global Nash equilibrium. Each UA is configured with two policy networks, which rely on the local observation state O i (t) Generate its own strategy and execute it in a distributed manner. Each UA does not exchange its own observation state. At the same time, an experience replay pool that aggregates global information is set up to assist in the training of the agent. An experience replay mechanism based on enhanced memory is proposed. The network settings of the NASH game-MADDPG algorithm are as follows: UA i The main policy network is represented as The centralized main Q network on the MO side is expressed as The corresponding target networks are φ′ i , Q′ i , the network parameters are The process of solving Nash equilibrium based on NASH game-MADDPG is as follows 1) UA i The agent generates a strategy based on the behaviorpolicyβ strategy 2) Each UA calculates its own real-time reward function r i (t), the state is transferred to O i (t+1), the experience chain Upload to MO and store in the experience replay pool Φ; 3) MO calculates the collaboration cost of each UA and asynchronously updates the Q network corresponding to each UA in the centralized main Q network based on the global environment state. The coefficient update formula is: Where Q i ′(S(t+1),π(t+1)) is the corresponding UA generated by the target Q network in state S(t+1) i The Q value is the cumulative return of all UAs adopting the Nash equilibrium strategy starting from S(t+1), in the form of Pareto optimality; Then, samples ||Φ(t)|| are extracted from Φ to update the main Q network parameters. An experience replay mechanism based on enhanced memory is proposed and measured by TD-error, which can be expressed as: The selection probability of different experience chains is expressed as Where ε is a small random number to avoid the situation where the TD-error is close to 0 and cannot be selected; The update formula of the centralized main Q network is: 4) Each UA extracts ||Φ from the MO's experience replay pool i (t)|| experience chains are used for training, and the parameter update process of the main policy network is expressed as: 5) Update the target Q network and target policy network, the formula is Where, is the weight of the soft update parameter.

9. A multi-load hierarchical and graded collaborative scheduling device based on multi-agent game learning, characterized by include: Framework construction module: used to build a multi-load energy interaction framework including UA; Modeling module: Based on the above multi-load energy interaction framework, the module considers economic incentives and user psychology, builds a user load response comfort model that takes into account multi-dimensional characteristics, and accurately reflects the willingness of multiple loads on the user side to participate in demand response. At the same time, the MO decision model and UA decision model are constructed respectively, and based on game theory, a multi-load Stackelberg-Nash game model is constructed to achieve multi-subject interest balance and mutual benefit and win-win results; Solution module: used to solve the upper-level Stackelberg game problem based on KKT conditions and reduce the computational complexity of the model; The upper-level game plan is transferred to the lower-level Nash game model. Based on the actual source-load uncertainty state, the Nash game-MADDPG algorithm is proposed to solve the Nash game problem, smooth out the impact of uncertainty, and obtain the optimal transaction result for cooperation among UAs. By introducing collaboration costs and an experience replay mechanism based on enhanced memory, the Nash equilibrium of the UA alliance is efficiently solved, ultimately achieving multi-layered and hierarchical collaborative scheduling of multiple loads.

10. The multi-load hierarchical and graded collaborative scheduling device based on multi-agent game learning according to claim 9, characterized in that: The framework building module includes a multi-load energy interaction framework with UA, which includes: the upper-level power grid and gas company are the sources of electricity and gas, and the microgrid operator (MO) uses combined heat and power (CHP) units and gas boilers (GB) to achieve electricity and heat coordination to meet the electricity and heat needs of a large number of users; the microgrid operator manages electricity and heat storage, as well as power-to-gas (P2G) and carbon capture (CCS) equipment to reduce carbon emissions; The controllable resources on the user side include household photovoltaics and residential electric loads and thermal loads that can participate in demand response. It is defined that there are I UAs in the multi-load microgrid, and the set is U={1,…,UA i ,...UA I }, each UA i Manage N in a specific area i users; the set of control time periods T = {1,...,t,…T}.

11. The multi-load hierarchical and graded collaborative scheduling device based on multi-agent game learning according to claim 9, characterized in that: In the modeling module, the user load response comfort model taking into account multi-dimensional characteristics is expressed as: Where, are the comfort cost coefficients of electrical load and thermal load respectively; UA i The economic incentive level EIL, incentive timing characteristics ITC and user herd mentality UCP of electric load, and the same applies to thermal load; Can be expressed as: Where, The energy price of the current period is used to express it. MO is UA for period t i The electricity sales price set; Consider the current period's energy price relative to the overall period's price; Affected by the number of users participating in demand response in previous periods, UA for period t-1 i The number of users participating in electricity load demand response within the control area; Similarly, Expressed as Where, is the selling price of MO during period t; UA for period t-1 i The number of users participating in heat load demand response within the control area; The number of users participating in demand response can be expressed as: Where, Indicates UA i The load response variable of the kth user in the tth period is: indicates participation in demand response, whereas non-participation is indicated. p(·) represents probability, where p0 is the probability of not participating in demand response in the current period due to the influence of herd mentality given that the participant participated in demand response in the previous period. p1 is the probability of not participating in demand response in the previous period due to the influence of herd mentality given that the participant participated in demand response in the current period.