Multi-agent bidding behavior analysis method and system based on enemy friend game framework
By combining the friend-enemy game framework with reinforcement learning, the problems of carbon emission impact and imperfect competition in the electricity market are solved, accurate analysis and strategy optimization of multi-agent bidding behavior are achieved, and more flexible bidding behavior auxiliary decision-making is provided.
Patent Information
- Application Number
- CN202510929213.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies are unable to effectively consider the impact of carbon emissions in the electricity market, resulting in inaccurate data analysis results and an inability to adapt to the complexity of imperfectly competitive markets, resulting in inflexible and inaccurate bidding behavior strategies.
A multi-agent bidding behavior analysis method based on the friend-enemy game framework is adopted. By obtaining the power generation and carbon emission data of each target entity, the marginal power generation cost is calculated, and the bidding strategy is optimized using reinforcement learning and the friend-enemy game framework. The reward mechanism of entities inside and outside the group is considered to optimize the bidding behavior of enterprises.
It achieves the accuracy and flexibility of data analysis results for multiple entities in an imperfectly competitive market, provides more accurate bidding behavior strategy assistance, and adapts to the complex power market environment.
Smart Images

Figure CN120807068A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of behavior information analysis, and in particular to a multi-agent bidding behavior analysis method and system based on an enemy-friend game framework. BACKGROUND
[0002] The current power industry has gradually changed to liberalization and marketization to encourage competition and improve efficiency. The transformation of the power market brings both opportunities for power generation companies to obtain more profits and various risks, including price fluctuations, market competition, etc. Therefore, it is necessary to analyze the corresponding data through data analysis methods so as to assist enterprises in making correct decisions based on the analysis results.
[0003] Currently, in order to effectively analyze data to provide corresponding auxiliary results, the current main reference is other meanings of a perfectly competitive market, research based on optimization models, game theory or agent-based modeling, and then analyze the power generation data and the data of the power market through the model, so as to determine the corresponding bidding behavior strategy according to the analysis results to provide a reference for enterprises.
[0004] However, these methods have the problems of only focusing on specific participants and being inflexible, so they cannot effectively guarantee the accuracy of the data analysis results. Moreover, these methods are based on the perspective of a perfectly competitive market, while the current power market is still in an imperfectly competitive market, and these methods do not take into account the impact of carbon emissions, so the accuracy of data analysis is low. Therefore, there is currently a need for a method that can accurately analyze data to provide accurate auxiliary information. SUMMARY
[0005] Based on the above-mentioned deficiencies of the prior art, the present application provides a multi-agent bidding behavior analysis method and system based on an enemy-friend game framework to solve the problem of inaccurate analysis results of the prior art.
[0006] In order to achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0007] The first aspect of the present application provides a multi-agent bidding behavior analysis method based on an enemy-friend game framework, comprising:
[0008] obtaining power generation data and carbon emission data of each target agent;
[0009] calculating the marginal power generation cost of each target agent based on the power generation data and the carbon emission data of each target agent, respectively;
[0010] respectively, based on the current updated state-action pair and the marginal generation cost of the target subject, a current reward of the state-action pair is calculated, and the expected cumulative reward value of the state-action pair is updated using the current reward of the state-action pair, until the expected cumulative reward values of all the state-action pairs converge;
[0011] wherein one of the state-action pairs comprises a state and an action selected in the state; the state comprises power demand information, clearing price and generation capacity; the action comprises a bidding strategy coefficient and reported generation capacity; if the target subject belongs to an object in a group, the current reward of the state-action pair is the total reward of each target subject in the current group in the state-action pair; if the target subject does not belong to an object in a group, the current reward of the state-action pair is the reward of the target subject in the state-action pair;
[0012] The action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state is determined as the optimal behavior strategy of the target subject in each state.
[0013] Optionally, in the multi-agent bidding behavior analysis method based on the enemy-friend game framework, after determining the action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state as the optimal behavior strategy of the target subject in each state, the method comprises:
[0014] Obtaining information of the current state of any one of the target subjects;
[0015] Based on the information of the current state of the target subject and the optimal behavior strategy of the target subject in each state, the optimal behavior strategy of the current state of the target subject is determined;
[0016] The bidding strategy coefficient in the optimal behavior strategy is used to calculate the current bid;
[0017] The optimal behavior strategy and the current bid are fed back.
[0018] Optionally, in the multi-agent bidding behavior analysis method based on the enemy-friend game framework, the marginal generation cost of each target subject is calculated based on the generation data and carbon emission data of each target subject, comprising:
[0019] For each target subject, the pure generation cost of the target subject is calculated using the generation data of the target subject;
[0020] if the carbon emission of the target subject is greater than the carbon emission benchmark, calculating a product of the carbon price and a difference between the carbon emission of the target subject and the carbon emission benchmark to obtain a carbon emission cost of the target subject;
[0021] if the carbon emission of the target subject is not greater than the carbon emission benchmark, determining that the carbon emission cost of the target subject is zero;
[0022] adding the power generation cost of the target subject and the carbon emission cost of the target subject to obtain a total power generation cost of the target subject;
[0023] calculating a marginal power generation cost of the target subject by using the total power generation cost of the target subject.
[0024] Optionally, in the multi-subject bidding behavior analysis method based on the enemy-friend game framework, the calculating the current reward of the state-action pair based on the current updated state-action pair and the marginal power generation cost of the target subject, and the updating the expected cumulative reward value of the state-action pair by using the current reward of the state-action pair are performed until the expected cumulative reward values of all the state-action pairs converge, and the method comprises the following steps.
[0025] initializing a current state and expected cumulative reward values of all state-action pairs;
[0026] selecting a current action in the current state to determine a current state-action pair; wherein the current state-action pair is a state-action pair comprising the current state and the current action;
[0027] calculating a current reward of the current state-action pair based on the current state-action pair and the marginal power generation cost of the target subject;
[0028] updating the expected cumulative reward value of the current state-action pair by using the reward of the current state-action pair;
[0029] updating the current state and returning to perform the selecting the current action in the current state to determine the current state-action pair until the expected cumulative reward values of all the state-action pairs converge.
[0030] Optionally, in the multi-subject bidding behavior analysis method based on the enemy-friend game framework, the calculating the current reward of the current state-action pair based on the current state-action pair and the marginal power generation cost of the target subject comprises the following steps.
[0031] if the target subject does not belong to the group, multiplying the marginal power generation cost of the target subject by the bidding strategy coefficient in the current action to obtain a current bid of the target subject;
[0032] multiplying the difference between the current offer of the target subject and the marginal generation cost of the target subject by the reported generation capacity in the current action to obtain a current initial reward of the target subject;
[0033] multiplying the current initial reward of the target subject by a first weight to obtain a current reward of the current state-action pair; wherein the first weight is greater than 0 and less than 1;
[0034] if the target subject belongs to a group, calculating a current initial reward of each target subject based on the current state-action pair and the marginal generation cost of each target subject in the group;
[0035] summing the results of multiplying the current initial reward of each target subject by a second weight to obtain a current reward of the current state-action pair; wherein the second weight is greater than 1.
[0036] The second aspect of the present application provides a multi-agent bidding behavior analysis system based on an enemy-friend game framework, comprising:
[0037] a first acquisition unit configured to acquire generation data and carbon emission data of each target subject;
[0038] a cost calculation unit configured to calculate the marginal generation cost of each target subject based on the generation data and carbon emission data of each target subject, respectively;
[0039] a learning unit configured to, for each target subject, repeatedly calculate a current reward of a state-action pair based on the current updated state-action pair and the marginal generation cost of the target subject, and update an expected cumulative reward value of the state-action pair using the current reward of the state-action pair, until the expected cumulative reward values of all state-action pairs converge;
[0040] wherein one state-action pair comprises a state and an action selected in the state; the state comprises power demand information, clearing price and generation capacity; the action comprises a bidding strategy coefficient and reported generation capacity; if the target subject belongs to an object in a group, the current reward of the state-action pair is the total reward of each target subject in the current group in the state-action pair; if the target subject does not belong to an object in a group, the current reward of the state-action pair is the reward of the target subject in the state-action pair;
[0041] a behavior determination unit configured to determine the action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state as the optimal behavior strategy of the target subject in each state.
[0042] Optionally, in the multi-agent bidding behavior analysis system based on the enemy-friend game framework, the system comprises:
[0043] a second acquisition unit configured to acquire information of a current state of any one of the target agents;
[0044] a strategy determination unit configured to determine a best behavior strategy of the current state of the target agent based on the information of the current state of the target agent and the best behavior strategy of the target agent in each of the states;
[0045] a first bid calculation unit configured to calculate a current bid by using a bid strategy coefficient in the best behavior strategy;
[0046] an information feedback unit configured to feed back the best behavior strategy and the current bid.
[0047] Optionally, in the multi-agent bidding behavior analysis system based on the enemy-friend game framework, the cost calculation unit comprises:
[0048] a first calculation unit configured to calculate, for each of the target agents, a pure power generation cost of the target agent by using power generation data of the target agent;
[0049] a second calculation unit configured to calculate, when the carbon emission of the target agent is greater than a carbon emission benchmark, a carbon emission cost of the target agent by multiplying a carbon price by a difference between the carbon emission of the target agent and the carbon emission benchmark;
[0050] a third calculation unit configured to determine that the carbon emission cost of the target agent is zero when the carbon emission of the target agent is not greater than the carbon emission benchmark;
[0051] a fourth calculation unit configured to calculate a total power generation cost of the target agent by adding the pure power generation cost of the target agent and the carbon emission cost of the target agent;
[0052] a cost determination unit configured to calculate a marginal power generation cost of the target agent by using the total power generation cost of the target agent.
[0053] Optionally, in the multi-agent bidding behavior analysis system based on the enemy-friend game framework, the learning unit comprises:
[0054] an initialization unit configured to initialize a current state and an expected cumulative reward value of an action pair in each state;
[0055] a selection unit configured to select a current action in the current state to determine a current state-action pair, wherein the current state-action pair is a state-action pair comprising the current state and the current action;
[0056] a reward calculation unit, configured to calculate a current reward for the current state-action pair based on the current state-action pair and the marginal power generation cost of the target entity;
[0057] a reward value updating unit, configured to update the expected cumulative reward value of the current state-action pair using the reward of the current state-action pair;
[0058] The state updating unit is used to update the current state and return to execute the selection unit until the expected cumulative reward values of each state-action pair converge.
[0059] Optionally, in the above-mentioned multi-agent bidding behavior analysis system based on the friend-enemy game framework, the reward calculation unit includes:
[0060] a second bid calculation unit, configured to, when the target entity does not belong to a group, multiply the bid strategy coefficient in the current action by the marginal power generation cost of the target entity to obtain a current bid of the target entity;
[0061] a first reward calculation unit, configured to multiply the difference between the current bid of the target entity and the marginal power generation cost of the target entity by the power generation reported in the current action to obtain the current initial reward of the target entity;
[0062] A first weighting unit, configured to multiply the current initial reward of the target subject by a first weight to obtain a current reward of the current state-action pair; wherein the first weight is greater than 0 and less than 1;
[0063] a second reward calculation unit, configured to calculate, when the target entity belongs to a group, a current initial reward of each target entity based on the current state-action pair and the marginal power generation cost of each target entity in the group to which it belongs;
[0064] A second weighting unit is used to sum the results of multiplying the current initial reward of each target subject by a second weight to obtain the current reward of the current state-action pair; wherein the second weight is greater than 1.
[0065] The application provides a multi-agent bidding behavior analysis method based on a friend-or-foe game framework, a power market clearing model based on agent modeling, so as to more flexibly analyze each agent. And the influence mechanism of carbon cost on enterprise bidding is introduced into the model. Therefore, when analysis is needed, the power generation data and carbon emission data of each target agent are obtained respectively. Based on the power generation data and carbon emission data of each target agent, the marginal power generation cost of the target agent is calculated. Then, for each target agent, the current reward of the state-action pair is calculated based on the current updated state-action pair and the marginal power generation cost of the target agent, and the expected cumulative reward value of the state-action pair is updated using the current reward of the state-action pair, until the expected cumulative reward values of all state-action pairs converge. A state-action pair includes a state and an action selected in the state. The state includes power demand information, clearing price and power generation. The action includes a bidding strategy coefficient and reported power generation. Therefore, the strategy selection of the enterprise is optimized through reinforcement learning. Moreover, if the target agent belongs to the object in the group, the current reward of the state-action pair is the total reward of each target agent in the current group under the state action. If the target agent does not belong to the object in the group, the current reward of the state-action pair is the reward of the current target agent under the state-action pair, so that the friend-or-foe game framework is introduced into the ABM-based power market bidding behavior simulation algorithm, so as to better adapt to the imperfect competition market. Finally, the action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state is determined as the best behavior strategy of the target agent in each state, so that the accurate analysis of the data of the target agent is realized, and accurate analysis results are obtained, so as to better assist the enterprise in action. BRIEF DESCRIPTION OF DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.
[0067] Figure 1 A flowchart of a multi-agent bidding behavior analysis method based on a friend-or-foe game framework is provided for the embodiments of the present application.
[0068] Figure 2 A flowchart of a method for calculating marginal power generation cost is provided for the embodiments of the present application.
[0069] Figure 3 A flowchart of a reinforcement learning method is provided for the embodiments of the present application.
[0070] Figure 4A flowchart of an information simulation auxiliary method provided for an embodiment of the present application;
[0071] Figure 5 A curve graph of an example of carbon price change trend provided for an embodiment of the present application;
[0072] Figure 6 A curve graph of an example of coal carbon price and natural gas price change trend provided for an embodiment of the present application;
[0073] Figure 7 A schematic diagram of an example of average load and renewable energy output of three places provided for an embodiment of the present application;
[0074] Figure 8 A schematic diagram of an example of day-ahead electricity price and real-time electricity price results of three places provided for an embodiment of the present application;
[0075] Figure 9 A curve graph of an example of average bidding of five large power generation groups and non-five large groups provided for an embodiment of the present application;
[0076] Figure 10 A histogram of an example of bidding of five large power generation groups and non-five large groups provided for an embodiment of the present application;
[0077] Figure 11 An architecture schematic diagram of a multi-agent bidding behavior analysis system based on an enemy and friend game framework provided for an embodiment of the present application. DETAILED DESCRIPTION
[0078] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0079] In this application, the relational terms such as first and second and the like are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0080] The embodiment of the present application provides a multi-agent bidding behavior analysis method based on an enemy and friend game framework, as shown in the figure, comprising the following steps. Figure 1
[0081] S101, acquiring power generation data and carbon emission data of each target agent.
[0082] The target agent can be each agent in the power market, i.e., each power generation enterprise. The power generation data is data related to the cost consumed in the power generation process of the target agent, such as power generation capacity, power generation fuel consumption, fuel price, etc. The carbon emission data is carbon emission, carbon price, and carbon emission related data. Alternatively, since the result calculated based on the data of a certain power generation may have errors, the data of power generation of multiple days can be acquired, and the errors can be eliminated by calculating the mean value or the like.
[0083] It should be noted that, in order to make the data analysis more flexible, in the embodiment of the present application, an electricity market clearing model based on agent-based modeling (ABM) is constructed, and the influence of the carbon emission cost is introduced into the model. Therefore, when the data needs to be analyzed, not only the power generation data of each target agent needs to be analyzed and acquired, but also the carbon emission data needs to be acquired.
[0084] The day-ahead market and the real-time market are considered in the model. The day-ahead market allows market participants to buy or sell electricity one day before the system operation day. Generally, the power spot market transaction of the tth day is first transacted and cleared in the day-ahead market of the t-1th day. In the day-ahead market, the power generation enterprise i makes a bid according to the power demand prediction result of the tth day provided by the grid operator , in combination with the marginal power generation cost MC i of the enterprise itself and the upper limit of the installed capacity . The power system operator clears the day-ahead market according to the bids and quantities reported by each power generation enterprise in accordance with the market clearing rules. Taking into account the technical nature of the power system, most power transactions are completed in the day-ahead market. However, due to the existence of uncertain factors such as unplanned power plant shutdowns, additional demand changes, and deviations in renewable energy output forecasts, there is a deviation between the planned power supply and demand situation and the actual power supply and demand situation when the market operates on the tth day. Therefore, the power supply and demand deviation not considered in the day-ahead market needs to be resolved through the real-time market. Transactions and clearing in the real-time market occur on the day the power system is operated. In the real-time market, power generation enterprise i will make decisions based on the actual power demand deviation in the power system. , combined with the company's own marginal power generation cost MC i and remaining power generation capacity Make a quotation and quantity Consistent with the day-ahead market, power system operators also clear the market based on the quoted volumes submitted by enterprises, following the corresponding clearing rules. Therefore, the rules of the two markets are identical; one is based on demand forecasts, while the other is based on demand deviations; in other words, the inputs differ.
[0085] Usually, power system operators usually clear the power market with the goal of minimizing market costs. The specific goals of clearing are:
[0086]
[0087]
[0088]
[0089] Where N is the total number of power generation enterprises participating in power market transactions, that is, the total number of target entities. T is the power market transaction period. t is the market clearing electricity price on day t. t is the electricity demand on day t. i,t represents the actual power generation of power generation enterprise i on day t. represents the upper limit of power generation capacity of power generation enterprise i. In addition, the clearing mechanism of the electricity market follows the priority effect.
[0090] S102: Calculate the marginal power generation cost of each target entity based on the power generation data and carbon emission data of each target entity.
[0091] Specifically, the power generation data and carbon emission data of each target entity are used to calculate the cost of power generation for each target entity. Then, by deriving the power generation of the power generation headquarters, the marginal power generation cost of the target entity is obtained, that is, the cost per kilowatt-hour of electricity produced.
[0092] Optionally, in another embodiment of the present application, a specific implementation of step S102 includes: Figure 2
[0093] S201, for each target subject, calculate the pure power generation cost of the target subject using the power generation data of the target subject.
[0094] Optionally, for a thermal power enterprise, the power generation coal consumption, fuel price and actual power generation in the power generation data can be multiplied, and then the fuel cost consumed can be obtained. Considering the self-use power of the power plant and other fixed costs, the fuel cost consumed is divided by the difference of 1 minus the self-use power rate coefficient, and finally the fixed cost is added to obtain the pure power generation cost without considering carbon emissions.
[0095] S202, judge whether the carbon emission of the target subject is greater than the carbon emission benchmark.
[0096] Since the cost needs to be paid only when the carbon emission exceeds the benchmark, it is necessary to judge whether the carbon emission of the target subject is greater than the carbon emission benchmark. If the carbon emission of the target subject is greater than the carbon emission benchmark, step S203 is executed. If the carbon emission of the target subject is not greater than the carbon emission benchmark, step S204 is executed.
[0097] S203, calculate the carbon emission cost of the target subject by multiplying the carbon price by the difference between the carbon emission of the target subject and the carbon emission benchmark.
[0098] S204, determine that the carbon emission cost of the target subject is zero.
[0099] S205, add the carbon emission cost of the target subject to the power generation cost of the target subject to obtain the total power generation cost of the target subject.
[0100] Therefore, for each power generation enterprise, the power generation cost mainly consists of two parts: fuel cost and fixed cost. At the same time, with the inclusion of the carbon emission trading system, the thermal power enterprise will receive additional carbon quota purchase cost, especially for the enterprise whose unit power emission is greater than the carbon emission benchmark. Therefore, the power generation cost of each thermal power enterprise is:
[0101]
[0102] Wherein, FC i is the power generation coal consumption of the power generation enterprise i, the higher the unit power generation coal consumption of the power plant, the higher the fuel cost. is the self-use power rate coefficient of the power plant, the higher the self-use power rate of the power plant, the lower the efficiency of the power plant, and the higher the cost. is the fuel price on the tth day. is the fixed cost of power generation enterprise i. is the actual power generation. The above data is the power generation data in the embodiment of this application. i is the unit carbon emission of power generation company i. β is the prescribed carbon emission benchmark. represents the carbon price on day t, and these data are carbon emission data.
[0103] S206: Calculate the marginal power generation cost of the target entity using the total power generation cost of the target entity.
[0104] Specifically, the total power generation cost of the target entity is derived by the actual power generation to obtain the marginal power generation cost. Therefore, the marginal power generation cost is:
[0105]
[0106] Therefore, the marginal cost of power generation enterprises is related to coal consumption, power consumption rate, fuel price and carbon cost. i <β), set e i -β is equal to 0, which means that efficient power generation companies do not incur additional carbon costs.
[0107] In order to obtain a certain profit, the target entity quotes a price based on a certain coefficient on the basis of the marginal power generation cost, so the target entity's quote is:
[0108]
[0109] Among them, it is the quotation strategy system.
[0110] Therefore, for power generation company i, its bidding strategy is to maximize corporate revenue under the installed capacity constraint:
[0111]
[0112]
[0113] S103. For each target entity, calculate the current reward of the state-action pair based on the currently updated state-action pair and the marginal power generation cost of the target entity in a loop, and use the current reward of the state-action pair to update the expected cumulative reward value of the state-action pair until the expected cumulative reward values of all state-action pairs converge.
[0114] In order to better analyze and understand the bidding behavior and strategy of each subject in the market, the reinforcement learning method is used to learn the subject-based power market transaction and clearing process in the embodiment of the present application. In the reinforcement learning process, the market subject interacts with the environment according to the time step in a certain order to maximize its cumulative income. Generally, reinforcement learning is described as a Markov decision process (MDP), which specifically includes state space, action space, state transition probability, reward function and other elements.
[0115] Among them, the state space depends on the power market environment, and in the embodiment of the present application, the state includes power demand information, clearing price and power generation. Specifically, the state space in the embodiment of the present application can be:
[0116]
[0117] Among them, D t and D t-1 are the power demand of the tth day and the (t-1)th day respectively. P t-1 is the market clearing price of the (t-1)th day. is the winning power generation of enterprise i on the (t-1)th day.
[0118] The action of the target subject in the market is to report the power generation and price. However, considering that the strategy coefficient can more accurately reflect the strategy selection of the power generation enterprise in the market, in the embodiment of the present application, the action includes the bidding strategy coefficient and the reported power generation. Therefore, the action space of the tth day can be expressed as:
[0119]
[0120] Among them, ( is the reported power generation of each power generation enterprise, and ( is the corresponding bidding strategy coefficient.
[0121] The reward function of the target subject i can be expressed as:
[0122]
[0123] Among them, r_(i,t) represents the income obtained by the power generation enterprise i in the market on the tth day.
[0124] And based on the reward function of the single-day income, the cumulative return function of the power generation enterprise i from the tth day in the whole market cycle is:
[0125]
[0126] Among them, G i,t is the cumulative return of the power generation enterprise i from the tth day. where γ is a discount factor and satisfies 0≤γ≤1, which is used to weigh the importance between current reward and future reward, ensuring that the agent can take long-term benefits into account.
[0127] Therefore, the bidding strategy of the target agent i can be defined as the conditional probability of the agent selecting an action in a given state, i.e., the probability of the agent taking action a in state s, denoted as Therefore, a state-action pair is defined to include a state and an action selected in the state.
[0128] The goal of reinforcement learning is to maximize the cumulative reward. Specifically, the expected value of the cumulative reward of a state or action pair can be evaluated based on the conditional probability through a state value function or an action value function. Therefore, the state value function (expected cumulative reward in state s) under policy is expressed as:
[0129]
[0130] The action value function (expected cumulative reward after taking action a) under policy is expressed as:
[0131]
[0132] Since in the electricity market, the main task is to analyze the corresponding competition strategy under certain market conditions, in the embodiments of the present application, learning is mainly performed through the action value function.
[0133] Considering that the electricity market is an imperfectly competitive market, there are not only independent enterprises, but also enterprises under the same large group, which will cooperate in market operation and will accordingly affect the market. Therefore, in order to accurately analyze the imperfectly competitive market, in the embodiments of the present application, an enemy-friend game framework is introduced.
[0134] Among them, the existing enemy-friend game framework divides the remaining agents into two groups for each agent, one group is regarded as the partner of the agent, and the goal is to maximize the total revenue of the partner agents, i.e., the target agents of the same group. The other group is regarded as the enemy, and the goal is to minimize the total revenue of the enemy agents. Based on this, the framework converts the general game of multiple agents into a zero-sum game of two agents. The enemy-friend Q-learning method has been proven to converge to the Nash equilibrium. In the convergence stage, the enemy-friend Q-learning method calculates the Nash equilibrium to obtain the optimal strategy:
[0135]
[0136] wherein, is the action set of the partner agent, and is the action set of all the opponent agents. In this framework, each agent i divides the remaining agents into two groups according to the camp to which it belongs. The agents belonging to the same camp are considered as partners, while the remaining agents are considered as opponents. Therefore, each power generation agent determines its bidding strategy according to the objective of maximizing the benefits of partners and minimizing the benefits of opponents.
[0137] In the friend-enemy game framework, the global reward function of the agent is:
[0138]
[0139] where the global reward function of the agent can be considered as maximizing the benefits of the own side and minimizing the benefits of the opponent . Wherein, and are the reward functions of the own agent and the opponent agent, respectively.
[0140] Therefore, when performing reinforcement learning, the cycle is based on the marginal power generation cost of the current updated state-action pair and the target agent to calculate the current reward of the state-action pair, and the expected cumulative reward value of the state-action pair is updated using the current reward of the state-action pair, until the expected cumulative reward values of all state-action pairs converge.
[0141] It should be noted that since the Q-Learning algorithm is used to solve the Markov decision problem with incomplete information. And this algorithm does not need an explicit environment model, the agent can find the optimal strategy through the experience obtained by directly interacting with the environment. Therefore, the Q-Learning algorithm is very suitable for processing decision-making problems in repeated games with unknown opponents, and therefore the embodiments of the present application select the Q-Learning algorithm to train the power generation agent in the market.
[0142] Specifically, in the learning process, at each time step t, the agent will inform the current environment state and select the corresponding action . Based on the action selected by the agent, the agent obtains the reward r t , and updates the Q value based on the reward and the environment state to , and the transition probability is .
[0143] Wherein, in the Q-learning algorithm, the value function of each state-action pair can be represented as:
[0144]
[0145] Wherein, the reward function of the target subject i is calculated.
[0146] In the embodiment of the present application, the enemy-friend game framework is introduced, so when calculating the current reward of the state-action pair, the reward of the enemy-friend pair needs to be considered. However, in the power market, the income of the opponent cannot be effectively affected. Therefore, in the embodiment of the present application, only the subject of the self party is concerned.
[0147] Therefore, in the embodiment of the present application, if the target subject belongs to the object in the group, the calculated current reward of the state-action pair is the total reward of each target subject in the current group under the state action, that is, the reward of the entire group is considered. If the target subject does not belong to the object in the group, the current reward of the state-action pair is the reward of the current target subject under the state-action pair, that is, only the reward of the individual is considered.
[0148] Specifically, each subject will update the Q value in the Q table according to the expected cumulative reward value of the state s t Taking action a t Obtaining the reward r t Update the Q value in the Q table of the self party, that is, the expected cumulative reward value. Therefore, for the target subject A not belonging to any power generation group, based on the Bellman equation, the value function of each state-action pair of the subject A The update rule of the value function
[0149]
[0150] Wherein, a represents the learning rate of the subject.
[0151] For the target subject A belonging to the group, the income of the other subjects in the group needs to be considered for measurement update. Therefore, in each state, the joint action of the subject A and the remaining subjects in the group will affect the Q value of the subject A. Therefore, for the existence of i subjects belonging to one power generation group, the Q value update rule of the subject A is:
[0152]
[0153] Wherein, The reward received by the remaining subjects in the same group.
[0154] Optionally, in another embodiment of the present application, a specific implementation of step S103 includes: Figure 3 As shown in the figure, it includes:
[0155] S301, initialize the current state and the expected cumulative reward value of each state-action pair.
[0156] S302, select the current action under the current state to determine the current state-action pair.
[0157] The current state-action pair is a state-action pair comprising the current state and the current action.
[0158] Optionally, to ensure that the agent has sufficient exploration degree for policy selection to avoid falling into a local optimum in the learning process, the agent does not directly select the action with the highest Q value, but selects the action according to a certain strategy. The market agent in this study follows an ε-greedy strategy to select an action:
[0159]
[0160] wherein ε decreases over time to simulate that the exploration degree of the agent is higher at the beginning of the action selection, so that the agent can continuously explore at the beginning and will not converge to a local optimal value. As the learning process matures, the agent's selection will tend to the optimal strategy, thereby ensuring that the optimal value can be obtained.
[0161] S303, based on the current state-action pair and the marginal generation cost of the target agent, calculate the current reward of the current state-action pair.
[0162] S304, update the expected cumulative reward value of the current state-action pair using the reward of the current state-action pair.
[0163] Optionally, in another embodiment of the present application, a specific implementation of step S304 comprises:
[0164] If the target agent does not belong to the group, multiply the marginal generation cost of the target agent by the bidding strategy coefficient in the current action to obtain the current bid of the target agent.
[0165] Multiply the difference between the current bid of the target agent and the marginal generation cost of the target agent by the reported generation capacity in the current action to obtain the current initial reward of the target agent, and multiply the current initial reward of the target agent by the first weight to obtain the current reward of the current state-action pair. The first weight is greater than 0 and less than 1.
[0166] If the target agent belongs to the group, calculate the current initial reward of each target agent in the group based on the current state-action pair and the marginal generation cost of each target agent in the group, and sum the results of multiplying the current initial reward of each target agent by the second weight to obtain the current reward of the current state-action pair. The second weight is greater than 1.
[0167] The calculation method of the current initial reward of each target agent in the group is consistent with the calculation method of the current initial reward of the target agent that does not belong to the group.
[0168] Therefore, in the embodiment of the present application, two adjustment coefficients are added in the existing reward function in combination with the idea of the friend or foe game framework, and thus the reward function of each agent is:
[0169]
[0170] wherein, and are the first weight and the second weight respectively, and represent the weight adjustment coefficients of the partner agent and the opponent agent respectively. It means that the income of cooperation within the same group is greater than the income of only considering the optimal strategy of itself, and encourages cooperation within the same group. It means that the agent not belonging to the power generation group needs to take a more aggressive market strategy to obtain more income.
[0171] S305, updating the current state.
[0172] S306, judging whether the expected cumulative reward values of each state-action pair converge.
[0173] If it is judged that the expected cumulative reward values of each state-action pair do not converge, step S302 is executed. If it is judged that the expected cumulative reward values of each state-action pair converge, step S307 is executed.
[0174] S307, ending the loop.
[0175] S104, determining the action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state as the best behavior strategy of the target agent in each state.
[0176] Since the Q function gives the value of the agent after taking action a in state s, the function considers the future income. The goal of the agent is to find the optimal strategy by maximizing the total reward it receives in each state s:
[0177]
[0178] And the Q values of all state-action pairs constitute the Q table. Each agent will gradually converge to the optimal strategy according to the Q table and the selection strategy Therefore, the action in the state-action pair with the maximum expected cumulative reward value in each state-action pair of each state is determined as the best behavior strategy of the target agent in each state.
[0179] Alternatively, in another embodiment of the present application, after step S104 is executed, an information simulation aid can be applied. Specifically, an information simulation aid method provided in the embodiment of the present application, as shown in Figure 4 includes:
[0180] S401, obtaining information of a current state of any target subject.
[0181] The information of the current state can include power demand information, clearing price and power generation.
[0182] Specifically, when it is needed to determine the optimal behavior strategy of any target subject in any state, the information of the current state of the target subject can be input.
[0183] S402, determining the optimal behavior strategy of the current state of the target subject based on the information of the current state of the target subject and the optimal behavior strategy of the target subject in each state.
[0184] Alternatively, the optimal behavior strategy of the current state of the target subject can be directly obtained from the optimal behavior strategy of the target subject in each state.
[0185] If the optimal behavior strategy corresponding to the current state does not exist, it can be determined according to the optimal behavior strategy of the current state of the target subject by interpolation or other ways.
[0186] S403, calculating the current bid price by using the bid price coefficient in the optimal behavior strategy.
[0187] S404, feeding back the optimal behavior strategy and the current bid price.
[0188] The embodiment of the application provides a multi-agent bidding behavior analysis method based on an enemy and friend game framework. A power market clearing model based on agent modeling is used to analyze each agent more flexibly. The influence mechanism of carbon cost on enterprise bidding is also introduced into the model. Therefore, when analysis is required, the power generation data and carbon emission data of each target agent are obtained. Based on the power generation data and carbon emission data of each target agent, the marginal power generation cost of the target agent is calculated. Then, for each target agent, the current reward of the state-action pair is calculated based on the updated current state-action pair and the marginal power generation cost of the target agent. The expected cumulative reward value of the state-action pair is updated using the current reward of the state-action pair until the expected cumulative reward values of all state-action pairs converge. A state-action pair includes a state and an action selected in the state. The state includes power demand information, clearing price and power generation. The action includes a bidding strategy coefficient and reported power generation. Thus, the strategy selection of the enterprise is optimized through reinforcement learning. Moreover, if the target agent belongs to the object in the group, the current reward of the state-action pair is the total reward of each target agent in the state action in the current group. If the target agent does not belong to the object in the group, the current reward of the state-action pair is the reward of the target agent in the state-action pair, thereby introducing the enemy and friend game framework into the ABM-based power market bidding behavior simulation algorithm to better adapt to the imperfect competition market. Finally, the action in the state-action pair with the maximum expected cumulative reward value in each state of each state-action pair is determined as the best behavior strategy of the target agent in each state, thereby realizing accurate analysis of the data of the target agent and obtaining accurate analysis results to better assist the enterprise in action.
[0189] In order to verify the method provided by the embodiment of the application, data is obtained from each power trading center. The frequency of the electricity price data is 15 minutes, which means that there are 96 real-time electricity price data per day. The data starts from the starting day of the long-period settlement trial run of the electricity spot market in A, B and B, so that the obtained electricity price data contain a large amount of data.
[0190] The carbon price is the core explanatory variable in empirical research. The daily carbon price data of the national carbon market can be obtained from the Wind database. For example, the carbon price in the time range of July 16, 2021 to August 31, 2024 is obtained, as shown in Figure 5
[0191] Fuel cost is one of the main sources of marginal cost of generating units and is considered as one of the key factors in the bidding strategy of units. Therefore, this paper takes fuel price as an explanatory variable and includes it in the regression. Given that almost all thermal power units in the three places are coal-fired or gas-fired units, the coal price and natural gas price data matching the time range of the electricity price are collected. Unlike most studies that use futures coal price data, this paper chooses the Hong Kong coal daily arrival price as the coal price variable. This is mainly because after the coal price soared in 2021-2022, the futures coal price data stopped updating from September 2022, and subsequent data cannot reflect the trend of coal price changes. Since one-third of the thermal power plants in C place are gas-fired, this paper further collects foreign natural gas futures market prices to represent the trend of natural gas prices. From Figure 6 It can be seen that the coal price and natural gas price both show significant fluctuations. Under the background of global economic recovery and significant increase in energy demand, coal prices were high during 2021-2022. The coal price in Hong Kong once reached nearly 2500 yuan / ton, almost 3 times the normal coal price level. Until 2023, the coal price gradually fell to around 1000 yuan / ton. The natural gas price rose significantly in 2022, reaching about 10 yuan / million British thermal units. With the subsequent recovery of stable supply, the natural gas price gradually fell to 2-4 yuan / million British thermal units.
[0192] Referring to the empirical research of relevant literature, the power demand and renewable energy output are considered as explanatory variables. The power demand and renewable energy output data are also from relevant websites, and the data collection time range matches the time range of the real-time electricity price data. Similar to real-time electricity price data, the data frequency of power demand and renewable energy output is also 15 minutes. Figure 7 The change trend of power demand and renewable energy output in A place, B place and B place after daily average is shown. According to Figure 7 The upper part of the area shows that the daily average load of A place, B place and B place is about 60000 megawatts, 30000 megawatts and 80000 megawatts respectively. Figure 7 The lower part of the area in represents the daily average renewable energy output of the three places. Renewable energy generation can cover about 10%-20% of the total load of each province.
[0193] While thermal power generation is still one of the most important types of power generation. Thermal power units are also the main participants in the electricity market and the carbon market. Thermal power generation in the three places accounts for about 70%-80% of the total power generation in the region, and is the most important source of electricity in the three places. The relevant data of unit level in the three places are sorted out.
[0194] Among them, A has the largest number of units, a total of 478 thermal power generating units. B and C each have 180 thermal power generating units. The thermal power structure of A and B is similar, about 80% of the units have a capacity of less than or equal to 300 megawatts, these units have lower efficiency and higher carbon emissions. However, compared with large units (>300 megawatts), these small units (≤300 megawatts) bear more power generation in A and B. A and B have almost no gas units, according to statistics, only B contains 6 gas units. The power structure of C is significantly different from that of A and B. In terms of coal-fired units, C has a similar number of large units (>300 megawatts) and small units (≤300 megawatts), and the annual total power generation of large units is higher. In addition, C has more gas power generating units, accounting for about one-third of the total number of generating units in the whole land, and bears one-seventh of the total power generation. Among them, if there are missing values in the collected data, the interpolation method is used to fill in the data.
[0195] And Figure 8 The results show that the real-time market and the day-ahead market clearing price in the simulation results. The results show that the real-time price of the three places is significantly higher than the day-ahead price, about 200 yuan / MWh to 800 yuan / MWh higher on average. This is mainly due to the fact that most of the electricity has been cleared in the day-ahead market, and the remaining adjustable electricity in the real-time market is usually priced higher. In addition, the real-time price is less volatile than the day-ahead price. Because the real-time market is mainly used to adjust the supply and demand deviation caused by factors such as load forecasting error and renewable energy output uncertainty, the small trading volume leads to limited price adjustment space for real-time prices.
[0196] And further analysis of the bidding strategy of power generation enterprises in the day-ahead and real-time markets. Figure 9 The average daily bidding curves of five large power generation groups and non-large power generation groups are shown. Among them, the bidding strategy of large power generation groups will change greatly in different time periods, while the bidding curve of non-large power generation groups is relatively flat, and the bidding of enterprises in different time periods tends to be consistent. This result directly reflects the difference in bidding strategies between groups and non-groups. Because the market share of large power generation groups is high, groups can influence the market clearing results by bidding and bidding, so power generation groups will adjust their bidding according to the supply and demand in the market to maximize profits. The power generation enterprises not belonging to any group are similar to the followers in the oligopoly market, their market share is small, the bidding space is limited, and the flexibility of strategy adjustment is low. Therefore, these enterprises often give similar bids based on market prices, and the bidding strategy selection is conservative.
[0197] In addition, the bidding strategies of different power generation groups are also different, such as Figure 10The bidding strategies of the groups are highly related to their marginal generation cost and market share. The groups with higher market share are more likely to get marginal pricing rights, and thus tend to bid high in the market. On the contrary, the groups with lower market share will choose more conservative bidding strategies to get stable market clearing shares. Most groups will choose to bid high in one market. The strategy of bidding low in the other market is to achieve the balance between revenue and risk. The groups with more gas power plants tend to bid high in the real-time market to take advantage of their more flexible peak shaving capability to get more profits. For the non-group power generation enterprises, bidding around the expected market clearing price is their optimal bidding strategy.
[0198] Overall, large power generation groups usually bid low in the day-ahead market to expand market share and ensure the stability of long-term revenue. In the real-time market, large power generation groups tend to increase their bids to take advantage of the opportunity of tight power supply and demand to get higher profits. Due to the characteristics of the power export province and cheaper coal prices, the strategy of large group enterprises in B is just the opposite of that in A and C. Further, this study explores the impact of power market and carbon market on enterprise bidding behavior by comparing the relationship between the average bid of the group and the average cost. Due to differences in factors such as unit efficiency and market share, the bidding strategies of different power generation groups also differ. Overall, groups with higher market share have more influence in the market and are more likely to adopt aggressive bidding strategies in the day-ahead or real-time market, while groups with lower market share choose more conservative strategies.
[0199] Another embodiment of the present application provides a multi-agent bidding behavior analysis system based on a friend-or-foe game framework, as shown in Figure 11 The system comprises:
[0200] A first acquisition unit 1101 is configured to acquire power generation data and carbon emission data of each target agent.
[0201] A cost calculation unit 1102 is configured to calculate the marginal generation cost of each target agent based on the power generation data and carbon emission data of the target agent, respectively.
[0202] A learning unit 1103 is configured to, for each target agent, calculate the current reward of a state-action pair based on the currently updated state-action pair and the marginal generation cost of the target agent, and update the expected cumulative reward value of the state-action pair using the current reward of the state-action pair, until the expected cumulative reward values of all state-action pairs converge.
[0203] The state-action pair includes a state and an action selected in the state. The state includes power demand information, a clearing price, and power generation. The action includes a bidding strategy coefficient and reported power generation. If the target subject belongs to the objects in the group, the current reward of the state-action pair is the total reward of each target subject in the current group in the state-action pair. If the target subject does not belong to the objects in the group, the current reward of the state-action pair is the reward of the target subject in the state-action pair.
[0204] The behavior determination unit 1104 is configured to determine, in each state-action pair of each state, the action in the state-action pair with the maximum expected cumulative reward value as the optimal behavior strategy of the target subject in each state.
[0205] Optionally, in the multi-agent bidding behavior analysis system based on the enemy-friend game framework provided in another embodiment of the present application, the system comprises:
[0206] The second acquisition unit is configured to acquire information of the current state of any target subject.
[0207] The strategy determination unit is configured to determine the optimal behavior strategy of the current state of the target subject based on the information of the current state of the target subject and the optimal behavior strategy of the target subject in each state.
[0208] The first bidding calculation unit is configured to calculate the current bid by using the bidding strategy coefficient in the optimal behavior strategy.
[0209] The information feedback unit is configured to feed back the optimal behavior strategy and the current bid.
[0210] Optionally, in the multi-agent bidding behavior analysis system based on the enemy-friend game framework provided in another embodiment of the present application, the cost calculation unit comprises:
[0211] The first calculation unit is configured to calculate the pure power generation cost of each target subject by using the power generation data of the target subject.
[0212] The second calculation unit is configured to calculate the carbon emission cost of the target subject by multiplying the carbon price by the difference between the carbon emission amount of the target subject and the carbon emission benchmark when the carbon emission amount of the target subject is greater than the carbon emission benchmark.
[0213] The third calculation unit is configured to determine that the carbon emission cost of the target subject is zero when the carbon emission amount of the target subject is not greater than the carbon emission benchmark.
[0214] The fourth calculation unit is configured to add the carbon emission cost of the target subject to the power generation cost of the target subject to obtain the total power generation cost of the target subject.
[0215] The cost determination unit is configured to calculate the marginal power generation cost of the target subject by using the total power generation cost of the target subject.
[0216] Optionally, in the multi-agent bidding behavior analysis system based on the friend-or-foe game framework provided in another embodiment of the present application, the learning unit comprises:
[0217] The initialization unit is configured to initialize the current state and the expected cumulative reward value of each state-action pair.
[0218] The selection unit is configured to select a current action in the current state to determine a current state-action pair. The current state-action pair is a state-action pair comprising the current state and the current action.
[0219] The reward calculation unit is configured to calculate a current reward of the current state-action pair based on the current state-action pair and the marginal power generation cost of the target subject.
[0220] The reward value updating unit is configured to update the expected cumulative reward value of the current state-action pair by using the reward of the current state-action pair.
[0221] The state updating unit is configured to update the current state and return to execute the selection unit until the expected cumulative reward values of the state-action pairs converge.
[0222] Optionally, in the multi-agent bidding behavior analysis system based on the friend-or-foe game framework provided in another embodiment of the present application, the reward calculation unit comprises:
[0223] The second offer calculation unit is configured to, when the target subject does not belong to the group, multiply the marginal power generation cost of the target subject by the offer strategy coefficient in the current action to obtain a current offer of the target subject.
[0224] The first reward calculation unit is configured to multiply the difference between the current offer of the target subject and the marginal power generation cost of the target subject by the reported power generation in the current action to obtain a current initial reward of the target subject.
[0225] The first weighting unit is configured to multiply the current initial reward of the target subject by a first weight to obtain the current reward of the current state-action pair. The first weight is greater than 0 and less than 1.
[0226] The second reward calculation unit is configured to, when the target subject belongs to the group, calculate a current initial reward of each target subject in the group based on the current state-action pair and the marginal power generation cost of each target subject in the group.
[0227] The second weighting unit is configured to sum the results of multiplying the current initial rewards of the target subjects by a second weight to obtain the current reward of the current state-action pair. The second weight is greater than 1.
[0228] It should be noted that the specific working process of each unit provided by the above-mentioned embodiments of the present application can be correspondingly referred to the implementation process of the corresponding steps in the above-mentioned method embodiments, which will not be described here.
[0229] Those skilled in the art will further appreciate that the individual steps of the example methods and algorithms described in connection with the embodiments disclosed herein can be embodied in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their general application to the functional operations of the examples. It will be apparent to those skilled in the art that the functions required to achieve the examples described herein can be implemented in software or hardware or a combination thereof, and that the examples described herein can be implemented by one or more computer programs executed on one or more programmable computing devices to perform the functions described herein. The various steps of the examples described herein can be performed in the order listed, or in any other order, or in parallel, or in any combination thereof.
[0230] The above description of the disclosed embodiments allows a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the examples shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-agent bidding behavior analysis method based on the friend-enemy game framework, characterized in that: include: Obtain power generation data and carbon emission data of each target entity; Calculating the marginal power generation cost of the target entity based on the power generation data and carbon emission data of each target entity; For each target entity, respectively, cyclically calculate the current reward of the state-action pair based on the currently updated state-action pair and the marginal power generation cost of the target entity, and use the current reward of the state-action pair to update the expected cumulative reward value of the state-action pair until the expected cumulative reward values of all the state-action pairs converge; Wherein, a state-action pair includes a state and an action selected under the state; the state includes electricity demand information, clearing price, and power generation; the action includes the bidding strategy coefficient and the reported power generation; if the target subject is an object in a group, the current reward of the state-action pair is the total reward of each target subject in the current group under the state action; if the target subject is not an object in the group, the current reward of the state-action pair is the current reward of the target subject under the state-action pair; Among the state-action pairs of each state, the action in the state-action pair with the largest expected cumulative reward value is used to determine the optimal behavior strategy of the target subject in each state.
2. The method according to claim 1, characterized in that After determining the optimal behavior strategy of the target subject in each state by taking the action in the state-action pair with the largest expected cumulative reward value among the state-action pairs in each state, the method includes: Obtaining information about the current status of any one of the target entities; Determining the optimal behavior strategy for the current state of the target subject based on information about the current state of the target subject and the optimal behavior strategy of the target subject in each of the states; Calculating the current quote using the quote strategy coefficient in the best behavior strategy; The optimal action strategy and the current quote are fed back.
3. The method according to claim 1, characterized in that The calculating of the marginal power generation cost of the target entity based on the power generation data and carbon emission data of each target entity includes: For each target entity, respectively, calculate the net power generation cost of the target entity using the power generation data of the target entity; If the target entity's carbon emissions are greater than the carbon emission benchmark, the carbon price is multiplied by the difference between the target entity's carbon emissions and the carbon emission benchmark to obtain the target entity's carbon emission cost; If the carbon emissions of the target entity are not greater than the carbon emission benchmark, the carbon emission cost of the target entity is determined to be zero; Adding the target entity's power generation cost to the target entity's carbon emission cost to obtain the target entity's total power generation cost; The marginal power generation cost of the target entity is calculated using the total power generation cost of the target entity.
4. The method according to claim 1, wherein The loop calculates the current reward of the state-action pair based on the currently updated state-action pair and the marginal power generation cost of the target entity, and updates the expected cumulative reward value of the state-action pair using the current reward of the state-action pair until the expected cumulative reward values of all the state-action pairs converge, including: Initialize the current state and the expected cumulative reward value of each state-action pair; Selecting a current action in the current state to determine a current state-action pair; wherein the current state-action pair is a state-action pair including the current state and the current behavior; Calculating a current reward for the current state-action pair based on the current state-action pair and the marginal power generation cost of the target entity; Using the reward of the current state-action pair, updating the expected cumulative reward value of the current state-action pair; The current state is updated, and the current action under the current state is selected and executed again to determine the current state-action pair until the expected cumulative reward values of the state-action pairs converge.
5. The method according to claim 4, characterized in that The calculating of the current reward of the current state action pair based on the current state action pair and the marginal power generation cost of the target entity includes: If the target entity does not belong to the group, multiplying the bidding strategy coefficient in the current action by the marginal power generation cost of the target entity to obtain the current bidding of the target entity; The current initial reward of the target entity is obtained by multiplying the difference between the target entity's current bid and the target entity's marginal power generation cost by the power generation reported in the current action; Multiplying the current initial reward of the target subject by a first weight to obtain the current reward of the current state-action pair; wherein the first weight is greater than 0 and less than 1; If the target entity belongs to a group, calculating the current initial reward of each target entity based on the current state-action pair and the marginal power generation cost of each target entity in the group; The current reward of the current state-action pair is obtained by multiplying the current initial reward of each target subject by the second weight and summing the results; wherein the second weight is greater than 1.
6. A multi-agent bidding behavior analysis system based on the friend-enemy game framework, characterized in that: include: A first acquisition unit is used to acquire power generation data and carbon emission data of each target entity; a cost calculation unit, configured to calculate the marginal power generation cost of the target entity based on the power generation data and carbon emission data of each target entity; a learning unit, configured to calculate, for each target entity, a current reward of the state-action pair based on a currently updated state-action pair and the marginal power generation cost of the target entity, and update the expected cumulative reward value of the state-action pair using the current reward of the state-action pair until the expected cumulative reward values of all the state-action pairs converge; Wherein, a state-action pair includes a state and an action selected under the state; the state includes electricity demand information, clearing price, and power generation; the action includes the bidding strategy coefficient and the reported power generation; if the target subject is an object in a group, the current reward of the state-action pair is the total reward of each target subject in the current group under the state action; if the target subject is not an object in the group, the current reward of the state-action pair is the current reward of the target subject under the state-action pair; The behavior determination unit is used to determine the optimal behavior strategy of the target subject in each state by taking the action in the state-action pair with the largest expected cumulative reward value in each state-action pair.
7. The system according to claim 6, characterized in that include: A second acquiring unit, configured to acquire information about the current state of any one of the target entities; a strategy determination unit, configured to determine an optimal behavior strategy for the current state of the target subject based on information about the current state of the target subject and the optimal behavior strategy of the target subject in each of the states; A first quote calculation unit, configured to calculate a current quote using a quote strategy coefficient in the optimal behavior strategy; The information feedback unit is used to feed back the optimal behavior strategy and the current quotation.
8. The system according to claim 6, wherein: The cost calculation unit includes: a first calculation unit, configured to calculate the net power generation cost of each target entity using the power generation data of the target entity; a second calculation unit, configured to calculate, when the carbon emissions of the target entity are greater than the carbon emission benchmark, the carbon price multiplied by the difference between the carbon emissions of the target entity and the carbon emission benchmark to obtain the carbon emission cost of the target entity; a third calculation unit, configured to determine that the carbon emission cost of the target entity is zero if the carbon emission amount of the target entity is not greater than the carbon emission benchmark; a fourth calculation unit, configured to add the power generation cost of the target entity to the carbon emission cost of the target entity to obtain the total power generation cost of the target entity; The cost determination unit is configured to calculate the marginal power generation cost of the target entity by using the total power generation cost of the target entity.
9. The system according to claim 6, wherein: The learning unit includes: Initialization unit, used to initialize the current state and the expected cumulative reward value of each state-action pair; a selection unit, configured to select a current action in the current state and determine a current state-action pair; wherein the current state-action pair is a state-action pair including the current state and the current behavior; a reward calculation unit, configured to calculate a current reward for the current state-action pair based on the current state-action pair and the marginal power generation cost of the target entity; a reward value updating unit, configured to update the expected cumulative reward value of the current state-action pair using the reward of the current state-action pair; A state updating unit is used to update the current state and return to execute the selection unit until the expected cumulative reward values of each state-action pair converge.
10. The system according to claim 9, characterized in that The reward calculation unit includes: a second bid calculation unit, configured to, when the target entity does not belong to a group, multiply the bid strategy coefficient in the current action by the marginal power generation cost of the target entity to obtain a current bid of the target entity; a first reward calculation unit, configured to multiply the difference between the current bid of the target entity and the marginal power generation cost of the target entity by the power generation reported in the current action to obtain the current initial reward of the target entity; A first weighting unit, configured to multiply the current initial reward of the target subject by a first weight to obtain a current reward of the current state-action pair; wherein the first weight is greater than 0 and less than 1; a second reward calculation unit, configured to calculate, when the target entity belongs to a group, a current initial reward of each target entity based on the current state-action pair and the marginal power generation cost of each target entity in the group to which it belongs; A second weighting unit is used to sum the results of multiplying the current initial reward of each target subject by a second weight to obtain the current reward of the current state-action pair; wherein the second weight is greater than 1.